Decide Live

On-device decisions from camera and microphone. Every frame is embedded in your browser by EmbeddingGemma 2, fine-tuned for decisions, and scored against the options you write. No text generation, no tokens, no server.

Model not loaded
0 bytes uploaded
Runs on your device. 0 bytes uploaded.
0.0 fps | 0 ms / frame

Loads the model on first start (about 260 to 290 MB at 4 bits, cached afterwards).

0decided on device
0would escalate
-on-device rate

When a card's top probability is below its threshold the frame is marked uncertain: in a production system it would go to a larger model. Here nothing is sent anywhere; it is only counted.

Decisions on every frame

Add a decision (up to 4)

1. Input

2. Question and options

Settings

Egg check

Candling photos (a light shone through the egg in a dark room). The decision uses the format the model was trained with: the photo with Is this egg good or bad? and the candling context, scored against Normal, healthy egg and Defective, spoiled or damaged egg. The 8 photos below are held out: they were not used in training. Pick one, run all 8, or use your own photo.

Photos: Good and Bad Eggs Identification Image Dataset (Mendeley Data, doi:10.17632/mdty358x8m.1), CC BY 4.0. See CREDITS.md.

Sample tests (18)

Counting, yes / no, spatial relations, a chart, a receipt, app screens, scene text, common sense and egg candling (the egg photos are held out from training; the other images may come from datasets the training mix drew on). Each test runs in this browser with the prompt format, image detail and temperature chosen under Single input > Settings. In the trained format, yes / no tests keep the question in the query and score the options "Yes" and "No"; the older formats use the yes / no mode (the question followed by "Yes." or "No.").

How it works

  1. The input (a camera frame, an image, a text, a 2 second sound window or a short video) and the question are turned into one 768-number vector (mean pooling, then L2 normalization) by EmbeddingGemma 2 (Apache-2.0), fine-tuned for decisions: the question and the media go into one embedding and each option into another, and training pushes the right option closest. It runs as ONNX (same file layout as onnx-community/embeddinggemma-2-ONNX, with the fine-tuned weights) with transformers.js 4.3.1 on WebGPU (or WASM on the CPU). The model files are served with this page.
  2. Each option description is embedded once and cached.
  3. The probabilities are softmax(scale × cosine) over the options. The scale was learned in training (about 19.96), so the default temperature is 1 / scale, about 0.050.
  4. In Live mode each card embeds the frame together with its own question (one pass per card). The alternative "frame alone" setting embeds the frame once and scores all cards from that vector, which is faster but is not the format the model was trained with.
  5. A card whose top probability is below its threshold is marked uncertain. That is where a larger model would take over; this page only counts it.

Privacy

The page is static. The model weights are downloaded from the Hugging Face Hub and cached by the browser. Camera frames, microphone audio, images and text are processed on this device and are never uploaded. You can check this in the browser's network panel: after the model is cached, decisions make no requests.

Prompts

Limits

Sample image credits: CREDITS.md. Code: Apache-2.0.