01
Sees
Reads the message and any images together, alongside the conversation, your assistant’s policy, and the current task state.
Text · Photos · Screenshots · Scanned documents
TOD · The universal decision model
TOD reads text and images, extracts the facts, and picks the next action. It sits in front of your existing assistant, takes the turns it is confident about, and hands the rest to your LLM.
Input
Text + images
Policy + task state
Available actions
Look up · Reply · Write
Your allowed choices
TOD
~400M parameters
Extract the facts. Rank the next actions.
Confidence gate
Confident → Serve
Take the selected action.
Skip the parent call.
Uncertain → Defer
Your existing LLM takes over.
Pass the turn unchanged.
The benchmark · Held-out data TOD never trained on
TOD starts from an open 12-billion-parameter model and is trained on our own decision data. Training taught it to read images, choose among hundreds of options, and report confidence you can act on.
100 held-out questions per set. Removing the image costs TOD 17 to 25 points, so it is reading the pixels, not guessing from the text.
Most pickers read answers as letters, which caps them at 26 options.
200 held-out decisions per set. Training improved accuracy by 6 to 21 points on each set and cut calibration error from 23.1% to 5.0%.
| Model | Accuracy | Easy / original / hard | Calibration error |
|---|---|---|---|
| Untrained base model | 86.2% | 100% / 94.4% / 74.8% | 3.6% |
| TOD | 84.4% | 100% / 91.7% / 73.0% | 3.3% (0.8% raw) |
On public JevBench, training cost TOD 4 of 231 decisions against the untrained base, all in general-reasoning questions. We ship the trained model because it adds image understanding, any number of options, and honest confidence. Recovering those decisions is planned for the next release.
JevBench was never used for training or calibration. Confidence temperatures were fit on a separate validation set. Calibration error is expected calibration error; “raw” is before any temperature correction.
TOD reads up to 48,000 tokens of context per decision and generates no text, so you pay for input only.
Source: TOD picker v1 release evaluation, 29 September 2026. The untrained base model is the open Gemma 4 12B model TOD is built on, scored the same way.
One model, any input
A photo can explain what a message leaves unsaid. TOD is built to read both together and decide what happens next. Customer support is the showcase; the decision model is designed to transfer across tasks.
01
Reads the message and any images together, alongside the conversation, your assistant’s policy, and the current task state.
Text · Photos · Screenshots · Scanned documents
02
Pulls out the facts that matter: the order, the item, the issue, the amount. A separate step resolves action arguments from the conversation and provider look-ups.
Unstructured input → Structured facts
03
Ranks the actions available on this turn. It chooses from that set, without inventing actions or writing free text, and only serves the turn when confident.
Look up · Select a reply · Choose a write
Customer input
“My boots arrived like this. Can I get a size-10 replacement?”
Structured facts
Next actions
Illustrates text-and-image capability. The measured benchmark below uses text conversations only, and all write operations are deferred to the parent model.
Measured on text support conversations · tau2-bench retail
Purpose-built to decide
A compact encoder of roughly 400 million parameters scores the available actions in two steps: which kind of action fits, then which candidate is best. A neural memory training objective teaches it to track where the conversation stands, with no extra serving cost.
The confidence gate makes the final call. Served turns skip the parent LLM entirely. Deferred turns pass through untouched to the model you already use.
An offline check on held-out decisions, using the same conversation and candidate actions.
Exact match means the identical action or argument values. These are decision-level measurements, separate from whole-task solve rates.
Per-decision pricing
A flat price for each decision, regardless of context length. Your existing model handles the turns TOD defers.
TOD fees are separate from the parent-model call costs reported in the benchmark above.
$0.30
per 1,000 decisions · $300 per million
25% lower decision price on the evaluated workload.
JEV’s token pricing works out to $0.40 per 1,000 decisions on this workload. TOD’s price does not rise with context length.
Sign up and try TOD on your own support conversations. Find the operating point that fits your model, your workload, and your quality bar.