TOD · The universal decision model

Small Model.
Clear Decisions.

TOD reads text and images, extracts the facts, and picks the next action. It sits in front of your existing assistant, takes the turns it is confident about, and hands the rest to your LLM.

Input

Text + images
Policy + task state

Available actions

Look up · Reply · Write
Your allowed choices

TOD

~400M parameters

Extract the facts. Rank the next actions.

Confidence gate

Confident → Serve

Take the selected action.
Skip the parent call.

Uncertain → Defer

Your existing LLM takes over.
Pass the turn unchanged.

One assistant experience. A decision on every turn.

The benchmark · Held-out data TOD never trained on

Sees the image.
Weighs every option.

TOD starts from an open 12-billion-parameter model and is trained on our own decision data. Training taught it to read images, choose among hundreds of options, and report confidence you can act on.

Public JevBench
84.4%
of 231 decisions correct
Easy 100% · original 91.7% · hard 73.0%.
Chart questions
97%
correct with the image
Up from 70% for the untrained base model.
Options per decision
300
without collapse
55% correct where chance is 0.3%. Typical pickers stop at 26.
Calibration error
0.8%
before any correction
When TOD says 90%, it is right about 90% of the time.
Accuracy on image questions · Higher is better
Untrained base modelTOD
Science questions · ScienceQA
Untrained base model:83%TOD:90%
TOD without the image: 73%
Diagrams · AI2D
Untrained base model:80%TOD:87%
TOD without the image: 62%
Charts · ChartQA
Untrained base model:70%TOD:97%
TOD without the image: 79%

100 held-out questions per set. Removing the image costs TOD 17 to 25 points, so it is reading the pixels, not guessing from the text.

Accuracy with many options · Higher is better

Most pickers read answers as letters, which caps them at 26 options.

Banking requests · 50 options
76%
Chance: 2.0%
Assistant requests · 50 options
93%
Chance: 2.0%
Legal provisions · 100 options
72.5%
Chance: 1.0%
Mixed task labels · 174 options
53%
Chance: 0.6%
Synthetic slate · 300 options
55%
Chance: 0.3%

200 held-out decisions per set. Training improved accuracy by 6 to 21 points on each set and cut calibration error from 23.1% to 5.0%.

Detailed results and methodology
Public JevBench · 231 decisions
ModelAccuracyEasy / original / hardCalibration error
Untrained base model86.2%100% / 94.4% / 74.8%3.6%
TOD84.4%100% / 91.7% / 73.0%3.3% (0.8% raw)

On public JevBench, training cost TOD 4 of 231 decisions against the untrained base, all in general-reasoning questions. We ship the trained model because it adds image understanding, any number of options, and honest confidence. Recovering those decisions is planned for the next release.

JevBench was never used for training or calibration. Confidence temperatures were fit on a separate validation set. Calibration error is expected calibration error; “raw” is before any temperature correction.

TOD reads up to 48,000 tokens of context per decision and generates no text, so you pay for input only.

Source: TOD picker v1 release evaluation, 29 September 2026. The untrained base model is the open Gemma 4 12B model TOD is built on, scored the same way.

One model, any input

Sees. Extracts. Acts.

A photo can explain what a message leaves unsaid. TOD is built to read both together and decide what happens next. Customer support is the showcase; the decision model is designed to transfer across tasks.

01

Sees

Reads the message and any images together, alongside the conversation, your assistant’s policy, and the current task state.

Text · Photos · Screenshots · Scanned documents

02

Extracts

Pulls out the facts that matter: the order, the item, the issue, the amount. A separate step resolves action arguments from the conversation and provider look-ups.

Unstructured input → Structured facts

03

Acts

Ranks the actions available on this turn. It chooses from that set, without inventing actions or writing free text, and only serves the turn when confident.

Look up · Select a reply · Choose a write

Support in action · Illustration, not a benchmark result

Customer input

“My boots arrived like this. Can I get a size-10 replacement?”

Attached photo
Damaged parcel

Structured facts

Item
Boots
Size
10
Issue
Damaged in transit
Request
Replacement

Next actions

  1. 01Look up the order and replacement.
  2. 02Select the replacement action and resolve its arguments.

Illustrates text-and-image capability. The measured benchmark below uses text conversations only, and all write operations are deferred to the parent model.

Measured on text support conversations · tau2-bench retail

Pro + TOD aggressive
25.8%
lower parent-model cost
92 of 114 tasks solved, versus 96 with Pro alone. TOD serves 58.0% of tool decisions.
Flash + TOD conservative
23.3%
lower parent-model cost
101 of 114 tasks solved, versus 102 with flash alone. TOD serves 30.8% of tool decisions.
TOD decision latency
69 ms
median model forward pass
121 ms at p95, measured on one H100 80GB in bf16. Model timing, before provider look-ups.

Purpose-built to decide

A smaller model.
A focused job.

A compact encoder of roughly 400 million parameters scores the available actions in two steps: which kind of action fits, then which candidate is best. A neural memory training objective teaches it to track where the conversation stands, with no extra serving cost.

The confidence gate makes the final call. Served turns skip the parent LLM entirely. Deferred turns pass through untouched to the model you already use.

Exact replication of a frontier reference

An offline check on held-out decisions, using the same conversation and candidate actions.

Action chosen
63.5%
582 / 916 decisions
Arguments resolved
94.7%
657 / 694 argument sets

Exact match means the identical action or argument values. These are decision-level measurements, separate from whole-task solve rates.

Per-decision pricing

One decision.
One flat price.

A flat price for each decision, regardless of context length. Your existing model handles the turns TOD defers.

TOD fees are separate from the parent-model call costs reported in the benchmark above.

$0.30

per 1,000 decisions · $300 per million

25% lower decision price on the evaluated workload.

JEV’s token pricing works out to $0.40 per 1,000 decisions on this workload. TOD’s price does not rise with context length.

Bring TOD to Your Support Desk.

Sign up and try TOD on your own support conversations. Find the operating point that fits your model, your workload, and your quality bar.