Parseclab has released TOD, a small decision model that sits in front of an existing LLM and handles the turns it is confident about, deferring the rest. Built on an open 12-billion-parameter Gemma 4 base, TOD reads text and images together, extracts structured facts, and chooses from a provided set of actions without generating free text. The model processes up to 48,000 tokens of context per decision and generates no output tokens, so pricing is based on input only.
How TOD Fits Into an Assistant Stack
TOD acts as a router. It receives the incoming message, any images, the conversation history, the assistant policy, and the current task state. It then pulls out the facts that matter — the order, the item, the issue, the amount — and resolves action arguments from the conversation and provider look-ups. A compact encoder of roughly 400 million parameters scores available actions in two steps: which kind of action fits, then which candidate is best. A neural memory training objective teaches it to track where the conversation stands with no extra serving cost.
A confidence gate makes the final call. Served turns skip the parent LLM entirely. Deferred turns pass through untouched to the model you already use. This architecture means TOD fees are separate from the parent-model call costs.
Benchmark Results on Held-Out Data
Public JevBench, 231 decisions, never used for training or calibration. The untrained base model scored 86.2 percent accuracy with a 3.6 percent calibration error. TOD scored 84.4 percent accuracy with a 3.3 percent calibration error (0.8 percent raw before temperature correction). Training cost TOD four of 231 decisions against the untrained base, all in general-reasoning questions. The company ships the trained model because it adds image understanding, any number of options, and honest confidence. Recovering those four decisions is planned for the next release.
On the tau2-bench retail workload for text support conversations, exact-match measurements show decision-level performance. These are separate from whole-task solve rates. The benchmark source is the TOD picker v1 release evaluation dated 29 September 2026.
Pricing and Trade-Offs
Flat per-decision pricing: $0.30 per 1,000 decisions ($300 per million). The price does not rise with context length. JEV's token pricing works out to $0.40 per 1,000 decisions on the evaluated workload, making TOD roughly 25 percent lower on that workload. Your existing model handles the turns TOD defers.
The trade-off is straightforward: TOD handles a narrower job — picking from a known set of actions — but does it with a smaller model, no output tokens, and calibrated confidence you can act on. It reads images alongside text, which the base model could not do. Customer support is the showcase, but the decision model is designed to transfer across tasks.
| Model | Accuracy | Easy / Orig / Hard | Calibration Error |
| Untrained base (12B) | 86.2% | 100% / 94.4% / 74.8% | 3.6% |
| TOD (trained) | 84.4% | 100% / 91.7% / 73.0% | 3.3% (0.8% raw) |
- Base: open Gemma 4 12B model
- Context window: up to 48,000 tokens per decision
- Output tokens: zero (input-only billing)
- Encoder size: ~400M parameters for action scoring
- Training data: proprietary decision data, not JevBench