When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. This makes total task consumption hard to predict before execution and requires revision as the run unfolds.
Composable cost representation for each execution segment
TokenCast learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Each segment's representation captures both the tokens it generates and how it expands the context that later segments must process. This granular recording allows the system to track cost contributions at fine granularity rather than treating the entire execution as a black box.
Cumulative estimation through adjacent segment composition
Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. Because each call's input includes the accumulated output of all preceding calls, the total consumption grows nonlinearly. TokenCast's composition operator accounts for this by summing segment costs while weighting each by the context multiplier introduced by prior segments.
Online forecast refinement during execution
As execution unfolds, newly observed evidence refreshes the forecast. The prediction updates without requiring additional LLM calls, incorporating real consumption data as it becomes available. On SWE-bench Verified, this cumulative prediction incurs a mean overhead of 32.8 ms per run, making the approach practical for online use.
Offline budget-control replay results
In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. Across 4 task suites and 6 agent models, the mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. These results demonstrate that the composable representation both improves prediction accuracy and enables tighter token budgeting.
Practical use for LLM developer workflows
A working developer can apply TokenCast to any LLM agent system without changing the underlying model or requiring additional inference calls. The composable cost representation integrates with existing execution pipelines and provides real-time forecast updates that improve budget management. This makes it easier to reason about token costs during agent design and to set practical token limits for production deployments.