infrastructure
An Agent's Token Budget Was Forecast While It Ran
TokenCast composed segment costs and context growth without making another model call.

Summary
TokenCast composed segment costs and context growth without making another model call.
TokenCast records each execution segment's own token use and the context growth it adds, then composes segments to estimate repeated input costs later in a run. On SWE-bench Verified, the authors report a mean cumulative prediction time of 32.8 milliseconds per run. Across four task suites and six agent models, mean absolute error improved by an average 14.5% over the strongest comparator; an offline replay used 21.3% fewer tokens than a fixed-budget policy at matched trace completion. Replay results do not prove identical savings in live production systems.
Why it matters
TokenCast composed segment costs and context growth without making another model call.
Limits and context
- Replay results do not prove identical savings in live production systems.
Key claims
TokenCast composed segment costs and context growth without making another model call.
Qualification: Replay results do not prove identical savings in live production systems.
Evidence: source-2026-09-29-008
Sources
- arXiv preprint 2609.35760arXiv · primary research
Corrections
No corrections have been recorded for this story.