benchmarks evals
The Agent Read the Numbers Before It Drew the Plot
TraceBench generates controlled physical time series so root-cause attribution can be tested against known parameter changes.
Summary
TraceBench generates controlled physical time series so root-cause attribution can be tested against known parameter changes.
Four evaluated agents benefited substantially from domain context and explored data mainly through numerical console output rather than visualizations. They also performed worse when asked to write a reusable sample-to-label Python program than when submitting predictions directly. The released simulations, trajectories and leaderboard make these behavioral differences auditable.
Why it matters
TraceBench generates controlled physical time series so root-cause attribution can be tested against known parameter changes.
Limits and context
No additional limitation was separately recorded.
Key claims
TraceBench generates controlled physical time series so root-cause attribution can be tested against known parameter changes.
Evidence: source-2026-08-29-009
Sources
- arXiv preprint 2608.27182arXiv · primary research
Corrections
No corrections have been recorded for this story.