benchmarks evals
The Final Paper Hid Thirty Times More Mistakes
A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.

Summary
A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.
OpenDiscoveryTrace records nine fields for every step—including tool calls, observations, errors, revision triggers and self-reported confidence—across 558 trajectories on 124 scientific tasks. In a pilot analysis of 363 LLM-judged runs, three frontier models finished at similar reported success rates of 84 to 89 percent, yet one produced 2.5 logged errors per trajectory versus 0.08 for another, a thirtyfold difference. Their failure shapes also diverged: tool misuse dominated one model while reasoning errors dominated another. The dataset is designed to make scientific-agent process auditable, though the headline comparisons rely on the paper’s trace schema and judge-based pilot rather than independent replication.
Why it matters
A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.
Limits and context
No additional limitation was separately recorded.
Key claims
A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.
Evidence: source-2026-09-10-001
Sources
- arXiv preprint 2609.09203arXiv · primary research
Corrections
No corrections have been recorded for this story.