TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Final Paper Hid Thirty Times More Mistakes

A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.

Published Updated Story ID: mp-2026-09-10-001
Read the complete editionStory JSON

Summary

A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.

OpenDiscoveryTrace records nine fields for every step—including tool calls, observations, errors, revision triggers and self-reported confidence—across 558 trajectories on 124 scientific tasks. In a pilot analysis of 363 LLM-judged runs, three frontier models finished at similar reported success rates of 84 to 89 percent, yet one produced 2.5 logged errors per trajectory versus 0.08 for another, a thirtyfold difference. Their failure shapes also diverged: tool misuse dominated one model while reasoning errors dominated another. The dataset is designed to make scientific-agent process auditable, though the headline comparisons rely on the paper’s trace schema and judge-based pilot rather than independent replication.

Why it matters

A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.

    Evidence: source-2026-09-10-001

Sources

  1. arXiv preprint 2609.09203arXiv · primary research

Corrections

No corrections have been recorded for this story.