benchmarks evals
The Search Kept Going After the First Plausible Failure
Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.

Summary
Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.
Long agent runs can bury the decisive evidence hundreds of actions away from the visible failure. A one-shot judge may settle early on a diagnosis that sounds plausible while leaving most of the trace unexplored. Continual Search instead asks the judge to keep looking across successive turns for unresolved evidence. Across four existing root-cause benchmarks and the new 50-trial MegaRCA-Mix set, the authors report consistent gains; on MegaRCA-Mix, GPT-5.5 F1 rose more than 40%, from 0.349 to 0.498. The benchmark is research evidence, not proof that automated diagnosis can replace production incident review.
Why it matters
Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.
Limits and context
- The benchmark is research evidence, not proof that automated diagnosis can replace production incident review.
Key claims
Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.
Qualification: The benchmark is research evidence, not proof that automated diagnosis can replace production incident review.
Evidence: source-2026-09-15-001
Sources
- arXiv preprint 2609.13463arXiv · primary research
Corrections
No corrections have been recorded for this story.