TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Search Kept Going After the First Plausible Failure

Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.

Published Updated Story ID: mp-2026-09-15-001
Read the complete editionStory JSON

Summary

Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.

Long agent runs can bury the decisive evidence hundreds of actions away from the visible failure. A one-shot judge may settle early on a diagnosis that sounds plausible while leaving most of the trace unexplored. Continual Search instead asks the judge to keep looking across successive turns for unresolved evidence. Across four existing root-cause benchmarks and the new 50-trial MegaRCA-Mix set, the authors report consistent gains; on MegaRCA-Mix, GPT-5.5 F1 rose more than 40%, from 0.349 to 0.498. The benchmark is research evidence, not proof that automated diagnosis can replace production incident review.

Why it matters

Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.

Limits and context

  • The benchmark is research evidence, not proof that automated diagnosis can replace production incident review.

Key claims

  1. Continual Search lifted GPT-5.5 root-cause F1 from 0.349 to 0.498 on 50 human-annotated, long-horizon agent failures.

    Qualification: The benchmark is research evidence, not proof that automated diagnosis can replace production incident review.

    Evidence: source-2026-09-15-001

Sources

  1. arXiv preprint 2609.13463arXiv · primary research

Corrections

No corrections have been recorded for this story.