benchmarks evals
Research Agents Struggled to Learn What Changed the Score
WhatWorkedBench asks agents to predict every configuration after a small experiment budget.
Published Updated Story ID: mp-2026-09-24-018
Summary
WhatWorkedBench asks agents to predict every configuration after a small experiment budget.
Across 36 tasks and 1,248 configurations, Gaussian-process fitting improved effect recovery in two cohorts. Encoding behaviorally equivalent code settings nearly doubled recovery in one six-workflow test.
Why it matters
WhatWorkedBench asks agents to predict every configuration after a small experiment budget.
Limits and context
No additional limitation was separately recorded.
Key claims
WhatWorkedBench asks agents to predict every configuration after a small experiment budget.
Evidence: source-2026-09-24-020
Sources
- arXiv preprint 2609.27490arXiv · primary research
Corrections
No corrections have been recorded for this story.