TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Research Agents Struggled to Learn What Changed the Score

WhatWorkedBench asks agents to predict every configuration after a small experiment budget.

Published Updated Story ID: mp-2026-09-24-018
Read the complete editionStory JSON

Summary

WhatWorkedBench asks agents to predict every configuration after a small experiment budget.

Across 36 tasks and 1,248 configurations, Gaussian-process fitting improved effect recovery in two cohorts. Encoding behaviorally equivalent code settings nearly doubled recovery in one six-workflow test.

Why it matters

WhatWorkedBench asks agents to predict every configuration after a small experiment budget.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. WhatWorkedBench asks agents to predict every configuration after a small experiment budget.

    Evidence: source-2026-09-24-020

Sources

  1. arXiv preprint 2609.27490arXiv · primary research

Corrections

No corrections have been recorded for this story.