benchmarks evals
The Harness Became an Optimization Benchmark
Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.
Summary
Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.
HarnessOpt-Bench scores normalized held-out gain while a trusted environment protects the test partition and preserves every candidate. Across four downstream tasks and 111 scored runs, optimizer models separated more than their coding harnesses, native harnesses were not consistently better, and gains varied sharply by task and starting point.
Why it matters
Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.
Limits and context
- Across four downstream tasks and 111 scored runs, optimizer models separated more than their coding harnesses, native harnesses were not consistently better, and gains varied sharply by task and starting point.
Key claims
Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.
Qualification: Across four downstream tasks and 111 scored runs, optimizer models separated more than their coding harnesses, native harnesses were not consistently better, and gains varied sharply by task and starting point.
Evidence: source-2026-08-08-021
Sources
- arXiv preprint 2608.06301arXiv · primary research
Corrections
No corrections have been recorded for this story.