TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Harness Became an Optimization Benchmark

Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.

Published Updated Story ID: mp-2026-08-08-019
Read the complete editionStory JSON

Summary

Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.

HarnessOpt-Bench scores normalized held-out gain while a trusted environment protects the test partition and preserves every candidate. Across four downstream tasks and 111 scored runs, optimizer models separated more than their coding harnesses, native harnesses were not consistently better, and gains varied sharply by task and starting point.

Why it matters

Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.

Limits and context

  • Across four downstream tasks and 111 scored runs, optimizer models separated more than their coding harnesses, native harnesses were not consistently better, and gains varied sharply by task and starting point.

Key claims

  1. Five frontier models edited prompts, tools, memory and orchestration under a metered evaluation budget.

    Qualification: Across four downstream tasks and 111 scored runs, optimizer models separated more than their coding harnesses, native harnesses were not consistently better, and gains varied sharply by task and starting point.

    Evidence: source-2026-08-08-021

Sources

  1. arXiv preprint 2608.06301arXiv · primary research

Corrections

No corrections have been recorded for this story.