TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

The Agent Studied Before It Knew the Exam

A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.

Published Updated Story ID: mp-2026-09-13-003
Read the complete editionStory JSON

Summary

A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.

Task-agnostic environment preprocessing lets an agent explore corpora and tools before it knows the downstream task distribution. Across six heterogeneous benchmarks, an archive-equipped meta-agent chose preparation artifacts and achieved the highest Avg@3 reward on five, while fixed corpus processing remained best on the largest corpus. Larger study budgets did not reliably improve reward, but the resulting artifacts reduced how many test-time samples were needed to reach a given score. The study shifts some computation from repeated attempts into reusable preparation without claiming one study strategy works everywhere.

Why it matters

A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.

Limits and context

  • Larger study budgets did not reliably improve reward, but the resulting artifacts reduced how many test-time samples were needed to reach a given score.

Key claims

  1. A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.

    Qualification: Larger study budgets did not reliably improve reward, but the resulting artifacts reduced how many test-time samples were needed to reach a given score.

    Evidence: source-2026-09-13-003

Sources

  1. arXiv preprint 2609.10824arXiv · primary research

Corrections

No corrections have been recorded for this story.