frontier models
The Agent Studied Before It Knew the Exam
A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.

Summary
A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.
Task-agnostic environment preprocessing lets an agent explore corpora and tools before it knows the downstream task distribution. Across six heterogeneous benchmarks, an archive-equipped meta-agent chose preparation artifacts and achieved the highest Avg@3 reward on five, while fixed corpus processing remained best on the largest corpus. Larger study budgets did not reliably improve reward, but the resulting artifacts reduced how many test-time samples were needed to reach a given score. The study shifts some computation from repeated attempts into reusable preparation without claiming one study strategy works everywhere.
Why it matters
A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.
Limits and context
- Larger study budgets did not reliably improve reward, but the resulting artifacts reduced how many test-time samples were needed to reach a given score.
Key claims
A meta-agent prepared reusable artifacts without task examples, then reduced the frozen solver’s test-time sampling on six benchmarks.
Qualification: Larger study budgets did not reliably improve reward, but the resulting artifacts reduced how many test-time samples were needed to reach a given score.
Evidence: source-2026-09-13-003
Sources
- arXiv preprint 2609.10824arXiv · primary research
Corrections
No corrections have been recorded for this story.