frontier models
The Small Model Borrowed a Better Skeleton
Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.

Summary
Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.
On four theory-of-mind benchmarks, stronger builder models used five percent of each dataset to iteratively design an inference harness for a weaker target model. Average target performance rose from 0.49 to 0.91 without parameter updates. The largest gains came from moving brittle reasoning into deterministic code, routing cases and enforcing answer formats rather than simply eliciting more chain-of-thought. The result is benchmark-specific scaffolding, not a general transfer of every capability.
Why it matters
Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.
Limits and context
- The result is benchmark-specific scaffolding, not a general transfer of every capability.
Key claims
Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.
Qualification: The result is benchmark-specific scaffolding, not a general transfer of every capability.
Evidence: source-2026-08-13-003
Sources
- arXiv preprint 2608.12307arXiv · primary research
Corrections
No corrections have been recorded for this story.