TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

The Small Model Borrowed a Better Skeleton

Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.

Published Updated Story ID: mp-2026-08-13-003
Read the complete editionStory JSON

Summary

Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.

On four theory-of-mind benchmarks, stronger builder models used five percent of each dataset to iteratively design an inference harness for a weaker target model. Average target performance rose from 0.49 to 0.91 without parameter updates. The largest gains came from moving brittle reasoning into deterministic code, routing cases and enforcing answer formats rather than simply eliciting more chain-of-thought. The result is benchmark-specific scaffolding, not a general transfer of every capability.

Why it matters

Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.

Limits and context

  • The result is benchmark-specific scaffolding, not a general transfer of every capability.

Key claims

  1. Builder models nearly doubled weaker models' average benchmark scores by writing deterministic test-time harnesses instead of changing their weights.

    Qualification: The result is benchmark-specific scaffolding, not a general transfer of every capability.

    Evidence: source-2026-08-13-003

Sources

  1. arXiv preprint 2608.12307arXiv · primary research

Corrections

No corrections have been recorded for this story.