frontier models
One Language Model Ran at Every One of Its Twenty Depths
Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.
Summary
Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.
Telescopic Language Models train one randomly truncated layer prefix alongside one full-capacity pass on every step. On a 200-million-parameter proxy trained over 20 billion FineWeb-Edu tokens, the resulting model remained usable at all 20 layer prefixes. The authors report a 43% to 44% reduction in area under the quality-budget curve versus fixed-exit suites, matching full-capacity quality at about 12% lower GPU cost per run. These are proxy-scale results, and deployment savings will depend on serving hardware and workload.
Why it matters
Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.
Limits and context
No additional limitation was separately recorded.
Key claims
Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.
Evidence: source-2026-09-29-004
Sources
- arXiv preprint 2609.35769arXiv · primary research
Corrections
No corrections have been recorded for this story.