TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

One Language Model Ran at Every One of Its Twenty Depths

Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.

Published Updated Story ID: mp-2026-09-29-004
Read the complete editionStory JSON

Summary

Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.

Telescopic Language Models train one randomly truncated layer prefix alongside one full-capacity pass on every step. On a 200-million-parameter proxy trained over 20 billion FineWeb-Edu tokens, the resulting model remained usable at all 20 layer prefixes. The authors report a 43% to 44% reduction in area under the quality-budget curve versus fixed-exit suites, matching full-capacity quality at about 12% lower GPU cost per run. These are proxy-scale results, and deployment savings will depend on serving hardware and workload.

Why it matters

Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Stochastic prefix supervision trained a continuum of capacity instead of a few fixed exits.

    Evidence: source-2026-09-29-004

Sources

  1. arXiv preprint 2609.35769arXiv · primary research

Corrections

No corrections have been recorded for this story.