TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

One Training Query Reached Seventy-One Percent of the Teacher's States

On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.

Published Updated Story ID: mp-2026-09-06-012
Read the complete editionStory JSON

Summary

On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.

A single query's rollouts visited 71.5 percent of the states reached by full-data training, mostly within the first 100 steps. Sixteen semantically diverse queries reached 98.9 percent coverage and matched full-data gains, while content-light and off-domain prompts approached the real-query baseline. The authors argue that on-policy distillation is data-overfed but algorithm-starved: rollouts expose broad supervision quickly, then alignment absorbs it slowly. The finding is bounded to the tested tasks, teachers and model families.

Why it matters

On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.

    Evidence: source-2026-09-06-012

Sources

  1. arXiv preprint 2609.04172arXiv · primary research

Corrections

No corrections have been recorded for this story.