frontier models
One Training Query Reached Seventy-One Percent of the Teacher's States
On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.
Summary
On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.
A single query's rollouts visited 71.5 percent of the states reached by full-data training, mostly within the first 100 steps. Sixteen semantically diverse queries reached 98.9 percent coverage and matched full-data gains, while content-light and off-domain prompts approached the real-query baseline. The authors argue that on-policy distillation is data-overfed but algorithm-starved: rollouts expose broad supervision quickly, then alignment absorbs it slowly. The finding is bounded to the tested tasks, teachers and model families.
Why it matters
On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.
Limits and context
No additional limitation was separately recorded.
Key claims
On-policy distillation kept improving for hundreds of steps even when the student repeatedly learned from a single prompt.
Evidence: source-2026-09-06-012
Sources
- arXiv preprint 2609.04172arXiv · primary research
Corrections
No corrections have been recorded for this story.