frontier models
The Teacher Learned What Privilege to Remember
Latent on-policy self-distillation made the teacher's private context learnable and used less than thirty percent of two comparison methods' rollout budgets.

Summary
Latent on-policy self-distillation made the teacher's private context learnable and used less than thirty percent of two comparison methods' rollout budgets.
LOPD retrieves relevant experiences, compresses them into continuous latent tokens for a privileged self-teacher and supplies dense token-level supervision along the student's own trajectories. On agentic tool-use and code-generation tasks, the authors report gains over reinforcement learning with verifiable rewards and several self-distillation baselines. Ablations attribute the improvement to learning the privileged context instead of prescribing answers, feedback or skills. The evidence is benchmark performance, not a demonstration of autonomous open-ended self-improvement.
Why it matters
Latent on-policy self-distillation made the teacher's private context learnable and used less than thirty percent of two comparison methods' rollout budgets.
Limits and context
- The evidence is benchmark performance, not a demonstration of autonomous open-ended self-improvement.
Key claims
Latent on-policy self-distillation made the teacher's private context learnable and used less than thirty percent of two comparison methods' rollout budgets.
Qualification: The evidence is benchmark performance, not a demonstration of autonomous open-ended self-improvement.
Evidence: source-2026-08-14-003
Sources
- arXiv preprint 2608.13040arXiv · primary research
Corrections
No corrections have been recorded for this story.