research
The Model Turned Its Own Failure Into the Training Signal
SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.
Summary
SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.
With a Qwen3-8B base, the authors report 73.3 percent on AIME 2024 using 8 percent of the training FLOPs of scaled supervised fine-tuning, alongside gains on WebShop, ALFWorld and SWE-Bench-Lite. The method uses reflection-conditioned teacher scores without a separate critic or larger teacher model.
Why it matters
SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.
Limits and context
No additional limitation was separately recorded.
Key claims
SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.
Evidence: source-2026-08-25-009
Sources
- arXiv preprint 2608.23493arXiv · primary research
Corrections
No corrections have been recorded for this story.