TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

The Model Turned Its Own Failure Into the Training Signal

SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

Published Updated Story ID: mp-2026-08-25-009
Read the complete editionStory JSON

Summary

SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

With a Qwen3-8B base, the authors report 73.3 percent on AIME 2024 using 8 percent of the training FLOPs of scaled supervised fine-tuning, alongside gains on WebShop, ALFWorld and SWE-Bench-Lite. The method uses reflection-conditioned teacher scores without a separate critic or larger teacher model.

Why it matters

SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

    Evidence: source-2026-08-25-009

Sources

  1. arXiv preprint 2608.23493arXiv · primary research

Corrections

No corrections have been recorded for this story.