safety
The Reasoning Was Readable. Its Important Steps Were Still Hidden
Rollout-based advantage tests found that language-model judges could identify consequential chain-of-thought steps better than chance—but far below the experiment's noise ceiling.

Summary
Rollout-based advantage tests found that language-model judges could identify consequential chain-of-thought steps better than chance—but far below the experiment's noise ceiling.
Researchers estimated each reasoning step's functional importance by measuring how much including it changed the expected final reward across Monte Carlo continuations. Capable language models beat a prevalence baseline when asked to identify high-advantage steps from the text alone, yet remained well short of the noise ceiling. Fine-tuning a step critic helped more on wrong answers than correct ones. The result does not show that reasoning text is useless; it shows that readable prose only partially reveals which step actually carries the answer. That distinction matters when traces are used for error diagnosis, process rewards or claims of interpretability.
Why it matters
Rollout-based advantage tests found that language-model judges could identify consequential chain-of-thought steps better than chance—but far below the experiment's noise ceiling.
Limits and context
- The result does not show that reasoning text is useless; it shows that readable prose only partially reveals which step actually carries the answer.
Key claims
Rollout-based advantage tests found that language-model judges could identify consequential chain-of-thought steps better than chance—but far below the experiment's noise ceiling.
Qualification: The result does not show that reasoning text is useless; it shows that readable prose only partially reveals which step actually carries the answer.
Evidence: source-2026-09-06-001
Sources
- arXiv preprint 2609.04194arXiv · primary research
Corrections
No corrections have been recorded for this story.