benchmarks evals
The Judge Couldn’t Find the Step That Changed the Outcome
An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.

Summary
An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.
The researchers resampled a policy’s own alternatives at each decision point in ALFWorld and rolled the trajectory forward, creating an executed counterfactual measure of what actually changed the outcome. Against that causal reference, LLM-judge scores, outcome-conditioned log-probability ratios and the policy’s confidence all performed at chance; the authors also report that measurable contribution was sparse and that the available counterfactuals depended on the policy. Their seven-arm training experiment found no arm that reliably beat the untrained policy, while differing sample counts explained apparent training signatures. The result is a preprint finding in one single-agent environment, but its warning is broader: a fluent-looking training signal can measure exposure or correctness without identifying the step that caused success.
Why it matters
An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.
Limits and context
No additional limitation was separately recorded.
Key claims
An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.
Evidence: source-2026-08-23-001
Sources
- arXiv preprint 2608.19760arXiv · primary research
Corrections
No corrections have been recorded for this story.