TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Judge Couldn’t Find the Step That Changed the Outcome

An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.

Published Updated Story ID: mp-2026-08-23-001
Read the complete editionStory JSON

Summary

An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.

The researchers resampled a policy’s own alternatives at each decision point in ALFWorld and rolled the trajectory forward, creating an executed counterfactual measure of what actually changed the outcome. Against that causal reference, LLM-judge scores, outcome-conditioned log-probability ratios and the policy’s confidence all performed at chance; the authors also report that measurable contribution was sparse and that the available counterfactuals depended on the policy. Their seven-arm training experiment found no arm that reliably beat the untrained policy, while differing sample counts explained apparent training signatures. The result is a preprint finding in one single-agent environment, but its warning is broader: a fluent-looking training signal can measure exposure or correctness without identifying the step that caused success.

Why it matters

An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. An executed-replay audit found that three common step-level credit signals for tool-using agents identified causal contribution no better than chance.

    Evidence: source-2026-08-23-001

Sources

  1. arXiv preprint 2608.19760arXiv · primary research

Corrections

No corrections have been recorded for this story.