TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

An Agent Forgot Its Reasoning After the State Was Saved

A training-free compressor cut input, output and cache tokens while raising reported reward on WorkBuddyBench.

Published Updated Story ID: mp-2026-09-27-003
Read the complete editionStory JSON

Summary

A training-free compressor cut input, output and cache tokens while raising reported reward on WorkBuddyBench.

The ICLR method ranks completed reasoning blocks by whether they still matter after an agent has externalized useful state into files, tools or the environment. On 260 WorkBuddyBench tasks, the authors report reward rising from 0.699 to 0.718 while input, output and cache tokens fell by 25.5%, 14.4% and 33.3%, respectively. The study argues that internal reasoning can become replaceable once its consequences are safely recorded, though the result is benchmark-specific and does not license deleting audit evidence.

Why it matters

A training-free compressor cut input, output and cache tokens while raising reported reward on WorkBuddyBench.

Limits and context

  • The study argues that internal reasoning can become replaceable once its consequences are safely recorded, though the result is benchmark-specific and does not license deleting audit evidence.

Key claims

  1. A training-free compressor cut input, output and cache tokens while raising reported reward on WorkBuddyBench.

    Qualification: The study argues that internal reasoning can become replaceable once its consequences are safely recorded, though the result is benchmark-specific and does not license deleting audit evidence.

    Evidence: source-2026-09-27-003

Sources

  1. arXiv preprint 2609.29875arXiv · primary research

Corrections

No corrections have been recorded for this story.