TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Literal Memory Won More Sessions but Proved No General Mechanism

DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.

Published Updated Story ID: mp-2026-08-24-011
Read the complete editionStory JSON

Summary

DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.

In the preregistered successor audit, no external memory passed 21 of 180 tasks, deterministic verbatim memory passed 82 and one pinned hosted configuration passed 97. The authors explicitly stop short of claiming mechanism, general product superiority or equivalence, making the benchmark a profile of exact conditions rather than a universal memory ranking.

Why it matters

DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.

    Evidence: source-2026-08-24-011

Sources

  1. arXiv preprint 2608.20664arXiv · primary research

Corrections

No corrections have been recorded for this story.