benchmarks evals
Literal Memory Won More Sessions but Proved No General Mechanism
DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.
Summary
DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.
In the preregistered successor audit, no external memory passed 21 of 180 tasks, deterministic verbatim memory passed 82 and one pinned hosted configuration passed 97. The authors explicitly stop short of claiming mechanism, general product superiority or equivalence, making the benchmark a profile of exact conditions rather than a universal memory ranking.
Why it matters
DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.
Limits and context
No additional limitation was separately recorded.
Key claims
DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.
Evidence: source-2026-08-24-011
Sources
- arXiv preprint 2608.20664arXiv · primary research
Corrections
No corrections have been recorded for this story.