frontier models
Raw Chat Logs Beat the Memory Architecture
An agent-controlled lexical search over unmodified conversations outscored graph and tree memory systems on a matched test suite.
Summary
An agent-controlled lexical search over unmodified conversations outscored graph and tree memory systems on a matched test suite.
ReFind leaves conversation archives unmodified and gives an agent controls for session-aware ranking, local context expansion, temporal narrowing and skipping inspected sessions. Across roughly 2,800 conversational-memory questions, it reported 58.2 mean accuracy versus 53.2 for the strongest graph- or tree-based comparison under the same GPT-4o-mini backbone. The result suggests structured preprocessing is not always the source of retrieval gains; it is specific to precise, evidence-grounded refinding tasks and the tested models.
Why it matters
An agent-controlled lexical search over unmodified conversations outscored graph and tree memory systems on a matched test suite.
Limits and context
- The result suggests structured preprocessing is not always the source of retrieval gains; it is specific to precise, evidence-grounded refinding tasks and the tested models.
Key claims
An agent-controlled lexical search over unmodified conversations outscored graph and tree memory systems on a matched test suite.
Qualification: The result suggests structured preprocessing is not always the source of retrieval gains; it is specific to precise, evidence-grounded refinding tasks and the tested models.
Evidence: source-2026-08-16-006
Sources
- arXiv preprint 2608.12888arXiv · primary research
Corrections
No corrections have been recorded for this story.