benchmarks evals
The Citation Graph Remembered Too Much
GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.

Summary
GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.
Across 4,440 main experiments, 600 cross-corpus runs and 1,200 paired judgments, GraphRAG returned 11 to 15 source IDs per answer with low citation precision in each tested configuration. Faithfulness fell across hops in a typed aviation-requirements corpus but rose on Wikipedia chains because extra passages were still topically useful; even the same judge changed verdicts often when retrieval state changed. The authors argue that architecture claims need tests across embedders, corpora and judges.
Why it matters
GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.
Limits and context
No additional limitation was separately recorded.
Key claims
GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.
Evidence: source-2026-08-09-013
Sources
- arXiv preprint 2608.05153arXiv · primary research
Corrections
No corrections have been recorded for this story.