TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Citation Graph Remembered Too Much

GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.

Published Updated Story ID: mp-2026-08-09-013
Read the complete editionStory JSON

Summary

GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.

Across 4,440 main experiments, 600 cross-corpus runs and 1,200 paired judgments, GraphRAG returned 11 to 15 source IDs per answer with low citation precision in each tested configuration. Faithfulness fell across hops in a typed aviation-requirements corpus but rose on Wikipedia chains because extra passages were still topically useful; even the same judge changed verdicts often when retrieval state changed. The authors argue that architecture claims need tests across embedders, corpora and judges.

Why it matters

GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. GraphRAG over-cited across corpora, while the damage to answer faithfulness depended on what those extra passages contained.

    Evidence: source-2026-08-09-013

Sources

  1. arXiv preprint 2608.05153arXiv · primary research

Corrections

No corrections have been recorded for this story.