benchmarks evals
Gold Evidence Added Fourteen to Twenty-Two Points
A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.
Summary
A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.
The best model on SciFact reached macro-F1 0.70 and fell to 0.31 on ClimateCheck. Replacing retrieved evidence with gold annotations improved veracity accuracy by 14 to 22 points across models, and noisy evidence sometimes made claim-only systems outperform more elaborate pipelines.
Why it matters
A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.
Limits and context
- Replacing retrieved evidence with gold annotations improved veracity accuracy by 14 to 22 points across models, and noisy evidence sometimes made claim-only systems outperform more elaborate pipelines.
Key claims
A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.
Qualification: Replacing retrieved evidence with gold annotations improved veracity accuracy by 14 to 22 points across models, and noisy evidence sometimes made claim-only systems outperform more elaborate pipelines.
Evidence: source-2026-08-27-006
Sources
- arXiv preprint 2608.25934arXiv · primary research
Corrections
No corrections have been recorded for this story.