TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Gold Evidence Added Fourteen to Twenty-Two Points

A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.

Published Updated Story ID: mp-2026-08-27-006
Read the complete editionStory JSON

Summary

A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.

The best model on SciFact reached macro-F1 0.70 and fell to 0.31 on ClimateCheck. Replacing retrieved evidence with gold annotations improved veracity accuracy by 14 to 22 points across models, and noisy evidence sometimes made claim-only systems outperform more elaborate pipelines.

Why it matters

A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.

Limits and context

  • Replacing retrieved evidence with gold annotations improved veracity accuracy by 14 to 22 points across models, and noisy evidence sometimes made claim-only systems outperform more elaborate pipelines.

Key claims

  1. A four-dataset evaluation found that automated fact-checking rankings changed with domain and metric, while retrieval remained the bottleneck.

    Qualification: Replacing retrieved evidence with gold annotations improved veracity accuracy by 14 to 22 points across models, and noisy evidence sometimes made claim-only systems outperform more elaborate pipelines.

    Evidence: source-2026-08-27-006

Sources

  1. arXiv preprint 2608.25934arXiv · primary research

Corrections

No corrections have been recorded for this story.