benchmarks evals
Spoken Claims Broke Text-Ready Fact Checkers
VeriSpeak found a text-to-speech gap across 3,879 balanced claims; retrieval plus explicit reasoning reached 86.1% accuracy.
Summary
VeriSpeak found a text-to-speech gap across 3,879 balanced claims; retrieval plus explicit reasoning reached 86.1% accuracy.
The benchmark presents the same kinds of temporal, geographic and relational facts as speech rather than text. Large audio-language models that handled written claims often failed when those claims were spoken, and retrieval alone brought limited improvement because systems confused retrieved evidence with the claim being checked. A thinking-tuned model performed best when retrieval was paired with explicit claim-evidence comparison. The result measures controlled benchmark claims, not end-to-end misinformation detection in noisy live audio.
Why it matters
VeriSpeak found a text-to-speech gap across 3,879 balanced claims; retrieval plus explicit reasoning reached 86.1% accuracy.
Limits and context
- The result measures controlled benchmark claims, not end-to-end misinformation detection in noisy live audio.
Key claims
VeriSpeak found a text-to-speech gap across 3,879 balanced claims; retrieval plus explicit reasoning reached 86.1% accuracy.
Qualification: The result measures controlled benchmark claims, not end-to-end misinformation detection in noisy live audio.
Evidence: source-2026-09-26-004
Sources
- arXiv preprint 2609.30227arXiv · primary research
Corrections
No corrections have been recorded for this story.