research
Correct-Looking Claims Survived Broken Evidence
In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.

Summary
In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.
SciRIGOR evaluates scientific coding agents as linked chains from executable analysis through results and figures to claims. Its 100 cases span six domains and 17 subfields, with typed evidence graphs that distinguish artifact fidelity from the validity of each supporting relation. Across 11 agent-model configurations, claims agreed with faithful results 91.8 percent of the time and with unfaithful results 91.0 percent of the time. No system exceeded 62.6 percent on the soft evidence-chain score or 18 percent on strict whole-chain success. Internal coherence therefore did not establish scientific correctness in this benchmark.
Why it matters
In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.
Limits and context
- Internal coherence therefore did not establish scientific correctness in this benchmark.
Key claims
In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.
Qualification: Internal coherence therefore did not establish scientific correctness in this benchmark.
Evidence: source-2026-09-09-008
Sources
- arXiv preprint 2609.06192arXiv · primary research
Corrections
No corrections have been recorded for this story.