TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

Correct-Looking Claims Survived Broken Evidence

In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.

Published Updated Story ID: mp-2026-09-09-008
Read the complete editionStory JSON

Summary

In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.

SciRIGOR evaluates scientific coding agents as linked chains from executable analysis through results and figures to claims. Its 100 cases span six domains and 17 subfields, with typed evidence graphs that distinguish artifact fidelity from the validity of each supporting relation. Across 11 agent-model configurations, claims agreed with faithful results 91.8 percent of the time and with unfaithful results 91.0 percent of the time. No system exceeded 62.6 percent on the soft evidence-chain score or 18 percent on strict whole-chain success. Internal coherence therefore did not establish scientific correctness in this benchmark.

Why it matters

In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.

Limits and context

  • Internal coherence therefore did not establish scientific correctness in this benchmark.

Key claims

  1. In SciRIGOR, claims agreed with faithful and unfaithful results at almost the same rate, while strict end-to-end evidence success stayed below 18 percent.

    Qualification: Internal coherence therefore did not establish scientific correctness in this benchmark.

    Evidence: source-2026-09-09-008

Sources

  1. arXiv preprint 2609.06192arXiv · primary research

Corrections

No corrections have been recorded for this story.