TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Right Answer Still Missed One-Fifth of the Evidence

Across 235 multimodal science tasks, answer accuracy exceeded complete evidence recovery by more than 20 points.

Published Updated Story ID: mp-2026-09-12-006
Read the complete editionStory JSON

Summary

Across 235 multimodal science tasks, answer accuracy exceeded complete evidence recovery by more than 20 points.

Sci-MMR links scientific claims to citations, figures and supporting image regions across 235 multi-hop tasks in four disciplines. Eight frontier multimodal models consistently answered more often than they recovered all required evidence, with a gap above 20 percentage points. Missing or incomplete figure evidence accounted for 57.2% of failures; gold evidence improved accuracy by as much as 37 points. Even with gold evidence, the strongest model reached 69.1% on the hardest tasks, exposing a second bottleneck in combining what had been found.

Why it matters

Across 235 multimodal science tasks, answer accuracy exceeded complete evidence recovery by more than 20 points.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Across 235 multimodal science tasks, answer accuracy exceeded complete evidence recovery by more than 20 points.

    Evidence: source-2026-09-12-006

Sources

  1. arXiv preprint 2609.11243arXiv · primary research

Corrections

No corrections have been recorded for this story.