TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Medical Model Read the Choices Better Than the Image

A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.

Published Updated Story ID: mp-2026-08-15-019
Read the complete editionStory JSON

Summary

A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.

The best tested model reached 79 percent on the full visual-question set, but most systems lagged the available human reference on a response subset. Models performed worse when images carried the answer and still exploited answer-choice cues without the image or question.

Why it matters

A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.

    Evidence: source-2026-08-15-021

Sources

  1. arXiv preprint 2608.12928arXiv · primary research

Corrections

No corrections have been recorded for this story.