benchmarks evals
The Medical Model Read the Choices Better Than the Image
A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.
Summary
A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.
The best tested model reached 79 percent on the full visual-question set, but most systems lagged the available human reference on a response subset. Models performed worse when images carried the answer and still exploited answer-choice cues without the image or question.
Why it matters
A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.
Limits and context
No additional limitation was separately recorded.
Key claims
A Polish board-exam benchmark found visual-language models underused image evidence and stayed above chance without key inputs.
Evidence: source-2026-08-15-021
Sources
- arXiv preprint 2608.12928arXiv · primary research
Corrections
No corrections have been recorded for this story.