benchmarks evals
Vision Tests Learned to Fight World Priors
SABRE generates controlled images and questions that punish answers based on expectation instead of pixels.
Summary
SABRE generates controlled images and questions that punish answers based on expectation instead of pixels.
Its first benchmark contains 600 images and 1,000 questions; six tested vision-language models averaged 22.6 percent accuracy. Human review remains part of the pipeline, so the work automates benchmark construction rather than removing validation.
Why it matters
SABRE generates controlled images and questions that punish answers based on expectation instead of pixels.
Limits and context
No additional limitation was separately recorded.
Key claims
SABRE generates controlled images and questions that punish answers based on expectation instead of pixels.
Evidence: source-2026-08-10-018
Sources
- arXiv preprint 2608.07435arXiv · primary research
Corrections
No corrections have been recorded for this story.