consumer products
The Best Vision Model Read Fewer Than Three in Ten Cultural Pairs
NormViz-Bench changed one culturally relevant behavior at a time across 6,536 images from 16 countries.
Summary
NormViz-Bench changed one culturally relevant behavior at a time across 6,536 images from 16 countries.
NormViz-Bench contains 3,268 human-validated contrastive image pairs spanning 16 countries. Each pair differs only in a behavior that changes whether the scene conforms to, violates or is irrelevant to a local social norm, and pair-level scoring requires both images to be correct. Gemini 3.0 Flash and Qwen2.5-VL-7B reached 26.6 and 21.6 percent pair accuracy, respectively. Fine-tuning smaller Qwen3-VL models on 64,000 explanation-linked images raised relative accuracy, but absolute performance remained below 30 percent. The benchmark measures its curated countries, behaviors and labels; it does not reduce culture to a single universal rulebook.
Why it matters
NormViz-Bench changed one culturally relevant behavior at a time across 6,536 images from 16 countries.
Limits and context
- Each pair differs only in a behavior that changes whether the scene conforms to, violates or is irrelevant to a local social norm, and pair-level scoring requires both images to be correct.
- The benchmark measures its curated countries, behaviors and labels; it does not reduce culture to a single universal rulebook.
Key claims
NormViz-Bench changed one culturally relevant behavior at a time across 6,536 images from 16 countries.
Qualification: Each pair differs only in a behavior that changes whether the scene conforms to, violates or is irrelevant to a local social norm, and pair-level scoring requires both images to be correct.
Evidence: source-2026-09-09-009
Sources
- arXiv preprint 2609.06831arXiv · primary research
Corrections
No corrections have been recorded for this story.