other
Grade 2 Braille Exposed the Accessibility Gap
BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances.
Summary
BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances.
Six evaluated models showed a persistent gap between print-English and Braille performance. Understanding and expression were asymmetric, contracted Grade 2 Braille was especially fragile on input, and fully Braille requests reduced performance further. The deterministic pipeline used no model-generated test instances, giving accessibility failures a clearer measurement target.
Why it matters
BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances.
Limits and context
No additional limitation was separately recorded.
Key claims
BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances.
Evidence: source-2026-08-28-010
Sources
- arXiv preprint 2608.27268arXiv · primary research
Corrections
No corrections have been recorded for this story.