benchmarks evals
The Translation Benchmark Collects Only Examples That Still Break Models
A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.
Summary
A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.
The Last Translation Benchmark starts from human-authored examples that defeat leading systems rather than a static set approaching saturation. Each text, image, audio or video case includes verification rules for concrete errors, aiming to make evaluation more reproducible and actionable than a single automatic score. Version one includes accepted contributions before September 1 and is designed to keep growing. Its value will depend on contribution quality, coverage and sustained review; it is not itself evidence that translation progress has stopped.
Why it matters
A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.
Limits and context
- Its value will depend on contribution quality, coverage and sustained review; it is not itself evidence that translation progress has stopped.
Key claims
A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.
Qualification: Its value will depend on contribution quality, coverage and sustained review; it is not itself evidence that translation progress has stopped.
Evidence: source-2026-09-06-011
Sources
- arXiv preprint 2609.04173arXiv · primary research
Corrections
No corrections have been recorded for this story.