TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Translation Benchmark Collects Only Examples That Still Break Models

A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.

Published Updated Story ID: mp-2026-09-06-011
Read the complete editionStory JSON

Summary

A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.

The Last Translation Benchmark starts from human-authored examples that defeat leading systems rather than a static set approaching saturation. Each text, image, audio or video case includes verification rules for concrete errors, aiming to make evaluation more reproducible and actionable than a single automatic score. Version one includes accepted contributions before September 1 and is designed to keep growing. Its value will depend on contribution quality, coverage and sustained review; it is not itself evidence that translation progress has stopped.

Why it matters

A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.

Limits and context

  • Its value will depend on contribution quality, coverage and sustained review; it is not itself evidence that translation progress has stopped.

Key claims

  1. A live, peer-reviewed dataset pairs difficult multimodal translations with handcrafted rules that say exactly what failure looks like.

    Qualification: Its value will depend on contribution quality, coverage and sustained review; it is not itself evidence that translation progress has stopped.

    Evidence: source-2026-09-06-011

Sources

  1. arXiv preprint 2609.04173arXiv · primary research

Corrections

No corrections have been recorded for this story.