benchmarks evals
The Answer Translated Better Than the Tool Policy
Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.
Summary
Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.
The study compares 2.38 million rollouts from eight models on six parallel tool-use benchmarks and measures actions rather than final text. After correcting for trace length, empty outputs, chance overlap, reproducibility and same-language variation, four frontier models converged on 71 to 73 percent cross-language policy retention under greedy decoding. The authors also found that non-English tasks route through an English pivot and that extraction code can manufacture apparent failures.
Why it matters
Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.
Limits and context
No additional limitation was separately recorded.
Key claims
Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.
Evidence: source-2026-08-12-012
Sources
- arXiv preprint 2608.11110arXiv · primary research
Corrections
No corrections have been recorded for this story.