TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Answer Translated Better Than the Tool Policy

Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.

Published Updated Story ID: mp-2026-08-12-012
Read the complete editionStory JSON

Summary

Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.

The study compares 2.38 million rollouts from eight models on six parallel tool-use benchmarks and measures actions rather than final text. After correcting for trace length, empty outputs, chance overlap, reproducibility and same-language variation, four frontier models converged on 71 to 73 percent cross-language policy retention under greedy decoding. The authors also found that non-English tasks route through an English pivot and that extraction code can manufacture apparent failures.

Why it matters

Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Across 41 languages, four frontier models retained only about 71 to 73 percent of their action policy after correcting five measurement confounds.

    Evidence: source-2026-08-12-012

Sources

  1. arXiv preprint 2608.11110arXiv · primary research

Corrections

No corrections have been recorded for this story.