TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Audio Test Added Six Languages and a View of the Scene

EXAM² combines speech, sound, music, mixed audio and images in one multilingual benchmark.

Published Updated Story ID: mp-2026-08-26-005
Read the complete editionStory JSON

Summary

EXAM² combines speech, sound, music, mixed audio and images in one multilingual benchmark.

The benchmark contains 5,667 multiple-choice questions, 22,614 image instances and 135,684 translations across six languages. Tested audio and multimodal models showed substantial multilingual and cross-modal gaps; a lightweight fusion model fine-tuned on the training split improved up to 12.4 percent in multilingual tests and 21.7 percent in multimodal evaluation over its stated baseline.

Why it matters

EXAM² combines speech, sound, music, mixed audio and images in one multilingual benchmark.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. EXAM² combines speech, sound, music, mixed audio and images in one multilingual benchmark.

    Evidence: source-2026-08-26-005

Sources

  1. arXiv preprint 2608.23758arXiv · primary research

Corrections

No corrections have been recorded for this story.