TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

media creative tools

The Sound Model Learned Where the Noise Was Going

Panoramic video and spatial audio support questions about a source's identity, direction, distance and motion.

Published Updated Story ID: mp-2026-08-11-013
Read the complete editionStory JSON

Summary

Panoramic video and spatial audio support questions about a source's identity, direction, distance and motion.

ST-OmniQA pairs 40,000 panoramic videos with synchronized first-order Ambisonics audio and supplies 400,000 questions across recognition, direction, distance and trajectory tasks. The accompanying ST-Omni-R1 system integrates spatial-audio representations with panoramic vision. The authors report 77.83 percent average semantic accuracy versus 37.28 percent for their strongest evaluated baseline. Those numbers describe this new benchmark, not general real-world listening or surveillance capability.

Why it matters

Panoramic video and spatial audio support questions about a source's identity, direction, distance and motion.

Limits and context

  • Those numbers describe this new benchmark, not general real-world listening or surveillance capability.

Key claims

  1. Panoramic video and spatial audio support questions about a source's identity, direction, distance and motion.

    Qualification: Those numbers describe this new benchmark, not general real-world listening or surveillance capability.

    Evidence: source-2026-08-11-013

Sources

  1. arXiv preprint 2608.09435arXiv · primary research

Corrections

No corrections have been recorded for this story.