research
The Assistant Heard the Concern but Didn’t Act on It
Audio barely changed decisions until prosody was converted into an explicit intermediate state.
Summary
Audio barely changed decisions until prosody was converted into an explicit intermediate state.
Hear2Act pairs 480 task scenarios with hidden concerns conveyed either in words or primarily through prosody. For two audio-capable models, adding audio to a transcript moved average optimal-solution rate only from 14.6 to 15.3 percent. When the model first inferred the concern into text and then selected an action, the rate rose to 39.6 percent, close to 40.7 percent with the ground-truth state. The benchmark tests two models and structured scenarios, not all spoken assistants.
Why it matters
Audio barely changed decisions until prosody was converted into an explicit intermediate state.
Limits and context
- For two audio-capable models, adding audio to a transcript moved average optimal-solution rate only from 14.6 to 15.3 percent.
- The benchmark tests two models and structured scenarios, not all spoken assistants.
Key claims
Audio barely changed decisions until prosody was converted into an explicit intermediate state.
Qualification: For two audio-capable models, adding audio to a transcript moved average optimal-solution rate only from 14.6 to 15.3 percent.
Evidence: source-2026-08-22-004
Sources
- arXiv preprint 2608.19515arXiv · primary research
Corrections
No corrections have been recorded for this story.