safety
Benign Inputs Combined Into Harm
Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.
Summary
Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.
The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context. Representative safety guards missed both kinds of compositional evidence across time and modality in the authors' evaluation. The dataset is scheduled for release in October 2026, so current claims rest on the paper's reported protocol rather than an independently inspectable public benchmark.
Why it matters
Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.
Limits and context
- The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context.
Key claims
Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.
Qualification: The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context.
Evidence: source-2026-08-30-012
Sources
- arXiv preprint 2608.26535arXiv · primary research
Corrections
No corrections have been recorded for this story.