TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety

Benign Inputs Combined Into Harm

Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.

Published Updated Story ID: mp-2026-08-30-012
Read the complete editionStory JSON

Summary

Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.

The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context. Representative safety guards missed both kinds of compositional evidence across time and modality in the authors' evaluation. The dataset is scheduled for release in October 2026, so current claims rest on the paper's reported protocol rather than an independently inspectable public benchmark.

Why it matters

Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.

Limits and context

  • The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context.

Key claims

  1. Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.

    Qualification: The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context.

    Evidence: source-2026-08-30-012

Sources

  1. arXiv preprint 2608.26535arXiv · primary research

Corrections

No corrections have been recorded for this story.