safety security
Harmless-Looking Skills Combined Into One Attack
SkillCascade split malicious intent across components so per-skill scanners missed the whole behavior.

Summary
SkillCascade split malicious intent across components so per-skill scanners missed the whole behavior.
The researchers define a skill-cascading attack as a harmful objective distributed across multiple agent skills, each modification appearing benign on its own. Their SkillCascade-Bench contains 213 validated test cases across several agent systems and domains. In the reported experiments, cascades induced harmful behavior while evading component-level scanners and runtime monitors. The preprint's examples and rates belong to its selected agents, backbones and threat models, but the structural warning is broader: reviewing packages independently can miss risk that exists only in composition.
Why it matters
SkillCascade split malicious intent across components so per-skill scanners missed the whole behavior.
Limits and context
- The preprint's examples and rates belong to its selected agents, backbones and threat models, but the structural warning is broader: reviewing packages independently can miss risk that exists only in composition.
Key claims
SkillCascade split malicious intent across components so per-skill scanners missed the whole behavior.
Qualification: The preprint's examples and rates belong to its selected agents, backbones and threat models, but the structural warning is broader: reviewing packages independently can miss risk that exists only in composition.
Evidence: source-2026-09-28-003
Sources
- arXiv preprint 2609.30383arXiv · primary research
Corrections
No corrections have been recorded for this story.