safety security
Malicious Tool Metadata Pulled Agents Off Course
A black-box attack optimized MCP listings and returns to attract calls and steer outcomes.

Summary
A black-box attack optimized MCP listings and returns to attract calls and steer outcomes.
A2M targets agents that choose Model Context Protocol tools through semantic matching. Its attraction stage rewrites attacker-controlled metadata to raise invocation probability; its manipulation stage refines tool returns from execution traces. On LiveMCPBench, attacks optimized for one model reached a 93.6% macro-average malicious invocation rate across four scenarios, while cross-model transfer was lower. These are benchmark results under a constructed threat model, but they support stronger tool vetting and runtime isolation.
Why it matters
A black-box attack optimized MCP listings and returns to attract calls and steer outcomes.
Limits and context
No additional limitation was separately recorded.
Key claims
A black-box attack optimized MCP listings and returns to attract calls and steer outcomes.
Evidence: source-2026-09-23-003
Sources
- arXiv preprint 2609.26761arXiv · primary research
Corrections
No corrections have been recorded for this story.