safety security
The Detector Recognized the Source More Than the Threat
MaliciousSkillBench consolidates 9,740 agent-skill packages and tests whether detection transfers beyond familiar feeds.

Summary
MaliciousSkillBench consolidates 9,740 agent-skill packages and tests whether detection transfers beyond familiar feeds.
The benchmark reduces 8,414 raw malicious records to 7,539 normalized identities and evaluates both learned detectors and off-the-shelf scanners. Random-split macro F1 reached as high as 0.932, but source-disjoint performance fell to 0.653–0.665; the strongest text model retained high malicious recall while falsely flagging 62.4 percent of benign skills from held-out sources.
Why it matters
MaliciousSkillBench consolidates 9,740 agent-skill packages and tests whether detection transfers beyond familiar feeds.
Limits and context
No additional limitation was separately recorded.
Key claims
MaliciousSkillBench consolidates 9,740 agent-skill packages and tests whether detection transfers beyond familiar feeds.
Evidence: source-2026-08-23-013
Sources
- arXiv preprint 2608.19901arXiv · primary research
Corrections
No corrections have been recorded for this story.