safety security
The Tool Server Waited Until the Agent Trusted It
TrustShift models an MCP server that behaves honestly through a conditioning window before switching to a malicious payload.
Summary
TrustShift models an MCP server that behaves honestly through a conditioning window before switching to a malicious payload.
Across four production-style domains and frontier proprietary and open-weight models, the staged attacks reached a reported 69.5 percent mean success rate. A transport-boundary defense learned behavioral baselines during clean windows and reduced that figure to 42.7 percent, leaving meaningful residual risk and showing why static deployment checks cannot see a later server-controlled defection.
Why it matters
TrustShift models an MCP server that behaves honestly through a conditioning window before switching to a malicious payload.
Limits and context
- A transport-boundary defense learned behavioral baselines during clean windows and reduced that figure to 42.7 percent, leaving meaningful residual risk and showing why static deployment checks cannot see a later server-controlled defection.
Key claims
TrustShift models an MCP server that behaves honestly through a conditioning window before switching to a malicious payload.
Qualification: A transport-boundary defense learned behavioral baselines during clean windows and reduced that figure to 42.7 percent, leaving meaningful residual risk and showing why static deployment checks cannot see a later server-controlled defection.
Evidence: source-2026-08-26-006
Sources
- arXiv preprint 2608.23763arXiv · primary research
Corrections
No corrections have been recorded for this story.