safety security
The Robot Followed the Rule—Until the Conversation Got Longer
A 40-session benchmark found that shorter context reduced ordinary behavior problems yet nearly doubled one model’s full safety violations.

Summary
A 40-session benchmark found that shorter context reduced ordinary behavior problems yet nearly doubled one model’s full safety violations.
Researchers built a Model Context Protocol test environment around five safety invariants grounded in ISO 10218-2:2025 protective measures, then ran four model backends through 40 sessions of 100 turns. Layer-one text results placed responses on a spectrum from correct compliance through overcompliance and undercompliance to full violation. Claude and Gemini stayed at or near zero violations, while GPT-4o-mini reached as many as 13 in a session. Sliding-window context reduced mean behavioral issues by 42 to 57 percent for every cloud backend, but GPT-4o-mini’s mean violations rose from 3.8 to 7.2. Simulation and physical validation remain incomplete, so the current result is a benchmark warning about orchestration, not a finished robot-safety certification.
Why it matters
A 40-session benchmark found that shorter context reduced ordinary behavior problems yet nearly doubled one model’s full safety violations.
Limits and context
- Simulation and physical validation remain incomplete, so the current result is a benchmark warning about orchestration, not a finished robot-safety certification.
Key claims
A 40-session benchmark found that shorter context reduced ordinary behavior problems yet nearly doubled one model’s full safety violations.
Qualification: Simulation and physical validation remain incomplete, so the current result is a benchmark warning about orchestration, not a finished robot-safety certification.
Evidence: source-2026-09-09-001
Sources
- arXiv preprint 2609.07288arXiv · primary research
Corrections
No corrections have been recorded for this story.