safety security
The Refusal Circuit Stayed Off Until the Attack Arrived
A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.
Summary
A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.
Tripwire identifies neurons associated with harmful inputs while filtering for utility specificity, then clamps them to harmful-conditional activations through a detector-gated intervention or an equivalent offline bias edit. Across four aligned models and four attacks, the paper reports average attack success no higher than 2.0 percent with MT-Bench utility drops of 0.5 to 5.3 percent. Those benchmark results do not establish immunity to unseen attacks, but they show a narrower intervention footprint.
Why it matters
A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.
Limits and context
- Those benchmark results do not establish immunity to unseen attacks, but they show a narrower intervention footprint.
Key claims
A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.
Qualification: Those benchmark results do not establish immunity to unseen attacks, but they show a narrower intervention footprint.
Evidence: source-2026-08-17-010
Sources
- arXiv preprint 2608.14392arXiv · primary research
Corrections
No corrections have been recorded for this story.