TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety security

The Refusal Circuit Stayed Off Until the Attack Arrived

A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.

Published Updated Story ID: mp-2026-08-17-010
Read the complete editionStory JSON

Summary

A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.

Tripwire identifies neurons associated with harmful inputs while filtering for utility specificity, then clamps them to harmful-conditional activations through a detector-gated intervention or an equivalent offline bias edit. Across four aligned models and four attacks, the paper reports average attack success no higher than 2.0 percent with MT-Bench utility drops of 0.5 to 5.3 percent. Those benchmark results do not establish immunity to unseen attacks, but they show a narrower intervention footprint.

Why it matters

A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.

Limits and context

  • Those benchmark results do not establish immunity to unseen attacks, but they show a narrower intervention footprint.

Key claims

  1. A training-free clamp selected safety-specific neurons under false-discovery control, then triggered the model’s learned refusal only when needed.

    Qualification: Those benchmark results do not establish immunity to unseen attacks, but they show a narrower intervention footprint.

    Evidence: source-2026-08-17-010

Sources

  1. arXiv preprint 2608.14392arXiv · primary research

Corrections

No corrections have been recorded for this story.