safety
Two Percent of Features Held the Safety Line
Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.
Summary
Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.
Across three model families and two evolution algorithms, the authors report that constraining an identified safety circuit representing under two percent of features preserved safety better than explicit reward constraints. The result is model- and evaluation-dependent evidence for circuit anchoring, not a general guarantee against unsafe adaptation.
Why it matters
Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.
Limits and context
- The result is model- and evaluation-dependent evidence for circuit anchoring, not a general guarantee against unsafe adaptation.
Key claims
Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.
Qualification: The result is model- and evaluation-dependent evidence for circuit anchoring, not a general guarantee against unsafe adaptation.
Evidence: source-2026-08-09-018
Sources
- arXiv preprint 2608.05158arXiv · primary research
Corrections
No corrections have been recorded for this story.