TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety

Two Percent of Features Held the Safety Line

Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.

Published Updated Story ID: mp-2026-08-09-016
Read the complete editionStory JSON

Summary

Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.

Across three model families and two evolution algorithms, the authors report that constraining an identified safety circuit representing under two percent of features preserved safety better than explicit reward constraints. The result is model- and evaluation-dependent evidence for circuit anchoring, not a general guarantee against unsafe adaptation.

Why it matters

Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.

Limits and context

  • The result is model- and evaluation-dependent evidence for circuit anchoring, not a general guarantee against unsafe adaptation.

Key claims

  1. Anchoring a small identified circuit preserved safety during model self-evolution with limited capability cost.

    Qualification: The result is model- and evaluation-dependent evidence for circuit anchoring, not a general guarantee against unsafe adaptation.

    Evidence: source-2026-08-09-018

Sources

  1. arXiv preprint 2608.05158arXiv · primary research

Corrections

No corrections have been recorded for this story.

Two Percent of Features Held the Safety Line · The Machine Press