safety security
The Model's Explanation May Not Describe Its Computation
A security argument says linguistic monitoring cannot supply complete guarantees when internal activation-space computation is only lossy-translated into words.
Summary
A security argument says linguistic monitoring cannot supply complete guarantees when internal activation-space computation is only lossy-translated into words.
James Mickens names the gap “linguistic illegibility”: external text or language-like probes may fail to represent how a model produced an outcome. The paper argues that chain-of-thought monitoring, self-critique and linguistically defined activation probes therefore need a non-linguistic security floor. It proposes output taint tracking, robust virtualization and independent configuration audits as complementary isolation measures; this is a position and design argument, not a measured proof that every monitor fails.
Why it matters
A security argument says linguistic monitoring cannot supply complete guarantees when internal activation-space computation is only lossy-translated into words.
Limits and context
- It proposes output taint tracking, robust virtualization and independent configuration audits as complementary isolation measures; this is a position and design argument, not a measured proof that every monitor fails.
Key claims
A security argument says linguistic monitoring cannot supply complete guarantees when internal activation-space computation is only lossy-translated into words.
Qualification: It proposes output taint tracking, robust virtualization and independent configuration audits as complementary isolation measures; this is a position and design argument, not a measured proof that every monitor fails.
Evidence: source-2026-09-03-004
Sources
- arXiv preprint 2609.02852arXiv · primary research
Corrections
No corrections have been recorded for this story.