frontier models
The Truth Probe Read Features the Model Barely Used
Probe alignment and behavioral sensitivity overlapped by about 12% in one Gemma 2 deception setting.
Summary
Probe alignment and behavioral sensitivity overlapped by about 12% in one Gemma 2 deception setting.
A linear probe can decode truthfulness without identifying the features that actually drive an answer. Decomposing one deployed truth probe into sparse-autoencoder features produced only about 12% overlap between geometric alignment and behavioral gradient sensitivity. Ablating features shared by the probe and the model flipped outputs as much as 27%, against 6% for equally sized probe-only sets and 1% for random sets at full coherence. The result held across five seeds and a held-out split, but it is evidence from one model and experimental setting, not a universal account of probe causality.
Why it matters
Probe alignment and behavioral sensitivity overlapped by about 12% in one Gemma 2 deception setting.
Limits and context
- Decomposing one deployed truth probe into sparse-autoencoder features produced only about 12% overlap between geometric alignment and behavioral gradient sensitivity.
- Ablating features shared by the probe and the model flipped outputs as much as 27%, against 6% for equally sized probe-only sets and 1% for random sets at full coherence.
- The result held across five seeds and a held-out split, but it is evidence from one model and experimental setting, not a universal account of probe causality.
Key claims
Probe alignment and behavioral sensitivity overlapped by about 12% in one Gemma 2 deception setting.
Qualification: Decomposing one deployed truth probe into sparse-autoencoder features produced only about 12% overlap between geometric alignment and behavioral gradient sensitivity.
Evidence: source-2026-09-17-004
Sources
- arXiv preprint 2609.18080arXiv · primary research
Corrections
No corrections have been recorded for this story.