safety
Attribution Could Not Reliably Filter Subliminal Learning
One gradient method mitigated part of the effect at token level, but results changed across models and preferences.
Summary
One gradient method mitigated part of the effect at token level, but results changed across models and preferences.
Subliminal learning can transmit behavioral traits through training examples that do not state those traits, limiting semantic filters. The study tested GradCos, a contrastive variant and EK-FAC against divergence tokens across three models. Token-level EK-FAC removed a meaningful part of the effect, while the other attribution methods offered little benefit and generally trailed the counterfactual-teacher baseline. Whole-sample filtering was weaker for every method, and no approach worked consistently across model-preference combinations.
Why it matters
One gradient method mitigated part of the effect at token level, but results changed across models and preferences.
Limits and context
- Subliminal learning can transmit behavioral traits through training examples that do not state those traits, limiting semantic filters.
Key claims
One gradient method mitigated part of the effect at token level, but results changed across models and preferences.
Qualification: Subliminal learning can transmit behavioral traits through training examples that do not state those traits, limiting semantic filters.
Evidence: source-2026-09-19-005
Sources
- arXiv preprint 2609.20027arXiv · primary research
Corrections
No corrections have been recorded for this story.