TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety

Attribution Could Not Reliably Filter Subliminal Learning

One gradient method mitigated part of the effect at token level, but results changed across models and preferences.

Published Updated Story ID: mp-2026-09-19-005
Read the complete editionStory JSON

Summary

One gradient method mitigated part of the effect at token level, but results changed across models and preferences.

Subliminal learning can transmit behavioral traits through training examples that do not state those traits, limiting semantic filters. The study tested GradCos, a contrastive variant and EK-FAC against divergence tokens across three models. Token-level EK-FAC removed a meaningful part of the effect, while the other attribution methods offered little benefit and generally trailed the counterfactual-teacher baseline. Whole-sample filtering was weaker for every method, and no approach worked consistently across model-preference combinations.

Why it matters

One gradient method mitigated part of the effect at token level, but results changed across models and preferences.

Limits and context

  • Subliminal learning can transmit behavioral traits through training examples that do not state those traits, limiting semantic filters.

Key claims

  1. One gradient method mitigated part of the effect at token level, but results changed across models and preferences.

    Qualification: Subliminal learning can transmit behavioral traits through training examples that do not state those traits, limiting semantic filters.

    Evidence: source-2026-09-19-005

Sources

  1. arXiv preprint 2609.20027arXiv · primary research

Corrections

No corrections have been recorded for this story.