safety security
Diffusion Models Carried Safety in a Few Neurons
Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.
Summary
Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.
Researchers tested diffusion language models that generate by iterative denoising rather than next-token prediction. They report that pruning mapped safety neurons raised attack success rates from 2.6 to 73.8 percent on LLaDA and from 1.9 to 86.6 percent on Dream; a separate offline steering method transferred attacks to several targets with reported success as high as 86.9 percent. These figures come from the authors' threat model and preprint codebase and should not be generalized to every diffusion model.
Why it matters
Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.
Limits and context
- These figures come from the authors' threat model and preprint codebase and should not be generalized to every diffusion model.
Key claims
Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.
Qualification: These figures come from the authors' threat model and preprint codebase and should not be generalized to every diffusion model.
Evidence: source-2026-08-10-015
Sources
- arXiv preprint 2608.07430arXiv · primary research
Corrections
No corrections have been recorded for this story.