TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety security

Diffusion Models Carried Safety in a Few Neurons

Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.

Published Updated Story ID: mp-2026-08-10-026
Read the complete editionStory JSON

Summary

Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.

Researchers tested diffusion language models that generate by iterative denoising rather than next-token prediction. They report that pruning mapped safety neurons raised attack success rates from 2.6 to 73.8 percent on LLaDA and from 1.9 to 86.6 percent on Dream; a separate offline steering method transferred attacks to several targets with reported success as high as 86.9 percent. These figures come from the authors' threat model and preprint codebase and should not be generalized to every diffusion model.

Why it matters

Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.

Limits and context

  • These figures come from the authors' threat model and preprint codebase and should not be generalized to every diffusion model.

Key claims

  1. Mechanistic attacks mapped and pruned sparse safety features inherited from autoregressive parent models.

    Qualification: These figures come from the authors' threat model and preprint codebase and should not be generalized to every diffusion model.

    Evidence: source-2026-08-10-015

Sources

  1. arXiv preprint 2608.07430arXiv · primary research

Corrections

No corrections have been recorded for this story.