safety security
The Attack Steered Noise Toward a Face Model’s Memory
A white-box inversion method injects identity gradients through a flow-matching trajectory to reconstruct representative target-class faces.
Summary
A white-box inversion method injects identity gradients through a flow-matching trajectory to reconstruct representative target-class faces.
SFMI first trains an unconditional flow-matching prior over faces, then backpropagates through the target recognition model to guide intermediate samples toward a selected identity class. Under an identity-disjoint CelebA evaluation against ArcFace, the paper reports 0.9248 attack accuracy, FID 22.61 and LPIPS 0.3874, with competitive results across additional targets. The images are representative reconstructions rather than recovered source photographs, but the experiment illustrates the privacy exposure created by white-box access to recognition models.
Why it matters
A white-box inversion method injects identity gradients through a flow-matching trajectory to reconstruct representative target-class faces.
Limits and context
No additional limitation was separately recorded.
Key claims
A white-box inversion method injects identity gradients through a flow-matching trajectory to reconstruct representative target-class faces.
Evidence: source-2026-08-18-011
Sources
- arXiv preprint 2608.16791arXiv · primary research
Corrections
No corrections have been recorded for this story.