benchmarks evals
Forty-Seven Unlearning Checkpoints Moved Without Changing a Weight
Refitting batch-normalization statistics shifted published measurements beyond seed spread in 47 of 221 released checkpoints.
Summary
Refitting batch-normalization statistics shifted published measurements beyond seed spread in 47 of 221 released checkpoints.
An unlearning verdict can depend on batch-normalization statistics that are shipped with the model but not written by gradient descent. At bit-identical weights, refitting those statistics on retained data moved 47 of 221 released checkpoints beyond the variability shown by their release seeds. Twelve published verdicts crossed their decision line; four exceeded a measured recalibration budget and two did so on every replicate. Swapping retained for removed records inside a fixed fitting pool barely moved the result, arguing against leftover removed data as the cause. The audit applies to batch-normalized vision models and calls for releases to state the fitting convention beside each number.
Why it matters
Refitting batch-normalization statistics shifted published measurements beyond seed spread in 47 of 221 released checkpoints.
Limits and context
- An unlearning verdict can depend on batch-normalization statistics that are shipped with the model but not written by gradient descent.
Key claims
Refitting batch-normalization statistics shifted published measurements beyond seed spread in 47 of 221 released checkpoints.
Qualification: An unlearning verdict can depend on batch-normalization statistics that are shipped with the model but not written by gradient descent.
Evidence: source-2026-09-11-007
Sources
- arXiv preprint 2609.11490arXiv · primary research
Corrections
No corrections have been recorded for this story.