chips infrastructure
Four-Bit Training Dropped the Extra Rotation
Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.
Summary
Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.
The recipe pairs E2M1 payloads with wider-range unsigned E5M3 scales, uses periodic tensor scaling and applies stochastic rounding selectively to backward gradients. In software-emulated training of an 8B Nemotron-H model for nearly 190 billion tokens, the authors report lower final-window and held-out validation loss than their Transformer Engine NVFP4 comparison. A separate ablation that removed the transform and final-block exemption raised measured model-body throughput by 21.2%, motivating native hardware support rather than proving a universal speedup.
Why it matters
Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.
Limits and context
No additional limitation was separately recorded.
Key claims
Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.
Evidence: source-2026-09-03-005
Sources
- arXiv preprint 2609.02846arXiv · primary research
Corrections
No corrections have been recorded for this story.