TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

chips infrastructure

Four-Bit Training Dropped the Extra Rotation

Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.

Published Updated Story ID: mp-2026-09-03-005
Read the complete editionStory JSON

Summary

Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.

The recipe pairs E2M1 payloads with wider-range unsigned E5M3 scales, uses periodic tensor scaling and applies stochastic rounding selectively to backward gradients. In software-emulated training of an 8B Nemotron-H model for nearly 190 billion tokens, the authors report lower final-window and held-out validation loss than their Transformer Engine NVFP4 comparison. A separate ablation that removed the transform and final-block exemption raised measured model-body throughput by 21.2%, motivating native hardware support rather than proving a universal speedup.

Why it matters

Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Unsigned E5M3 block scales replace randomized Hadamard transforms and BF16 final-layer exemptions in one FP4 pretraining recipe.

    Evidence: source-2026-09-03-005

Sources

  1. arXiv preprint 2609.02846arXiv · primary research

Corrections

No corrections have been recorded for this story.