Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon
Summary
This article introduces Gram Newton-Schulz, a faster and more hardware-aware version of the Newton-Schulz routine used inside the Muon optimizer. The method shifts computation from rectangular matrices to smaller Gram matrices, which cuts FLOPs and reduces optimizer step time on Hopper and Blackwell GPUs. The authors also address half-precision instability with a restart strategy and custom symmetric GEMM kernels. In experiments, GramMuon preserves training quality while reducing orthogonalization runtime by 40-50% and speeding up overall optimizer work in large language model training.
Classifications
industries
No industries detected
applications
No applications detected
AskAI Classifications
Labels
No AI classifications detected