Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon

General News

Summary

This article introduces Gram Newton-Schulz, a faster and more hardware-aware version of the Newton-Schulz routine used inside the Muon optimizer. The method shifts computation from rectangular matrices to smaller Gram matrices, which cuts FLOPs and reduces optimizer step time on Hopper and Blackwell GPUs. The authors also address half-precision instability with a restart strategy and custom symmetric GEMM kernels. In experiments, GramMuon preserves training quality while reducing orthogonalization runtime by 40-50% and speeding up overall optimizer work in large language model training.

Classifications

industries
No industries detected
applications
No applications detected

AskAI Classifications

Labels
No AI classifications detected

Linked Companies