国際会議論文
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
International Conference on Machine Learning · 2026年7月
- arXiv
- 2602.03001
概要(原文)
To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune. Existing adaptive strategies based on gradient noise scale (GNS) offer a principled alternative. However, their assumption of SGD's Euclidean geometry creates a fundamental mismatch with popular optimizers based on generalized norms, such as signSGD / Signum ($\ell_\infty$) and stochastic spectral descent (specSGD) / Muon ($\mathcal{S}_\infty$). In this work, we derive gradient noise scales for signSGD and specSGD that naturally emerge from the geometry of their respective dual norms. To practically estimate these non-Euclidean metrics, we propose an efficient variance estimation procedure that leverages the local mini-batch gradients on different ranks in distributed data-parallel systems. Our experiments demonstrate that adaptive batch size strategies using non-Euclidean GNS enable us to match the validation loss of constant-batch baselines while reducing training steps by up to 66\% for Signum and Muon on a 160 million parameter Llama model.
研究の要点
- 課題
- ユークリッド幾何に基づくgradient noise scaleは、SignumやMuon系のスペクトル最適化と幾何的に整合しません。
- 手法
- 各最適化手法の双対ノルムから非ユークリッドなnoise scaleを導出し、分散学習の各rankのmini-batch勾配から効率的に推定します。
- 主結果
- 1.6億パラメータのLlamaで、constant-batchと同等のvalidation lossを保ちながら、SignumとMuonの学習stepを最大66%削減しました。
- 意義
- 適応的batch sizeを現代的な非ユークリッド最適化の幾何と整合させ、分散システムで直接利用できる形にします。
- 限界
- 主要結果は1.6億パラメータのモデルでの報告であり、それを超える規模での挙動には追加研究が必要です。
関連リンク
引用
Hiroki Naganuma, Shagun Gupta, Youssef Briki, Ioannis Mitliagkas, Irina Rish, Parameswaran Raman, Hao-Jun Michael Shi. “Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent.” International Conference on Machine Learning, 2026.
@inproceedings{Naganuma2026Adaptive,
title = {Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent},
author = {Hiroki Naganuma and Shagun Gupta and Youssef Briki and Ioannis Mitliagkas and Irina Rish and Parameswaran Raman and Hao-Jun Michael Shi},
year = {2026},
booktitle = {International Conference on Machine Learning},
eprint = {2602.03001},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2602.03001}
}