
Hiroki Naganuma
Research Engineer (Deep Learning Software Engineer), NVIDIA · Convergence for High-Efficiency Formulations (CHEF)
Hiroki Naganuma is a Research Engineer (formal job title: Deep Learning Software Engineer) on the CHEF team at NVIDIA. His research focuses on large-scale optimization, distributed deep learning, and efficient training of large language models.
hiroki11x@gmail.com · ORCID · Google Scholar · researchmap
Research areas
- large-scale optimization
- distributed deep learning
- large language model training
- high-performance computing
- semi-synchronous training
- large-batch training
- critical batch size
- non-Euclidean optimization
Selected and recent publications
- Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark — Transactions on Machine Learning Research (2026)
- Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent — International Conference on Machine Learning (2026)
- What do near-optimal learning rate schedules look like? — Transactions on Machine Learning Research (2026)
- Which Geometry on Which Layer? A Principled Criterion for Mixed-Optimizer Training — Under Review (2026)
- Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization — Under Review (2026)
- The Geometry of Spectral Gradient Descent: Layerwise Criteria for SignSGD vs SpecSGD — ICLR 2026 Workshop on Geometry-grounded Representation Learning and Generative Modeling (2026)
- On Fairness of Task Arithmetic: The Role of Task Vectors — International Conference on Learning Representations (2026)
- DiTaC: Conditioning Task Vectors via Distillation for Robust Model Merging — International Conference on Learning Representations (2026)
- Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training — Transactions on Machine Learning Research (2025)
- Convergence Bound and Critical Batch Size of Muon Optimizer — Under Review (2025)