# Hiroki Naganuma / 長沼 大樹 > Official research profile for Hiroki Naganuma, a Research Engineer (formal job title: Deep Learning Software Engineer) at NVIDIA, working on machine learning optimization and efficient large-scale training. The canonical HTML pages and their structured data are authoritative. This LLM-readable index is generated from the same identity, research-topic, and publication records. It contains no private affiliations or unsupported claims. ## Identity - [Official English profile](https://hiroki11x.github.io/): Hiroki Naganuma is an NVIDIA Research Engineer studying machine learning optimization, scalable LLM pretraining, distributed training, and the Muon optimizer. - [Official Japanese profile](https://hiroki11x.github.io/ja/): 長沼 大樹 (Hiroki Naganuma) はNVIDIAのResearch Engineerで、機械学習の最適化、LLM・基盤モデルのスケーラブルな事前学習、分散学習、Muonオプティマイザを研究しています。 - Japanese name: 長沼 大樹 - Current role category: Research Engineer - Formal job title: Deep Learning Software Engineer at NVIDIA; team: Convergence for High-Efficiency Formulations (CHEF) - Degrees: B.Eng. / M.Eng. / Ph.D. - Research: Large-scale optimization, distributed deep learning, and efficient training of large language models - [Ph.D. thesis: Toward Efficient and Scalable Optimization: Theoretical Insights and Practical Challenges](https://umontreal.scholaris.ca/items/9c19791c-97c2-46d4-8769-3be805e65bda); [DOI](https://doi.org/10.71781/34808) - [ORCID](https://orcid.org/0000-0002-4595-8381) · [Google Scholar](https://scholar.google.com/citations?user=xx7O2voAAAAJ) · [researchmap](https://researchmap.jp/hiroki11x) · [OpenAlex](https://openalex.org/A5010523882) · [DBLP](https://dblp.org/pid/206/0082) ## Research areas - [Scalable Pretraining and Distributed Learning](https://hiroki11x.github.io/research/#scalable-pretraining): Efficient training of large language and foundation models through distributed, semi-synchronous, and large-batch methods. Supporting publications: [Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training](https://hiroki11x.github.io/publications/pseudo-asynchronous-local-sgd/); [Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent](https://hiroki11x.github.io/publications/adaptive-batch-sizes-using-non-euclidean-gradient-noise-scales/); [Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization](https://hiroki11x.github.io/publications/orth-dion/). - [Optimization for Machine Learning](https://hiroki11x.github.io/research/#optimization): Optimization algorithms and training dynamics, including learning-rate schedules, non-Euclidean and spectral methods, Muon, and critical batch size analysis. Supporting publications: [What do near-optimal learning rate schedules look like?](https://hiroki11x.github.io/publications/what-do-near-optimal-learning-rate-schedules-look-like/); [Convergence Bound and Critical Batch Size of Muon Optimizer](https://hiroki11x.github.io/publications/convergence-bound-and-critical-batch-size-of-muon/); [No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths](https://hiroki11x.github.io/publications/no-wrong-turns-neural-network-optimization-paths/). - [Generalization and Foundation Models](https://hiroki11x.github.io/research/#generalization-foundation-models): Generalization, calibration, fairness, and model composition for pre-trained and foundation models under distribution shift. Supporting publications: [An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration](https://hiroki11x.github.io/publications/pre-trained-model-selection-for-ood-generalization-and-calibration/); [Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark](https://hiroki11x.github.io/publications/generalization-measures-beyond-iid/); [Empirical Study on Optimizer Selection for Out-of-Distribution Generalizations](https://hiroki11x.github.io/publications/optimizer-selection-for-ood-generalization/). ## Primary pages - [Research overview](https://hiroki11x.github.io/research/): research areas with supporting primary publications - [Biography](https://hiroki11x.github.io/bio/): current role, education, research experience, honors, and public writing - [Publications](https://hiroki11x.github.io/publications/): complete categorized publication list - [Machine-readable profile](https://hiroki11x.github.io/profile.json): identity, research areas, evidence links, and persistent profiles - [Publication data](https://hiroki11x.github.io/publications.json): structured catalog - [BibTeX](https://hiroki11x.github.io/publications.bib): generated citations - [CV](https://hiroki11x.github.io/files/CV_HirokiNAGANUMA.pdf): public curriculum vitae - [Research notes](https://hiroki11x.github.io/posts/research_topics/): separately maintained technical notes, including Muon and large-scale training topics ## Publication pages - [What do near-optimal learning rate schedules look like?](https://hiroki11x.github.io/publications/what-do-near-optimal-learning-rate-schedules-look-like/): TMLR 2026 (2026); Optimization for Machine Learning - [Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training](https://hiroki11x.github.io/publications/pseudo-asynchronous-local-sgd/): TMLR 2025 (2025); Scalable Pretraining and Distributed Learning - [An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration](https://hiroki11x.github.io/publications/pre-trained-model-selection-for-ood-generalization-and-calibration/): TMLR 2025 (2025); Generalization and Foundation Models - [Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration](https://hiroki11x.github.io/publications/geometric-insights-into-focal-loss/): PRL (2025); Optimization for Machine Learning; Generalization and Foundation Models - [Towards Understanding Variants of Invariant Risk Minimization from the Perspective of Calibration](https://hiroki11x.github.io/publications/invariant-risk-minimization-variants-and-calibration/): TMLR 2024 (2024); Generalization and Foundation Models - [Empirical Study on Optimizer Selection for Out-of-Distribution Generalizations](https://hiroki11x.github.io/publications/optimizer-selection-for-ood-generalization/): TMLR 2023 (2023); Optimization for Machine Learning; Generalization and Foundation Models - [Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent](https://hiroki11x.github.io/publications/adaptive-batch-sizes-using-non-euclidean-gradient-noise-scales/): ICML 2026 (2026); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning - [On Fairness of Task Arithmetic: The Role of Task Vectors](https://hiroki11x.github.io/publications/fairness-of-task-arithmetic/): ICLR 2026 (2026); Generalization and Foundation Models - [DiTaC: Conditioning Task Vectors via Distillation for Robust Model Merging](https://hiroki11x.github.io/publications/ditac-conditioning-task-vectors-via-distillation/): ICLR 2026 (2026); Generalization and Foundation Models - [Mastering Task Arithmetic: τJp as a Key Indicator for Weight Disentanglement](https://hiroki11x.github.io/publications/mastering-task-arithmetic-tau-jp/): ICLR 2025 (2025); Generalization and Foundation Models - [No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths](https://hiroki11x.github.io/publications/no-wrong-turns-neural-network-optimization-paths/): ICML 2024 (2024); Optimization for Machine Learning - [How Image Corruption and Perturbation Affect Out-Of-Distribution Generalization and Calibration](https://hiroki11x.github.io/publications/image-corruption-ood-generalization-and-calibration/): IJCNN 2023 (2023); Generalization and Foundation Models - [Conjugate Gradient Method for Generative Adversarial Networks](https://hiroki11x.github.io/publications/conjugate-gradient-method-for-generative-adversarial-networks/): AISTATS 2023 (2023); Optimization for Machine Learning - [Optimal Transport Meets Noisy Label Robust Loss and MixUp Regularization for Domain Adaptation](https://hiroki11x.github.io/publications/optimal-transport-noisy-label-robust-loss-and-mixup/): CoLLAs 2022 (2022); Generalization and Foundation Models - [Accelerating Convolutional Neural Networks Using Low Precision Arithmetic](https://hiroki11x.github.io/publications/accelerating-convolutional-neural-networks-using-low-precision-arithmetic/): HPC Asia 2018 (2018); Scalable Pretraining and Distributed Learning - [Accelerating Matrix Multiplication in Deep Learning by using Low-Rank Approximation](https://hiroki11x.github.io/publications/accelerating-matrix-multiplication-using-low-rank-approximation/): HPCS 2017 (2017); Scalable Pretraining and Distributed Learning - [Which Geometry on Which Layer? A Principled Criterion for Mixed-Optimizer Training](https://hiroki11x.github.io/publications/which-geometry-on-which-layer/): Under Review (2026); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning - [Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization](https://hiroki11x.github.io/publications/orth-dion/): Under Review (2026); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning - [Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark](https://hiroki11x.github.io/publications/generalization-measures-beyond-iid/): TMLR 2026 (2026); Generalization and Foundation Models - [Convergence Bound and Critical Batch Size of Muon Optimizer](https://hiroki11x.github.io/publications/convergence-bound-and-critical-batch-size-of-muon/): Under Review (2025); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning - [When Does Alignment Help? A Comparative Study of DCCA and Fusion-Based Approaches for Multi-modal Chest X-ray Analysis](https://hiroki11x.github.io/publications/when-does-alignment-help-multimodal-chest-x-ray-analysis/): SSRN (2025); Generalization and Foundation Models - [Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation](https://hiroki11x.github.io/publications/augmenting-ner-datasets-with-llms/): arXiv (2024); Generalization and Foundation Models - [Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime](https://hiroki11x.github.io/publications/takeuchi-information-criteria-for-dnns-close-to-ntk/): arXiv (2021); Optimization for Machine Learning; Generalization and Foundation Models - [The Geometry of Spectral Gradient Descent: Layerwise Criteria for SignSGD vs SpecSGD](https://hiroki11x.github.io/publications/geometry-of-spectral-gradient-descent/): ICLR 2026 Workshop (2026); Optimization for Machine Learning - [Smoothness-Adaptive Sharpness Aware Minimization for Finding Flatter Minima](https://hiroki11x.github.io/publications/smoothness-adaptive-sharpness-aware-minimization/): ICLR 2024 Workshop (2024); Optimization for Machine Learning; Generalization and Foundation Models - [Story-to-Images Translation: Leveraging Diffusion Models and Large Language Models for Sequence Image Generation](https://hiroki11x.github.io/publications/story-to-images-translation/): NarSUM 2023 (2023); Generalization and Foundation Models - [On the Interplay of Curvature, Calibration and Out-of-Distribution Generalization: Insights from SAM and Focal Loss Analyses](https://hiroki11x.github.io/publications/curvature-calibration-and-ood-generalization/): UNCV 2023 (2023); Optimization for Machine Learning; Generalization and Foundation Models - [Necessary and Sufficient Hypothesis of Curvature: Understanding Connection Between Out-of-Distribution Generalization and Calibration](https://hiroki11x.github.io/publications/necessary-and-sufficient-hypothesis-of-curvature/): ICLR 2023 Workshop (2023); Optimization for Machine Learning; Generalization and Foundation Models - [Towards Understanding the Relationship of Batch Size and Iterations in Deep Learning](https://hiroki11x.github.io/publications/batch-size-and-iterations-in-deep-learning/): MLSS 2020 (2020); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning - [On Empirical Analysis of Layer-wise Learning Rate Schedule](https://hiroki11x.github.io/publications/layer-wise-learning-rate-schedule/): ACML 2019 Workshop (2019); Optimization for Machine Learning - [A Performance Improvement Approach for Second-Order Optimization in Large Mini-batch Training](https://hiroki11x.github.io/publications/second-order-optimization-in-large-mini-batch-training/): HPML 2019 (2019); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning - [Noise Injection Leads to Better Generalization in Large Mini-Batch Training](https://hiroki11x.github.io/publications/noise-injection-for-large-mini-batch-training/): Tokyo Tech-Stony Brook Joint Meeting (2019); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning Last generated: 2026-09-12