# Hiroki Naganuma

> Hiroki Naganuma is a Research Engineer (Deep Learning Software Engineer) at NVIDIA. He works on large-scale optimization, distributed deep learning, and efficient training of large language models.

Canonical HTML: https://hiroki11x.github.io/

## Current affiliation

- Research Engineer (Deep Learning Software Engineer), NVIDIA — Convergence for High-Efficiency Formulations (CHEF)

## Research areas

- [Scalable Pretraining and Distributed Learning](https://hiroki11x.github.io/research/#scalable-pretraining): Efficient training of large language and foundation models through distributed, semi-synchronous, and large-batch methods.
- [Optimization for Machine Learning](https://hiroki11x.github.io/research/#optimization): Optimization algorithms and training dynamics, including learning-rate schedules, non-Euclidean and spectral methods, Muon, and critical batch size analysis.
- [Generalization and Foundation Models](https://hiroki11x.github.io/research/#generalization-foundation-models): Generalization, calibration, fairness, and model composition for pre-trained and foundation models under distribution shift.

## Selected and recent publications

- [Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark](https://hiroki11x.github.io/publications/generalization-measures-beyond-iid/) — TMLR 2026 (2026); Generalization and Foundation Models
- [Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent](https://hiroki11x.github.io/publications/adaptive-batch-sizes-using-non-euclidean-gradient-noise-scales/) — ICML 2026 (2026); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning
- [What do near-optimal learning rate schedules look like?](https://hiroki11x.github.io/publications/what-do-near-optimal-learning-rate-schedules-look-like/) — TMLR 2026 (2026); Optimization for Machine Learning
- [Which Geometry on Which Layer? A Principled Criterion for Mixed-Optimizer Training](https://hiroki11x.github.io/publications/which-geometry-on-which-layer/) — Under Review (2026); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning
- [Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization](https://hiroki11x.github.io/publications/orth-dion/) — Under Review (2026); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning
- [The Geometry of Spectral Gradient Descent: Layerwise Criteria for SignSGD vs SpecSGD](https://hiroki11x.github.io/publications/geometry-of-spectral-gradient-descent/) — ICLR 2026 Workshop (2026); Optimization for Machine Learning
- [On Fairness of Task Arithmetic: The Role of Task Vectors](https://hiroki11x.github.io/publications/fairness-of-task-arithmetic/) — ICLR 2026 (2026); Generalization and Foundation Models
- [DiTaC: Conditioning Task Vectors via Distillation for Robust Model Merging](https://hiroki11x.github.io/publications/ditac-conditioning-task-vectors-via-distillation/) — ICLR 2026 (2026); Generalization and Foundation Models
- [Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training](https://hiroki11x.github.io/publications/pseudo-asynchronous-local-sgd/) — TMLR 2025 (2025); Scalable Pretraining and Distributed Learning
- [Convergence Bound and Critical Batch Size of Muon Optimizer](https://hiroki11x.github.io/publications/convergence-bound-and-critical-batch-size-of-muon/) — Under Review (2025); Scalable Pretraining and Distributed Learning; Optimization for Machine Learning

## Researcher profiles

- [ORCID](https://orcid.org/0000-0002-4595-8381)
- [Google Scholar](https://scholar.google.com/citations?user=xx7O2voAAAAJ)
- [researchmap](https://researchmap.jp/hiroki11x)
- [OpenAlex](https://openalex.org/A5010523882)
- [DBLP](https://dblp.org/pid/206/0082)
