# 長沼 大樹 / Hiroki Naganuma

> 長沼 大樹 (Hiroki Naganuma) は NVIDIA の Research Engineer (Deep Learning Software Engineer) です。2026 年 10 月より慶應義塾大学の Visiting Assistant Professor を兼任します。大規模最適化、分散深層学習、大規模言語モデルの効率的な学習に取り組んでいます。

Canonical HTML: https://hiroki11x.github.io/ja/

## 現在の所属

- Research Engineer (Deep Learning Software Engineer), NVIDIA — Convergence for High-Efficiency Formulations (CHEF)

## 研究領域

- [スケーラブルな事前学習と分散学習](https://hiroki11x.github.io/ja/research/#scalable-pretraining): 分散・半同期・大規模バッチ手法を通じて、大規模言語モデルと基盤モデルの学習効率を高める研究です。
- [機械学習の最適化](https://hiroki11x.github.io/ja/research/#optimization): 学習率スケジュール、非ユークリッド・スペクトル手法、Muon、クリティカルバッチサイズを含む最適化アルゴリズムと学習ダイナミクスの研究です。
- [汎化と基盤モデル](https://hiroki11x.github.io/ja/research/#generalization-foundation-models): 分布シフト下の事前学習済みモデルと基盤モデルを対象に、汎化、キャリブレーション、公平性、モデル合成を研究しています。

## 主要・最新論文

- [From Inner Randomness to Outer Stability](https://hiroki11x.github.io/ja/publications/from-inner-randomness-to-outer-stability/) — Under Review (2026); 機械学習の最適化
- [Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark](https://hiroki11x.github.io/ja/publications/generalization-measures-beyond-iid/) — TMLR 2026 (2026); 汎化と基盤モデル
- [Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent](https://hiroki11x.github.io/ja/publications/adaptive-batch-sizes-using-non-euclidean-gradient-noise-scales/) — ICML 2026 (2026); スケーラブルな事前学習と分散学習; 機械学習の最適化
- [What do near-optimal learning rate schedules look like?](https://hiroki11x.github.io/ja/publications/what-do-near-optimal-learning-rate-schedules-look-like/) — TMLR 2026 (2026); 機械学習の最適化
- [Which Geometry on Which Layer? A Principled Criterion for Mixed-Optimizer Training](https://hiroki11x.github.io/ja/publications/which-geometry-on-which-layer/) — Under Review (2026); スケーラブルな事前学習と分散学習; 機械学習の最適化
- [Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization](https://hiroki11x.github.io/ja/publications/orth-dion/) — Under Review (2026); スケーラブルな事前学習と分散学習; 機械学習の最適化
- [The Geometry of Spectral Gradient Descent: Layerwise Criteria for SignSGD vs SpecSGD](https://hiroki11x.github.io/ja/publications/geometry-of-spectral-gradient-descent/) — ICLR 2026 Workshop (2026); 機械学習の最適化
- [On Fairness of Task Arithmetic: The Role of Task Vectors](https://hiroki11x.github.io/ja/publications/fairness-of-task-arithmetic/) — ICLR 2026 (2026); 汎化と基盤モデル
- [DiTaC: Conditioning Task Vectors via Distillation for Robust Model Merging](https://hiroki11x.github.io/ja/publications/ditac-conditioning-task-vectors-via-distillation/) — ICLR 2026 (2026); 汎化と基盤モデル
- [Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training](https://hiroki11x.github.io/ja/publications/pseudo-asynchronous-local-sgd/) — TMLR 2025 (2025); スケーラブルな事前学習と分散学習

## 研究者プロフィール

- [ORCID](https://orcid.org/0000-0002-4595-8381)
- [Google Scholar](https://scholar.google.com/citations?user=xx7O2voAAAAJ)
- [researchmap](https://researchmap.jp/hiroki11x)
- [OpenAlex](https://openalex.org/A5010523882)
- [DBLP](https://dblp.org/pid/206/0082)
- [J-GLOBAL](https://jglobal.jst.go.jp/detail?JGLOBAL_ID=201901008501473681)
