日本語

Hiroki Naganuma / 長沼 大樹

Research Overview

Research on efficient and scalable machine learning, connecting optimization theory with large-scale training systems.

Scalable Pretraining and Distributed Learning

Efficient training of large language and foundation models through distributed, semi-synchronous, and large-batch methods.

Topics: large language model pretraining · distributed deep learning · large-batch training · high-performance computing

Selected related publications

Research notes

Optimization for Machine Learning

Optimization algorithms and training dynamics, including learning-rate schedules, non-Euclidean and spectral methods, Muon, and critical batch size analysis.

Topics: machine learning optimization · Muon optimizer · non-Euclidean optimization · learning-rate schedules · critical batch size

Selected related publications

Muon optimizer research note

Generalization and Foundation Models

Generalization, calibration, fairness, and model composition for pre-trained and foundation models under distribution shift.

Topics: foundation models · out-of-distribution generalization · model calibration · distribution shift · model merging

Selected related publications