ICML2026 参加記
Day1 (7/6 Mon)
HND->GMP
- Is numerical optimization theory irrelevant to machine learning practice in 2026?
- https://www.cs.ubc.ca/~schmidtm/Documents/2026_ICML_Tutorial.pdf
- Mark Schmidt (UBC)
Day2 (7/7 Tue)
- 朝食 w/ Atish Agarwal (GDM)
-
コーヒーチャット w/ Matsunaga Daiki (KAIST)
- ポスターセッション
- Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
- URL: https://icml.cc/virtual/2026/poster/62214
- 著者: Hee-Sung Kim ⋅ Hyeonseong Kim ⋅ Sungyoon Lee
- まとめ: ラベルなしで計算できる local inconsistency を汎化指標として導入し、IAM に組み込むことで supervised、semi-supervised、self-supervised learning の汎化性能を改善する研究。
- FOAM: Frequency and Operator-Error Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo
- URL: https://icml.cc/virtual/2026/poster/63117
- 著者: Kyunghun Nam ⋅ Sumyeong Ahn
- まとめ: Shampoo の stale preconditioner update が生む効率と安定性の trade-off を解析し、damping と eigendecomposition frequency を適応制御する FOAM により wall-clock training を高速化する研究。
- Flat Minima and Generalization: Insights from Stochastic Convex Optimization
- URL: https://icml.cc/virtual/2026/poster/61057
- 著者: Matan Schliserman ⋅ Shira Vansover-Hager ⋅ Tomer Koren
- まとめ: Stochastic convex optimization の設定で flat minima が必ずしも良い汎化を保証せず、sharp minima が最適に汎化し得ることを示し、SA-GD と SAM の限界も解析する研究。
- One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
- URL: https://icml.cc/virtual/2026/poster/62737
- 著者: Philip Zmushko ⋅ Egor Petrov ⋅ Nursultan Abdullaev ⋅ Khrushchev Mikhail ⋅ Samuel Horváth
- まとめ: Asynchronous pipeline parallelism における one-step gradient delay は optimizer 選択と Error-Feedback により緩和でき、10B パラメータ級 LLM でも synchronous training との差を縮められることを示す研究。
- Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
- URL: https://icml.cc/virtual/2026/poster/66517
- 著者: Hyunji Jung ⋅ Sungbin Shin ⋅ Namhoon Lee
- まとめ: Asynchronous pipeline parallelism で pipeline depth とともに増える gradient staleness を basis rotation で補正し、1B パラメータ LLM で少ない iteration により同等の loss に到達する研究。
- On the Role of Batch Size in Stochastic Conditional Gradient Methods
- URL: https://icml.cc/virtual/2026/poster/60616
- 著者: Rustem Islamov ⋅ Roman Machacek ⋅ Aurelien Lucchi ⋅ Antonio Silveti-Falls ⋅ Eduard Gorbunov ⋅ Volkan Cevher
- まとめ: Stochastic conditional gradient methods における batch size、step size、noise の相互作用を解析し、固定 token budget 下での batch size と sequence length schedule の指針を与える研究。
- Why Do We Need Warm-up? A Theoretical Perspective
- URL: https://icml.cc/virtual/2026/poster/63104
- 著者: Foivos Alimisis ⋅ Rustem Islamov ⋅ Aurelien Lucchi
- まとめ: 一般化 $(L_0, L_1)$-smoothness により training 初期の curvature 変化を説明し、learning-rate warm-up が自然に導かれ固定 learning rate より収束を改善することを示す研究。
- Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers
- URL: https://icml.cc/virtual/2026/poster/65724
- 著者: Samuel Erickson ⋅ Mikael Johansson
- まとめ: Asynchronous SGD において straggler による大きな delay への有害な依存性を gradient clipping が取り除けることを、heavy-tailed noise を含む設定で理論的に示す研究。
- Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCo
- URL: https://icml.cc/virtual/2026/poster/66683
- 著者: Chamin Hewa Koneputugodage ⋅ Thalaiyasingam Ajanthan ⋅ Sameera Ramasinghe ⋅ Hadi Mohaghegh Dolatabadi ⋅ Shamane Siriwardhana ⋅ Gil Avraham ⋅ Violetta Shevchenko ⋅ Karol Pajak ⋅ James Snewin ⋅ Alexander Long
- まとめ: DiLoCo の outer synchronization を gossip-style mixing に緩和し、non-blocking communication と blocking communication を分けることで低帯域環境での compute utilization と安定性を両立する研究。
- Adaptive Momentum and Nonlinear Damping for Neural Network Training
- URL: https://icml.cc/virtual/2026/poster/64813
- 著者: Aikaterini Karoni ⋅ Rajit Rajpal ⋅ Benedict Leimkuhler ⋅ Gabriel Stoltz
- まとめ: Parameter ごとの kinetic energy に基づく adaptive momentum と cubic damping を導入し、ViT、BERT、GPT-2 training で Adam と同等以上の性能を示す研究。
- プロジェクトページ: https://github.com/rajit906/higher-order-damping-in-ml-v0
- Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
- クイックチャット w/ Params Raman (Meta)
- レセプション
- Sakana AI Social
Day3 (7/8 Wed)
- How Far Can Quadratics Take Us? Lessons for LLM Pretraining
- https://icml.cc/virtual/2026/invited-talk/67264
-
チャット w/ Charles (IFM)
- Controlled LLM Training on Spectral Sphere
- URL: https://icml.cc/virtual/2026/poster/66212
- 著者: Tian Xie ⋅ Haoming Luo ⋅ Haoyu Tang ⋅ Hu Yiwen ⋅ Jason Liu ⋅ Qingnan Ren ⋅ Yang Wang ⋅ Xin Zhao ⋅ Rui Yan ⋅ Bing Su ⋅ Chong Luo ⋅ Baining Guo
- まとめ: Weight と update に module-wise な spectral constraint を課す Spectral Sphere Optimizer を提案し、AdamW や Muon を上回る LLM training の安定性を実現する研究。
- スライド: https://icml.cc/media/icml-2026/Slides/71055.pdf
- プロジェクトページ: https://github.com/Unakar/Spectral-Sphere-Optimizer
- LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
- URL: https://icml.cc/virtual/2026/poster/63818
- 著者: Andrej Jovanović ⋅ Alex Iacob ⋅ Mher Safaryan ⋅ Ionut-Vlad Modoranu ⋅ Lorenzo Sani ⋅ Shen ⋅ Xinchi Qiu ⋅ Dan Alistarh ⋅ Nicholas Lane
- まとめ: LoRDO により low-rank optimization と infrequent synchronization を組み合わせ、distributed foundation-model training で性能を維持しながら通信量を約10分の1に削減する研究。
- On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
- URL: https://icml.cc/virtual/2026/poster/64227
- 著者: Enea Monzio Compagnoni ⋅ Rustem Islamov ⋅ Frank Proske ⋅ Aurelien Lucchi ⋅ Antonio Orvieto ⋅ Eduard Gorbunov
- まとめ: $(L_0,L_1)$-smoothness の下で SDE の観点から batch noise、communication compression、adaptive / normalized updates が DCSGD と DSignSGD の安定性に与える影響を解析する研究。
- Per-example Gradients: a New Frontier for Understanding and Improving Optimizers
- URL: https://icml.cc/virtual/2026/poster/63313
- 著者: Vincent Roulet ⋅ Atish Agarwala
- まとめ: Per-example / per-token gradient statistics を低コストで取得できることを示し、それを用いて signSGD と Adam preconditioner の設計を再検討する研究。
- Exploiting weight-space symmetries for approximating curvature
- URL: https://icml.cc/virtual/2026/poster/60589
- 著者: Artem Artemev ⋅ Rui Xia ⋅ Benjamin M. Boyd ⋅ Youjing Yu ⋅ Felix Dangel ⋅ Guillaume Hennequin ⋅ Alberto Bernacchia
- まとめ: Loss を不変に保つ weight-space symmetries を利用して single gradient から structured Hessian approximation を構成し、Shampoo や Muon 的な curvature estimate と結びつける研究。
- OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit Biases
- URL: https://icml.cc/virtual/2026/poster/62586
- 著者: Zixiao Wang ⋅ Yifei Shen ⋅ Huishuai Zhang
- まとめ: Orthogonalized updates による spectral control と sign updates による $\ell_\infty$ 的 coordinate control を組み合わせた軽量 optimizer OLion を提案し、AdamW や Muon と同等以上の性能を示す研究。
- GradientStabilizer: Fix the Norm, Not the Gradient
- URL: https://icml.cc/virtual/2026/poster/63695
- 著者: Tianjin Huang ⋅ Zhangyang “Atlas” Wang ⋅ Haotian Hu ⋅ Zhenyu Zhang ⋅ Gaojie Jin ⋅ Xiang Li ⋅ Li Shen ⋅ Jiaxing Shang ⋅ Tianlong Chen ⋅ Ke Li ⋅ Lu Liu ⋅ Qingsong Wen ⋅ Shiwei Liu
- まとめ: Gradient direction を保ったまま running gradient-norm statistics により update magnitude を安定化し、clipping より少ない threshold tuning で training instability を抑える研究。
- Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability
- URL: https://icml.cc/virtual/2026/poster/65256
- 著者: Vincent Bürgin ⋅ Daniel Herbst ⋅ Ya-Wei Eileen Lin ⋅ Stefanie Jegelka
- まとめ: Structural symmetries だけでは説明できない approximate equivalent solutions と linear mode connectivity を、neuron identifiability と effective function classes の観点から解析する研究。
- プロジェクトページ: https://github.com/Vuenc/neuron-identifiability
- WildCat: Near-Linear Attention in Theory and Practice
- URL: https://icml.cc/virtual/2026/poster/61920
- 著者: Tobias Schröder ⋅ Lester Mackey
- まとめ: Randomly pivoted Cholesky で選んだ weighted coreset により attention を近似し、理論保証、near-linear runtime、高精度を両立する研究。
- プロジェクトページ: https://github.com/microsoft/wildcat
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
- URL: https://icml.cc/virtual/2026/poster/61050
- 著者: Guoxia Wang ⋅ Shuai Li ⋅ Congliang Chen ⋅ Jinle Zeng ⋅ Jiabin Yang ⋅ Dianhai Yu ⋅ Yanjun Ma ⋅ Li Shen
- まとめ: 過去の clipped gradient norm に基づく per-tensor adaptive clipping により、LLM pretraining の loss spike と optimizer-state contamination を抑える研究。
- MODEL MERGING SCALING LAWS IN LARGE LANGUAGE MODELS
- URL: https://icml.cc/virtual/2026/poster/63590
- 著者: Yuanyi Wang ⋅ Yanggan Gu ⋅ Yiming Zhang ⋅ Qi Zhou ⋅ Zhaoyi Yan ⋅ Congkai Xie ⋅ Xinyao Wang ⋅ Jianbo Yuan ⋅ Hongxia Yang
- まとめ: Model size と expert count に関する power law を示し、LLM model merging の効果や stopping point を予測しやすくする研究。
- プロジェクトページ: https://infix.io/research/MergingScalingLaw
- Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
- URL: https://icml.cc/virtual/2026/poster/60685
- 著者: Zhijun Tu ⋅ Hanting Chen ⋅ Jian Li ⋅ Yuanyuan Xi ⋅ Siqi Liu ⋅ Chuanjian Liu ⋅ Jie Hu ⋅ Yunhe Wang
- まとめ: Pretrained LLM を forward / backward の両方で full precision から 1-bit representation へ段階的に変換し、ゼロからの高コスト training を避ける研究。
- A Geometry-Based View of Mahalanobis OOD Detection
- URL: https://icml.cc/virtual/2026/poster/62433
- 著者: Denis Janiak ⋅ Jakub Binkowski ⋅ Tomasz Kajdanowicz
- まとめ: Mahalanobis OOD detection の性能が feature geometry に強く依存することを示し、within-class spectral structure と local intrinsic dimensionality を用いて改善する研究。
- When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
- URL: https://icml.cc/virtual/2026/poster/63669
- 著者: Haoran Zhao ⋅ Caren Han ⋅ Eduard Hovy
- まとめ: Vision-language models における rank-1 LoRA の不安定性を modality-gap direction との misalignment として説明し、Gap-Init により low-rank fine-tuning を安定化する研究。
- Path-conditioned training: a principled way to rescale ReLU neural networks
- URL: https://icml.cc/virtual/2026/poster/61112
- 著者: Arthur Lebeurrier ⋅ Titouan Vayer ⋅ Rémi Gribonval
- まとめ: ReLU networks の rescaling symmetries を path-lifting framework で扱い、path space での kernel alignment により training を高速化する研究。
- プロジェクトページ: https://github.com/Artim436/pathcond
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
- URL: https://icml.cc/virtual/2026/poster/61601
- 著者: Jatin Chhugani ⋅ Geonhwa Jeong ⋅ Bor-Yiing Su ⋅ Yunjie Pan ⋅ Hanmei Yang ⋅ Aayush Ankit ⋅ Jiecao Yu ⋅ Summer Deng ⋅ Yunqing Chen ⋅ Nadathur Satish ⋅ Changkyu Kim
- まとめ: Overflow-Aware Scaling と Macro Block Scaling により MXFP4 の quantization error を減らし、hardware 変更なしで NVFP4 に近い精度を実現する研究。
- PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
- URL: https://icml.cc/virtual/2026/poster/62288
- 著者: Shenghao Yang ⋅ Zhichao Wang ⋅ Oleg Balabanov ⋅ N. Benjamin Erichson ⋅ Michael Mahoney
- まとめ: Adaptive polynomial approximation と randomized sketching により、Shampoo や Muon で使われる matrix square-root と orthogonalization の反復計算を高速化する研究。
- LiMuon: Light and Fast Muon Optimizer for Large Models
- URL: https://icml.cc/virtual/2026/poster/61819
- 著者: Feihu Huang ⋅ Yuning Luo ⋅ Songcan Chen
- まとめ: Variance reduction と randomized SVD を用いる large-model optimizer LiMuon を提案し、Muon と比べて memory usage と sample complexity を削減する研究。
- 昼食 w/
- Mahdi and Vincent (Mila)
- ポスターセッションでの会話
- Meta Social
Day4 (7/9 Thu)
- 午前セッション
- MixQuant: Pushing the Limits of Block Rotations in Post-Training Quantization
- URL: https://icml.cc/virtual/2026/poster/61670
- 著者: Sai Sanjeet ⋅ Ian Colbert ⋅ Pablo Monteagudo-Lago ⋅ Giuseppe Franco ⋅ Yaman Umuroglu ⋅ Nicholas Fraser
- まとめ: Block Hadamard rotations の前に permutation で activation mass を再配置する MixQuant により、INT4 post-training quantization を改善する研究。
- プロジェクトページ: https://github.com/Xilinx/brevitas/tree/dev/src/brevitas_examples/papers/perq
- Float8@2bits: Entropy Coding Enables Data-Free Model Compression
- URL: https://icml.cc/virtual/2026/poster/66714
- 著者: Patrick Putzky ⋅ Martin Genzel ⋅ Mattes Mollenhauer ⋅ Sebastian Schulze ⋅ Thomas Wollmann ⋅ Stefan Dietzel
- まとめ: Entropy coding により numerical precision と storage cost を切り離す EntQuant を提案し、70B model を data-free かつ極低 bit-rate で圧縮する研究。
- プロジェクトページ: https://github.com/merantix-momentum/entquant
- From Flat Facts to Sharp Hallucinations: Detecting Stubborn Errors via Gradient Sensitivity
- URL: https://icml.cc/virtual/2026/poster/62350
- 著者: Liew Yee Zhing ⋅ Andrew Tan ⋅ Anwar Majeed
- まとめ: Embedding perturbation 下の gradient-sensitivity signal である EPGS を用いて、LLM の高信頼な stubborn hallucination を検出する研究。
- プロジェクトページ: https://github.com/potato0o/Stubborn_Hallucinations
- Compute When Worth It: Risk Control for Reasoning on a Compute Budget
- URL: https://icml.cc/virtual/2026/poster/61680
- 著者: Xi Wang ⋅ Anushri Suresh ⋅ Alvin Zhang ⋅ Rishi More ⋅ William Jurayj ⋅ Mehrdad Farajtabar ⋅ Daniel Khashabi ⋅ Eric Nalisnick
- まとめ: Reasoning LLM の token-budget decision を distribution-free risk control として定式化し、目標 error rate を守りながら compute を節約する研究。
- Anatomy of Massive Activations and Attention Sinks
- URL: https://icml.cc/virtual/2026/poster/63610
- 著者: Shangwen Sun ⋅ Alfredo Canziani ⋅ Yann LeCun ⋅ Jiachen Zhu
- まとめ: Transformers における massive activations と attention sinks を inference-time mechanism として統一的に説明し、normalization と head dimension の役割を解析する研究。
- プロジェクトページ: https://github.com/savinasun/SpikeSparseSink
- FPTQuant: Function-Preserving Transforms for LLM Quantization
- URL: https://icml.cc/virtual/2026/poster/66544
- 著者: Boris van Breugel ⋅ Yelysei Bondarenko ⋅ Paul Whatmough ⋅ Markus Nagel
- まとめ: Function-preserving Transformer transforms により activation distribution を quantization-friendly に整え、ほぼ inference overhead なしで static INT4 quantization を可能にする研究。
- https://eshyperscale.github.io/
- Zeroth-Order Optimization at the Edge of Stability
- URL: https://icml.cc/virtual/2026/poster/61252
- 著者: Minhak Song ⋅ Liang Zhang ⋅ Bingcong Li ⋅ Niao He ⋅ Michael Muehlebach ⋅ Sewoong Oh
- まとめ: Two-point-estimator zeroth-order methods の安定性が full Hessian spectrum に依存することを示し、deep learning における edge-of-stability behavior を解析する研究。
- RubricRobustness: A Simple Framework for Evaluating the Robustness of Rubrics-Based Benchmarks
- URL: https://icml.cc/virtual/2026/poster/65214
- 著者: Manasi Sharma
- まとめ: Negation、deletion、irrelevant-addition perturbations を加えて rubric-based benchmarks の頑健性を検証し、LLM-as-a-judge evaluation の弱点を明らかにする研究。
- Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
- URL: https://icml.cc/virtual/2026/poster/62621
- 著者: Jake Cunningham ⋅ Nicola Muca Cirone
- まとめ: Attention weights、value magnitude、output-direction alignment を組み合わせた Contribution Weights を導入し、token influence と attention sinks の機能的役割を解析する研究。
- ECO: Quantized Training without Full-Precision Master Weights
- URL: https://icml.cc/virtual/2026/poster/62753
- 著者: Mahdi Nikdan ⋅ Amir Zandieh ⋅ Dan Alistarh ⋅ Vahab Mirrokni
- まとめ: Quantization error を optimizer momentum に戻すことで full-precision master weights なしに quantized parameters を直接更新する ECO を提案する研究。
- Geometric Rate-Distortion Invariance for Domain Generalization
- URL: https://icml.cc/virtual/2026/poster/65890
- 著者: Tong Liu ⋅ Sen Liang ⋅ Shuo Bai
- まとめ: Class-conditional representations を Grassmann manifold 上の subspace として扱い、alignment と complexity control を組み合わせる RDI により domain generalization を改善する研究。
- MixQuant: Pushing the Limits of Block Rotations in Post-Training Quantization
- What will be left for us to work on?
- https://icml.cc/virtual/2026/invited-talk/67274
- https://icml.cc/virtual/2026/oral/71156
- Rustem Islamov ⋅ Michael Crawshaw ⋅ Jeremy Cohen ⋅ Robert Gower
- 午後セッション
- Riemannian Dueling Optimization
- URL: https://icml.cc/virtual/2026/poster/61755
- 著者: Yuxuan Ren ⋅ Abhishek Roy ⋅ Shiqian Ma
- まとめ: Comparison-oracle dueling optimization を Riemannian manifolds に拡張し、RDNGD と projection-free RDFW の complexity result を示す研究。
- Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in Reinforcement Learning
- URL: https://icml.cc/virtual/2026/poster/63355
- 著者: Ruijia Zhang ⋅ Jiacheng Zhu ⋅ Hanqing Zhu ⋅ Laixi Shi
- まとめ: RLVR 向けの orthonormal LoRA initialization を導出し、LoRA-RLPO と LoRA-RLMO により training stability と reasoning performance を改善する研究。
- プロジェクトページ: https://github.com/Richard-ZZZ/geometry-preserving-orthonormal-init-rlvr
- Improved Convergence Analysis of Topology Dependence in Decentralized SGD
- URL: https://icml.cc/virtual/2026/poster/61500
- 著者: Yuki Takezawa ⋅ Anastasiia Koloskova ⋅ Sebastian Stich
- まとめ: Mixing matrix の全 eigenvalues を通じて Decentralized SGD の topology dependence を解析し、spectral gap のみに基づく解析より精密な convergence picture を与える研究。
- L-SR1: Learned Symmetric-Rank-One Preconditioning
- URL: https://icml.cc/virtual/2026/poster/60853
- 著者: Gal Lifshitz ⋅ Shahar Zuler ⋅ Ori Fouks ⋅ Dan Raviv
- まとめ: Classical SR1 に軽量な trainable preconditioning unit を追加し、secant condition に沿った learned second-order optimizer を構築する研究。
- プロジェクトページ: https://gallif.github.io/lsr1/
- Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
- URL: https://icml.cc/virtual/2026/poster/60702
- 著者: Haoming Meng ⋅ Anton Sugolov ⋅ Vardan Papyan
- まとめ: Transformers や ResNets の repeated blocks 間で layer-wise gradient updates を smooth し、preconditioning method として training と generalization を改善する研究。
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
- URL: https://icml.cc/virtual/2026/poster/64390
- 著者: Michael Crawshaw ⋅ Chirag Modi ⋅ Mingrui Liu ⋅ Robert Gower
- まとめ: Muon や Adam-style methods を layer-wise norms と aggregations 上の non-Euclidean gradient descent として整理し、より robust な MuonMax と Momo variants を導く研究。
- An Embarrasingly Simple Way to Optimize Orthogonal Matrices at Scale
- URL: https://icml.cc/virtual/2026/poster/61285
- 著者: Adrián Javaloy ⋅ Antonio Vergari
- まとめ: 少数の matrix products だけで orthogonal constraints を保つ GPU-friendly optimizer POGO を導入し、orthogonal matrix optimization を大規模化する研究。
- プロジェクトページ: https://github.com/adrianjav/pogo
- https://icml.cc/virtual/2026/oral/71156
- Understanding SAM through Minimax Perspective
- URL: https://icml.cc/virtual/2026/poster/65197
- 著者: Ying Chen ⋅ Aoxi Li ⋅ Javad Lavaei
- まとめ: SAM を bilevel minimax problem として解析し、より大きな radius と複数 inner updates を用いる Multi-step SAM の理論的根拠を与える研究。
- Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler
- URL: https://icml.cc/virtual/2026/poster/64318
- 著者: Dimitris Oikonomou ⋅ Nicolas Loizou
- まとめ: SAM-style updates に stochastic Polyak step sizes を導入し、learning-rate tuning を減らす adaptive SAM scheduler を提案する研究。
- General Analysis of LMO-based Optimizers: Beyond Bounded Variance
- URL: https://icml.cc/virtual/2026/poster/61267
- 著者: Egor Shulgin ⋅ Mohamed Awad ⋅ Peter Richtarik ⋅ Eduard Gorbunov
- まとめ: Momentum LMO methods を expected-smoothness の下で統一解析し、Muon と sign-based updates に対する batch-size scaling と optimal momentum rule を導く研究。
- Lions and Muons: Optimization via Stochastic Frank-Wolfe
- URL: https://icml.cc/virtual/2026/poster/62403
- 著者: Maria-Eleni Sfyraki ⋅ Jun-Kun Wang
- まとめ: Lion と Muon を Stochastic Frank-Wolfe の特殊例として解釈し、Frank-Wolfe gap convergence を示すとともに heavy-tailed noise に頑健な variants を開発する研究。
- Can Muon Fine-tune Adam-Pretrained Models?
- URL: https://icml.cc/virtual/2026/poster/64467
- 著者: Xingyu Qu ⋅ Peigeng Huang ⋅ Samuel Horváth
- まとめ: Adam-pretrained models を Muon で fine-tuning する際の optimizer mismatch を解析し、LoRA がその性能劣化を緩和できることを示す研究。
- プロジェクトページ: https://github.com/XingyuQu/muon-finetune
- Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
- URL: https://icml.cc/virtual/2026/poster/60591
- 著者: Fabian Schaipp ⋅ Robert Gower ⋅ Adrien Taylor
- まとめ: Stochastic optimization methods ごとの step-size sensitivity を定量化し、SPS や NGN などの adaptive step-size methods が SGD より頑健な理由を説明する研究。
- Delving into Muon and Beyond: Deep Analysis and Extensions
- URL: https://icml.cc/virtual/2026/poster/60694
- 著者: Xianbiao Qi ⋅ Marco Chen ⋅ Jiaquan Ye ⋅ Yelin He ⋅ Rong Xiao
- まとめ: Muon を spectral transformations の endpoint として捉え、RMS-normalized variants と spectral variants を通じてその安定化効果と限界を解析する研究。
- プロジェクトページ: https://github.com/marcotchen/BeyondMuon
- A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization
- URL: https://icml.cc/virtual/2026/poster/66768
- 著者: Xiyuan Wei ⋅ Linli Zhou ⋅ Bokun Wang ⋅ Chih-Jen Lin ⋅ Tianbao Yang
- まとめ: Compositional entropic risk minimization を min-min dual problem として定式化し、geometry-aware stochastic proximal mirror descent により効率的に最適化する研究。
- プロジェクトページ: https://github.com/Optimization-AI/SCENT
- Riemannian Dueling Optimization
- クイックチャット
- GMP->HND
所感
自分が参加した optimization 関連のセッションでは、Stanford や CMU の affiliation を見る機会が想像より少なかったが、これは単に自分のセッション選択の偏りかもしれない。
印象に残った研究は UBassel、ISTA、EPFL、Flatiron、Mila、Vector、Meta、Google、KAUST など、かなり幅広い所属から出ていた。
自分の研究関心にかなり近い論文もいくつかあり、positioning、related work、citation をもっと丁寧に扱う必要があると感じた。
繰り返し感じたのは、強い研究は頻繁な浅い pivot というより、その領域への深い理解から出てくることが多いということだった。
Optimization の領域では、成果を出している研究者の多くが leading group での研究経験を持っているように見え、カジュアルな会話でも research lineage、internship、共有している技術的文脈がよく話題になる。
異なる研究環境の人たちとの会話からは、強い publication pressure と、より長期的で選択的な研究方向に取り組む余地との間に trade-off があることも感じた。
ここには書けないけど、非公式の採用イベントなどが結構ある印象だった。
今後研究をフォローすべきと個人的に思った若手の人
- ロシア勢・オーストリア勢が強い
最適化理論系
- Fabian Schaipp (Flatiron)
- Robert M. Gower での Guest Researcher
- Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
- Stochastic optimization methods ごとの step-size sensitivity を定量化し、SPS や NGN などの adaptive step-size methods が SGD より頑健な理由を説明する研究。
大規模システム・最適化理論の間系
- Rustem Islamov (UBasel)
- Aurelien Lucchi の PhD, Robert M. Gower とも共著, Dan Alistarh のとこで Master, Peter Richtárik のとこで Undergrad
- On the Role of Batch Size in Stochastic Conditional Gradient Methods
- まとめ: Stochastic conditional gradient methods における batch size、step size、noise の相互作用を解析し、固定 token budget 下での batch size と sequence length schedule の指針を与える研究。
- Why Do We Need Warm-up? A Theoretical Perspective
- まとめ: 一般化 $(L_0, L_1)$-smoothness により training 初期の curvature 変化を説明し、learning-rate warm-up が自然に導かれ固定 learning rate より収束を改善することを示す研究。
- On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
- まとめ: $(L_0,L_1)$-smoothness の下で SDE の観点から batch noise、communication compression、adaptive / normalized updates が DCSGD と DSignSGD の安定性に与える影響を解析する研究。
- Non-Euclidean Gradient Descent Operates at the Edge of Stability
- まとめ: Non-Euclidean gradient descent が full Hessian spectrum に依存する edge-of-stability behavior を示すことを解析する研究。
- Egor Shulgin (KAUST)
- Peter Richtárik のとこで PhD
- General Analysis of LMO-based Optimizers: Beyond Bounded Variance
- まとめ: Momentum LMO methods を expected-smoothness の下で統一解析し、Muon と sign-based updates に対する batch-size scaling と optimal momentum rule を導く研究。
- From Muon to Gluon: Bridging Theory and Practice of LMO-based Optimizers for LLMs
- まとめ: Gluonは、MuonやScionなどのLMOベース最適化手法を実際の層ごとの実装に即して理論化し、より現実的な仮定のもとで最先端の収束保証を与える新しい解析フレームワークを提案する。
- Xi Wang (JHU, MSR AI Frontier)
- John Langford のとこで Intern
- Depth scaling and Muon enable balanced expert usage in MoE training
- MoEのルーターの負荷分散は初期化時の隠れ状態の多様性に本質的に依存し、深いTransformerやMuonは表現崩壊を抑えることで、学習初期から学習中まで負荷分散を改善することを示した。
大規模システム系
- Benjamin Thérien (Mila)
- Irina Rish,Eugene Belilovskyのとこの PhD, GDM の Zachary Charlesのとこで、Student Researcher, FAIR の Aaron Defazio のとこで Intern
- MuLoCo: Muon is a practical inner optimizer for DiLoCo
- まとめ: MuLoCo は、DiLoCo の inner optimizer として Muon を用いることで、通信量を削減しつつ、低帯域環境での compute utilization と安定性を両立する研究。
- Philip Zmushko (ISTA)
- Dan Alistarh のとこの PhD
- One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
- まとめ: Asynchronous pipeline parallelism における one-step gradient delay は optimizer 選択と Error-Feedback により緩和でき、10B パラメータ級 LLM でも synchronous training との差を縮められることを示す研究。
深く理解しとかないとなという論文
- Controlled LLM Training on Spectral Sphere
- 実務では Muon 系はみんな (Flontier AI Lab は) Hyperbolic Sphere 系を使ってる
- Spectral Sphere Optimizer は weight と update に module-wise な spectral constraint を課すことで、AdamW や Muon を上回る LLM training の安定性を実現する研究。
- LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
- LoRDO は low-rank optimization と infrequent synchronization を組み合わせ、distributed foundation-model training で性能を維持しながら通信量を約10分の1に削減する研究。
- Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCo
- DiLoCo の outer synchronization を gossip-style mixing に緩和し、non-blocking communication と blocking communication を分けることで低帯域環境での compute utilization と安定性を両立する研究。
- OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit Biases
- Orthogonalized updates による spectral control と sign updates による $\ell_\infty$ 的 coordinate control を組み合わせた軽量 optimizer OLion を提案し、AdamW や Muon と同等以上の性能を示す研究。
- Exploiting weight-space symmetries for approximating curvature
- Loss を不変に保つ weight-space symmetries を利用して single gradient から structured Hessian approximation を構成し、Shampoo や Muon 的な curvature estimate と結びつける研究。
- Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
- Pretrained LLM を forward / backward の両方で full precision から 1-bit representation へ段階的に変換し、ゼロからの高コスト training を避ける研究。
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
- Overflow-Aware Scaling と Macro Block Scaling により MXFP4 の quantization error を減らし、hardware 変更なしで NVFP4 に近い精度を実現する研究。
- PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
- Adaptive polynomial approximation と randomized sketching により、Shampoo や Muon で使われる matrix square-root と orthogonalization の反復計算を高速化する研究。
- FPTQuant: Function-Preserving Transforms for LLM Quantization
- Function-preserving Transformer transforms により activation distribution を quantization-friendly に整え、ほぼ inference overhead なしで static INT4 quantization を可能にする研究。
- Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
- Transformers や ResNets の repeated blocks 間で layer-wise gradient updates を smooth し、preconditioning method として training と generalization を改善する研究。
細かな研究アイディアへの発展
- A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
- 解説投稿
- pretraining が fine-tuning の inductive bias に与える影響を理論的に解析する研究。
- 下記を補完する内容、これを使って、どう改善するかを考える?
- Can Muon Fine-tune Adam-Pretrained Models?
- Adam-pretrained models を Muon で fine-tuning する際の optimizer mismatch を解析し、LoRA がその性能劣化を緩和できることを示す研究。
- あくまでもこれは探索的な研究、どう緩和するかを考える?
- ECO: Quantized Training without Full-Precision Master Weights
- Quantization error を optimizer momentum に戻すことで full-precision master weights なしに quantized parameters を直接更新する ECO を提案する研究。
- この辺りのアイディア(error feedback)を Muon 系で試すとどうなるかを考える
- GradientStabilizer: Fix the Norm, Not the Gradient
- Gradient direction を保ったまま running gradient-norm statistics により update magnitude を安定化し、clipping より少ない threshold tuning で training instability を抑える研究。
- これは、正しい norm space でやられている? Muon に拡張できる? Hyperbolic Sphere Optimizer に拡張できる? などの疑問が湧いた。
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
- 過去の clipped gradient norm に基づく per-tensor adaptive clipping により、LLM pretraining の loss spike と optimizer-state contamination を抑える研究。
- これは、正しい norm space でやられている? Muon に拡張できる? Hyperbolic Sphere Optimizer に拡張できる? などの疑問が湧いた。
謝辞
ICML への参加を支援してくださった PhD supervisor の Ioannis Mitliagkas (Mila, UdeM) と Irina Rish に御礼申し上げます。