AlgoPerf Workshop 2025
Links
Day1
Opening Remark / David Kanter / MLCommons
- ML Perf Training and Moore’s Laws

- How it started (algorithm)
- e.g. Shampoo paper is rejected at first (ICLR2021)
-
What is next?
- e.g. SOAP optimizer
-
success of MLPerf
- Caspr Adaptive / Combining Axes Preconditioners through Kronecker Approximation for Deep Learning
Intro Talk: The AlgoPerf Benchmark / Frank Schneider / University of Tübingen

- Shampoo 28% faster than compared to baseline
- Schedule free andam achieve 10% faster than baseline
- In ResNet workload, excpt for generalized adam, all the submiddion failed to beat the baseline
- Workload of ResNet is well studied, and well established benchmark, so it is hard to beat the baseline.
- No single algorithm is the best for all workloads.
QA
- How to make the benchmark?
- Pick several popular training algorithms with extensive hyperparameter tuning
Next TODO:
- SOAP, MUON, AdEMAMIX
- Pierre Ablin, Apple
Spotlight Talk I: External Tuning Track Winner / Scaling Beyond Diagonal Preconditioners for Training Neural Networks At-Scale / Anna Cai & Michael Shi / Meta AI
-
Two approximation strategies of shampoo
- Block diagonal
- kronecker factored approximation
-
Why kronecker factored approximation?
- upper bound or coarse kronecker rank-one approximation (1900s)
-
Learning Rate Grafting: Transferability of Optimizer Tuning / Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, Cyril Zhang (ICLR2022 Rejected) is a key to success of shampoo

- Why do we need grafting?
- correcting scaling across blocks
- corecting preconditioner staleness
Eigenvalue-corrected shampoo
- technical contribution
- enable warm-started QR algorithm to decrease precondition frequency significantly
Spotlight Talk II: Self-Tuning Track Winner Schedules & Schedule-Free Learning / Aaron Defazio / Meta AI
Two theory practice mismatch
-
theory: look nothing like the best-performing schedules used by experiments
-
we never use sgd in precise form as we analyze
- sgd with averaging
[1] Folk-law: sqrt t schedule is bad, a flat schedule is worse
-
from gradient norm observation, we can derive refined schedules for any particular task
-
Making the Last Iterate of SGD Information Theoretically Optimal

.%20When%20considering%20only%20worst-case%20analysis%2C%20our%20theory%20predicts%20that%20the%20optimal%20choice%20is%20the%20linear%20decay%20schedule%20where%20the%20step-size%20is%20set%20proportional%20to%201%20-%20t%2FT%2C%20where%20t%20is%20the%20current%20iteration%20and%20T%20is%20the%20total%20number%20of%20steps.%20To%20go%20beyond%20this%20worst-case%20analysis%2C%20we%20use%20the%20observed%20gradient%20norms%20to%20derive%20schedules%20refined%20for%20any%20particular%20task.%20These%20refined%20schedules%20exhibit%20learning%20rate%20warm-up%20and%20rapid%20learning%20rate%20annealing%20near%20the%20end%20of%20training.%20Ours%20is%20the%20first%20systematic%20approach%20to%20automatically%20yield%20both%20of%20these%20properties.%20We%20perform%20the%20most%20comprehensive%20evaluation%20of%20learning%20rate%20schedules%20to%20date%2C%20evaluating%20across%2010%20diverse%20deep%20learning%20problems%2C%20a%20series%20of%20LLMs%2C%20and%20a%20suite%20of%20logistic%20regression%20problems.%20We%20validate%20that%20overall%2C%20the%20linear-decay%20schedule%20outperforms%20all%20commonly%20used%20default%20schedules%20including%20cosine%20annealing.%20Our%20adaptive%20schedule%20refinement%20method%20gives%20further%20improvements.)
- why does not polyak averaging work well
Lightning Talks / AlgoPerf Submissions & their Follow-Ups
Niccolò Ajroldi: Weight Averaging Techniques on AlgoPerf
- large evaluation of weight averaging on algo perf
- speedup training
- works well, but there is diminishing returns
- cannot beat resnet, but can beat other workload
- shampoo + lawa works well
- drawback
- CPU-GPU communication is slow
- improve generalization
- replace lr schedule
- averaging: LAWA, EMA
- baseline: nadamw + linear warmup + cosine decay
David Tweedle: Applying Randomized Singular Value Decomposition During Training
Sourabh Medapati: Lessons from competing in AlgoPerf v0.5
- conduct extensive hyperparameter search
-
the configuration which beats the baseline of ResNet is behave like a momentum sgd optimizer
- nadam
- power = 1 : sign sgd
- power = 2 : adam
- make this power a hyperparameter for tuning
Roundtable Discussion
A moderated open audience discussion on “The Future of Training Algorithms”
- Aaron Defazio
- Jeremy Cohen
- current and previous neural network architechture is co-designed with the optimizer
- e.g.
- ResNet w/ SGD
- Transformer w/ Adam
- e.g.
-
composition of optimizer and model
- e.g. shampoo ignores the model’s inter-layer dependencies since it drop non-diagonal blocks for approximation
-
overlooked components
- label smoothing
- EMA
- metrics
- L1 norm of the gradient
- L2
- entropy
Day2
Invited Talk I / How does Gradient Descent Work? / Jeremy Cohen - Flatiron Institute
Invited Talk II / What is the best O(n) Hessian query? / Madeleine Udell - Stanford University
- PINN
Challenges in Training PINNs: A Loss Landscape Perspective
low rank aproximation of curvature
Invited Talk III / Stochastic-Gradient-based Algorithms for Nonconvex Constrained Optimization and Learning / Frank E. Curtis - Lehigh University
Roundtable Discussion / A moderated open audience discussion on “Training Algorithms in Production” / Michael Rabbat - Meta AI
- Michael Rabbat (Meta)
- Panel
- Rohan Anil (GDM -> Meta)
- Zachary Nado (GDM)
- Hao-Jun Michael Shi (Meta)
- Guna Lakshminarayanan (Meta -> LinkedIn)
- …..
Panel Discussion / The Future of AlgoPerf / George Dahl - Google DeepMind
- George Dahl (Google DeepMind)
- Panel
- Runa Eschenhagen (Meta, Cambridge)
- Priya Kasimbeg (Google DeepMind)
- Niccolò Ajroldi (MPI-IS)
- Michael Rabbat (Meta)
Topics:
- Modern workloads
Other Referece
Acknowlegements
- GDM team and Meta AI team for providing the opportunity to attend the workshop.