日本語

Journal articles

Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training

Hiroki Naganuma (equal contribution), Xinzhi Zhang (equal contribution), Man-Chung Yue, Ioannis Mitliagkas, Russell J. Hewett, Philipp Andre Witte, Yin Tat Lee

Equal contribution

Transactions on Machine Learning Research · August 2025

arXiv
2504.18454
OpenReview
8VTrvS5vN7

Abstract

Following AI scaling trends, frontier models continue to grow in size and continue to be trained on larger datasets. Training these models requires huge investments in exascale computational resources, which has in turn driven developtment of distributed deep learning methods. Data parallelism is an essential approach to speed up training, but it requires frequent global communication between workers, which can bottleneck training at the largest scales. In this work, we propose a method called Pseudo-Asynchronous Local SGD (PALSGD) to improve the efficiency of data-parallel training. PALSGD is an extension of Local SGD (Stich, 2018) and DiLoCo (Douillard et al., 2023), designed to further reduce communication frequency by introducing a pseudo-synchronization mechanism. PALSGD allows the use of longer synchronization intervals compared to standard Local SGD. Despite the reduced communication frequency, the pseudo-synchronization approach ensures that model consistency is maintained, leading to performance results comparable to those achieved with more frequent synchronization. Furthermore, we provide a theoretical analysis of PALSGD, establishing its convergence and deriving its convergence rate. This analysis offers insights into the algorithm's behavior and performance guarantees. We evaluated PALSGD on image classification and language modeling tasks. Our results show that PALSGD achieves better performance in less time compared to existing methods like Distributed Data Parallel (DDP), and DiLoCo. Notably, PALSGD trains 18.4% faster than DDP on ImageNet-1K with ResNet-50, 24.4% faster than DDP on TinyStories with GPT-Neo-125M, and 21.1% faster than DDP on TinyStories with GPT-Neo-8M.

Research summary

Problem
Frequent global communication can bottleneck data-parallel training at large scale.
Method
Pseudo-Asynchronous Local SGD extends Local SGD and DiLoCo with pseudo-synchronization so workers can use longer synchronization intervals while maintaining model consistency.
Result
The paper proves convergence and reports training-time improvements over DDP of 18.4% on ImageNet-1K/ResNet-50 and 24.4% and 21.1% on TinyStories with GPT-Neo-125M and GPT-Neo-8M.
Significance
It links communication-efficient distributed optimization with both theoretical guarantees and measured end-to-end speedups.
Limitations
The reported experiments cover image classification and small language-model workloads; behavior at frontier-model scale is not established by these results.

Research links

How to cite

Hiroki Naganuma, Xinzhi Zhang, Man-Chung Yue, Ioannis Mitliagkas, Russell J. Hewett, Philipp Andre Witte, Yin Tat Lee. “Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training.” Transactions on Machine Learning Research, 2025.

@article{Naganuma2025PseudoAsynchronous,
  title = {Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training},
  author = {Hiroki Naganuma and Xinzhi Zhang and Man-Chung Yue and Ioannis Mitliagkas and Russell J. Hewett and Philipp Andre Witte and Yin Tat Lee},
  year = {2025},
  journal = {Transactions on Machine Learning Research},
  eprint = {2504.18454},
  archivePrefix = {arXiv},
  url = {https://openreview.net/forum?id=8VTrvS5vN7}
}