AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the communication overhead bottleneck in large-scale distributed training of large language models (LLMs) by proposing an adaptive synchronization framework. The method introduces a novel adaptive local update interval mechanism based on scalar statistics, which dynamically adjusts the synchronization frequency according to the training state. Additionally, an outer optimizer correction algorithm is designed to effectively compensate for biases introduced by varying update steps. By integrating momentum and learning rate correction techniques, the proposed framework preserves model convergence performance while reducing communication frequency by 27% compared to the DiLoCo baseline, thereby significantly improving communication efficiency in distributed training.
📝 Abstract
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Distributed Training
Communication Efficiency
Adaptive Synchronization
Local Update Methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Synchronization
Distributed LLM Training
Local Update Methods
Communication Efficiency
Optimizer Correction
💼 Related Jobs
No related jobs found.
P
Pengyu He
Department of Electrical Engineering, California Institute of Technology
Yan Zhang
Yan Zhang
Tsinghua University
Computer VisionMulti-Modal Learning
R
Ruien Li
Department of Computer Sciences, University of Wisconsin - Madison
Guangwen Yang
Guangwen Yang
Professor of Computer Science and Technology, Tsinghua University