communication-efficient training

Designs, implements, and evaluates algorithms and system components that reduce the volume of data exchanged during collaborative or distributed model training and related distributed computations, including methods for compressing gradients, tokens, payloads, scan streams, and other transmitted data and for integrating hardware-aware encoding/decoding. Analyzes and trades off compression techniques such as quantization, sparsification, coding and protocol-level payload reduction against effects on accuracy, convergence, latency and resource use, and builds implementations in communication stacks or accelerators to validate efficiency.

communication-efficienttraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the accumulation of compression errors caused by cross-layer propagation of activations and gradients in decentralized model-parallel training, this paper proposes the first forward/backward joint compression framework tailored for model parallelism. Our method introduces a predefined low-dimensional subspace based on the recursive structure of Transformers, enabling lossless reconstruction of activations and gradients with zero convergence degradation. We further design an inter-layer error correction mechanism and a model-parallelism-aware communication scheduling strategy. Experiments demonstrate that our approach achieves up to 99% communication compression ratio and a 100× improvement in communication efficiency. Notably, it successfully trains billion-parameter models on consumer-grade networks with only 80 Mbps bandwidth—matching the convergence performance attained in 100 Gbps data-center environments. This work bridges the gap between high-fidelity distributed training and resource-constrained edge or wide-area network settings.

Addressing communication bottlenecks in decentralized model-parallel trainingCompressing activations and gradients without convergence degradationEnabling efficient billion-parameter training on low-bandwidth networks

Distributed Low-Communication Training with Decoupled Momentum Optimization

Oct 03, 2025
SN
Sasho Nedelkoski
🏛️ CAMPUS.TU-BERLIN.DE | LOGSIGHT.AI | TU-BERLIN.DE

To address the high-bandwidth interconnect dependency and substantial communication overhead in distributed training of large-scale models, this paper proposes a low-communication training method based on momentum frequency-domain sparsification. The core innovation lies in modeling optimizer momentum as a time-series signal, applying the discrete cosine transform (DCT) to separate its high- and low-frequency components, and synchronizing only the information-rich high-frequency part—thereby decoupling momentum updates from gradient synchronization. Combined with periodic model replica synchronization and Nesterov momentum compression, the method significantly reduces communication load. Extensive evaluations across Transformer and CNN architectures demonstrate up to 16× reduction in communication volume compared to the DiLoCo baseline, while preserving convergence stability and model accuracy—even under low-bandwidth conditions.

Enabling large model training with low-bandwidth interconnectsOptimizing momentum synchronization across distributed compute nodesReducing communication overhead in distributed model training

This work addresses the challenge of efficiently transmitting high-dimensional features from edge devices under stringent constraints on bandwidth, latency, and energy consumption. To this end, the authors propose a trainable, bit-wise soft quantization layer that approximates discrete step functions using multiple sigmoid functions, enabling end-to-end differentiability and task-oriented lossy compression. The method allows users to specify the desired bit-width and can be seamlessly integrated as a lightweight module at the data acquisition stage of neural networks, where it is jointly optimized with downstream tasks. Experimental results across multiple datasets demonstrate that the approach achieves 5–16× compression ratios (relative to 32-bit floating-point representations) using only 2–6 bits per feature, while maintaining accuracy nearly on par with full-precision models—significantly outperforming conventional quantization baselines.

bandwidth constraintedge computingfeature compression

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Nov 20, 2024
JF
Jared Fernandez
🏛️ Meta | Carnegie Mellon University

Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.

Assessing diminishing returns in scaling accelerators for large model trainingEvaluating parallelization strategies to minimize distributed communication overheadOptimizing hardware configuration for efficient large-scale model training

Effective Interplay between Sparsity and Quantization: From Theory to Practice

May 31, 2024
SB
Simla Burcu Harma
🏛️ EPFL | MangoBoost Inc. | Google | Korea University | Google DeepMind

To address efficient deployment of large language and vision models on resource-constrained devices, this work systematically investigates the non-orthogonality between sparsification and quantization—two key model compression techniques—and characterizes their joint accuracy degradation mechanism. We provide the first rigorous mathematical proof of their non-orthogonality, revealing that error accumulation is intrinsic and critically dependent on operational ordering: quantizing before sparsifying severely degrades accuracy, whereas sparsifying before quantizing mitigates error propagation. Grounded in theoretical analysis and validated across diverse models (OPT, LLaMA-125M–8B, ViT, ResNet) and hardware platforms, we establish “sparsify-then-quantize” as the optimal practice. This principle achieves high compression ratios while markedly improving the accuracy–efficiency trade-off, offering both theoretically grounded insights and a practical, empirically verified paradigm for edge deployment of foundation models.

Model CompressionResource-constrained DevicesSparse Quantization

Latest Papers

What's happening recently
View more

General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.

application-specific compressorslossless compressionmaintainability

Hot Scholars

WL

Weisi Lin

President's Chair Professor in Computer Science, CCDS, Nanyang Technological Unversity
Perception-inspired signal modelingperceptual multimedia quality evaluationvideo compressionimage processing & analysis
SC

Symeon Chatzinotas

Full Professor | IEEE Fellow | SIGCOM Head, SnT, University of Luxembourg
Wireless CommunicationsNon-Terrestrial NetworksInternet of Things6G
LL

Ligang Liu

University of Science and Technology of China
Computer GraphicsGeometry Processing3D Printing
YP

Yuval Pinter

Ben-Gurion University of the Negev
Natural Language ProcessingMachine LearningInformation RetrievalLinguistics
FC

Franck Cappello

Argonne National Laboratory, IEEE Fellow
Parallel ProcessingParallel ComputingHigh Performance ComputingFault Tolerance