Score
Designs, implements, and evaluates algorithms and system components that reduce the volume of data exchanged during collaborative or distributed model training and related distributed computations, including methods for compressing gradients, tokens, payloads, scan streams, and other transmitted data and for integrating hardware-aware encoding/decoding. Analyzes and trades off compression techniques such as quantization, sparsification, coding and protocol-level payload reduction against effects on accuracy, convergence, latency and resource use, and builds implementations in communication stacks or accelerators to validate efficiency.
To address the accumulation of compression errors caused by cross-layer propagation of activations and gradients in decentralized model-parallel training, this paper proposes the first forward/backward joint compression framework tailored for model parallelism. Our method introduces a predefined low-dimensional subspace based on the recursive structure of Transformers, enabling lossless reconstruction of activations and gradients with zero convergence degradation. We further design an inter-layer error correction mechanism and a model-parallelism-aware communication scheduling strategy. Experiments demonstrate that our approach achieves up to 99% communication compression ratio and a 100× improvement in communication efficiency. Notably, it successfully trains billion-parameter models on consumer-grade networks with only 80 Mbps bandwidth—matching the convergence performance attained in 100 Gbps data-center environments. This work bridges the gap between high-fidelity distributed training and resource-constrained edge or wide-area network settings.
To address the high-bandwidth interconnect dependency and substantial communication overhead in distributed training of large-scale models, this paper proposes a low-communication training method based on momentum frequency-domain sparsification. The core innovation lies in modeling optimizer momentum as a time-series signal, applying the discrete cosine transform (DCT) to separate its high- and low-frequency components, and synchronizing only the information-rich high-frequency part—thereby decoupling momentum updates from gradient synchronization. Combined with periodic model replica synchronization and Nesterov momentum compression, the method significantly reduces communication load. Extensive evaluations across Transformer and CNN architectures demonstrate up to 16× reduction in communication volume compared to the DiLoCo baseline, while preserving convergence stability and model accuracy—even under low-bandwidth conditions.
This work addresses the challenge of efficiently transmitting high-dimensional features from edge devices under stringent constraints on bandwidth, latency, and energy consumption. To this end, the authors propose a trainable, bit-wise soft quantization layer that approximates discrete step functions using multiple sigmoid functions, enabling end-to-end differentiability and task-oriented lossy compression. The method allows users to specify the desired bit-width and can be seamlessly integrated as a lightweight module at the data acquisition stage of neural networks, where it is jointly optimized with downstream tasks. Experimental results across multiple datasets demonstrate that the approach achieves 5–16× compression ratios (relative to 32-bit floating-point representations) using only 2–6 bits per feature, while maintaining accuracy nearly on par with full-precision models—significantly outperforming conventional quantization baselines.
Modern large-scale distributed training faces sharply diminishing returns in hardware scaling: as GPU counts reach thousands, communication overhead dominates performance bottlenecks, rendering conventional parallelism strategies—data, tensor, and pipeline parallelism—suboptimal. Method: Leveraging real-world LLM training workloads, this project establishes an empirical analytical framework spanning diverse model scales, hardware configurations, and parallelization strategies. It quantifies the nonlinear relationship between accelerator count and performance gain, precisely identifying critical inflection points across model, data, and compute scaling dimensions. Contribution/Results: We discover that low-communication “suboptimal” strategies become optimal at extreme scale; we empirically determine hardware selection criteria, cluster topology requirements, and optimal parallelism combinations for training billion-parameter models. Our findings provide actionable, deployment-ready optimization guidelines for trillion-parameter LLM training infrastructures.
To address efficient deployment of large language and vision models on resource-constrained devices, this work systematically investigates the non-orthogonality between sparsification and quantization—two key model compression techniques—and characterizes their joint accuracy degradation mechanism. We provide the first rigorous mathematical proof of their non-orthogonality, revealing that error accumulation is intrinsic and critically dependent on operational ordering: quantizing before sparsifying severely degrades accuracy, whereas sparsifying before quantizing mitigates error propagation. Grounded in theoretical analysis and validated across diverse models (OPT, LLaMA-125M–8B, ViT, ResNet) and hardware platforms, we establish “sparsify-then-quantize” as the optimal practice. This principle achieves high compression ratios while markedly improving the accuracy–efficiency trade-off, offering both theoretically grounded insights and a practical, empirically verified paradigm for edge deployment of foundation models.
General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.