Collaborative Compressors in Distributed Mean Estimation with Limited Communication Budget

πŸ“… 2026-01-26
πŸ›οΈ Trans. Mach. Learn. Res.
πŸ“ˆ Citations: 1
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of reducing communication overhead while maintaining estimation accuracy in distributed high-dimensional mean estimation under communication constraints, particularly when the correlation among node vectors is unknown. The authors propose four collaborative compression schemes that, for the first time, adaptively exploit inter-node vector similarity for efficient quantization and encoding without requiring prior knowledge of similarity. These methods achieve low computational complexity and substantial communication savings, with theoretical guarantees on estimation error under β„“β‚‚, β„“_∞, and cosine distance metrics. Crucially, the error bounds systematically decrease as vector similarity increases, demonstrating the algorithms’ effective utilization of structural similarity among node data.

Technology Category

Constraint Satisfaction and Optimization: Distributed CSP/OptimizationSearch and Optimization: Distributed SearchData Mining & Knowledge Management: Data Compression

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsSystems and Infrastructure for Web, Mobile and WoT: Experiences and lessons learnt from Web-based algorithms and system deployments
πŸ“ Abstract
Distributed high dimensional mean estimation is a common aggregation routine used often in distributed optimization methods. Most of these applications call for a communication-constrained setting where vectors, whose mean is to be estimated, have to be compressed before sharing. One could independently encode and decode these to achieve compression, but that overlooks the fact that these vectors are often close to each other. To exploit these similarities, recently Suresh et al., 2022, Jhunjhunwala et al., 2021, Jiang et al, 2023, proposed multiple correlation-aware compression schemes. However, in most cases, the correlations have to be known for these schemes to work. Moreover, a theoretical analysis of graceful degradation of these correlation-aware compression schemes with increasing dissimilarity is limited to only the $\ell_2$-error in the literature. In this paper, we propose four different collaborative compression schemes that agnostically exploit the similarities among vectors in a distributed setting. Our schemes are all simple to implement and computationally efficient, while resulting in big savings in communication. The analysis of our proposed schemes show how the $\ell_2$, $\ell_\infty$ and cosine estimation error varies with the degree of similarity among vectors.
Problem

Research questions and friction points this paper is trying to address.

distributed mean estimation
communication budget
correlation-aware compression
vector similarity
limited communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

collaborative compression
distributed mean estimation
communication-constrained learning
similarity-aware encoding
error analysis
πŸ”Ž Similar Papers