Score
Designs and analyzes algorithms and systems that compress high-dimensional vectors by projecting them into lower-dimensional subspaces (using sparse or random projections and related compressed representations) to enable efficient aggregation. Builds projection-aware aggregation routines and weighting schemes that compute client reliability or preserve distance-based robustness while reducing server computation and communication complexity.
This work addresses the trade-off between communication overhead and aggregation accuracy in distributed learning when transmitting sparse local models. It proposes the first compression scheme that integrates covering codes with sketching techniques and establishes a dual information-theoretic lower bound based on f-divergence. This bound is tight for binary alphabets and strictly stronger than conventional Fano-type bounds. The proposed method achieves a communication–accuracy trade-off approaching the theoretical limit in frequency estimation tasks. Although it does not attain the bound for general alphabets, it opens a new avenue for future research in this direction.
本文通过将多种聚类方法统一表达为受约束的低秩投影,建立了一个通用优化框架,并提供了关于这些方法稳定性和恢复保证的理论分析。
Existing approaches to large-scale matrix multiplication in elastic computing suffer from poor straggler tolerance, high data upload overhead, and excessive storage redundancy—requiring full local storage of input matrices. To address these challenges, this paper proposes a novel two-stage Lagrange coding scheme, the first to jointly optimize both storage and download efficiency in elastic environments. The method supports VM preemption, dynamic scaling, and heterogeneous node scheduling, reducing storage overhead to 1/L of Zhong et al.’s scheme and overcoming limitations of conventional full-storage or fault-intolerant designs. By integrating distributed matrix blocking with heterogeneous and cyclic task assignment strategies, our approach achieves load balancing, minimizes total computation time and computational redundancy, and significantly improves fault tolerance and resource utilization—as empirically validated on AWS EC2.
This work addresses the issue of uncontrolled reconstruction errors in SVD-based compression of large collections of matrices when heuristic grouping is employed prior to concatenation. To overcome this limitation, the authors propose a theory-driven compressive clustering framework grounded in spectral analysis of horizontally concatenated matrices. They establish, for the first time, a globally provable upper bound on SVD reconstruction error and derive two novel spectral bounds based on a lower bound for singular value growth. Building upon these theoretical guarantees, they design three clustering algorithms with explicit error control, integrated with incremental approximate SVD to efficiently estimate compression error without explicitly forming the full concatenated matrix. The resulting approach achieves a favorable balance among speed, accuracy, and scalability, significantly enhancing the reliability and practicality of SVD compression in applications such as multi-view learning, signal processing, and neural network compression.
This work addresses efficient decomposition of high-dimensional sparse tensors—common in healthcare and cybersecurity—on modern parallel processors, overcoming restrictive assumptions about mode structure or sparsity distribution inherent in conventional compressed formats. We propose ALTO, an adaptive linearization tensor representation that is agnostic to both mode structure and sparsity distribution. Built upon ALTO, we design a parallel decomposition algorithm featuring low synchronization overhead and high data reuse, augmented by dynamic performance modeling and scheduling heuristics for automatic hardware adaptation. Leveraging cache- and memory-aware optimizations on Intel Xeon Scalable platforms, experiments demonstrate that ALTO achieves over 10× speedup versus the best structure-agnostic format and a 5.1× geometric mean speedup versus the best structure-aware format, while incurring only 25% of the latter’s storage overhead.
This work addresses the high computational overhead of existing robust aggregation methods in federated learning under Byzantine attacks, which stems from processing high-dimensional gradients and hinders scalability to large models. The authors propose a Projection-based Dimensionality Reduction (PDR) framework that compresses client gradients into a low-dimensional subspace via sparse random projections, substantially reducing server-side aggregation complexity while preserving robustness. PDR is the first method to enable generic acceleration of distance-based robust aggregators at the vector level, approaching the theoretical lower bound on computation. It guarantees optimal convergence rates of $O(1/\sqrt{T})$ in non-convex settings and $O(1/T)$ under strong convexity. Experiments demonstrate that PDR achieves speedups of several orders of magnitude in aggregation time, introducing only controllable approximation error while maintaining both efficiency and convergence performance.
论文研究了大规模数据的欧几里得空间和$\ell_p$-范数表示,提出高效算法计算这些表示,并讨论了它们在数据压缩和洞察提取中的应用。
Clustering high-dimensional discrete data is often hindered by high computational cost, sensitivity to sparsity, and limited methodological applicability. This work proposes a deterministic dimensionality reduction framework that compresses binary, categorical, or count-based high-dimensional discrete data into low-dimensional continuous representations via weighted positional encoding. The resulting mapping is injective, preserving the discriminative structure of the original data; under mild conditions, the compressed variables approximately follow a Gaussian distribution while maintaining inter-cluster distances, thereby ensuring identifiable clustering structures. Empirical evaluations on real-world datasets—including infant names and microbiome profiles—demonstrate that the method achieves high clustering accuracy and substantially outperforms mainstream dimensionality reduction techniques in computational efficiency, offering both strong theoretical guarantees and practical utility.
This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.
This work addresses the activation communication bottleneck that limits pipeline-parallel training of large language models under low-bandwidth network conditions. The authors propose MAPL, a novel method that enables each pipeline stage to independently learn and dynamically optimize an orthogonal compression subspace. Task-adaptive low-rank compression is achieved through orthogonal projection learning constrained on the Stiefel manifold, while reconstruction fidelity is enhanced via factorized anchor embeddings combined with residual vector quantization. A streaming codebook synchronization protocol is further introduced to reduce communication overhead. Evaluated on LLaMA models ranging from 150M to 1B parameters, MAPL achieves negligible performance degradation even at high compression ratios, significantly outperforming existing approaches such as Subspace Networks.