🤖 AI Summary
To address high communication overhead, weak privacy preservation, and low geometric fidelity in distributed large-scale linear regression, this paper proposes the first distributed hybrid sketching framework: local random projections are applied at each node, followed by a secondary sketching step at the central node to construct an efficient and robust $ell_2$ subspace embedding. The method jointly optimizes embedding dimension and computational time, overcoming the inherent trade-off limitations of conventional single-layer sketching. Under rigorous theoretical guarantees on embedding accuracy, it significantly reduces both communication cost and target dimensionality—experiments demonstrate reductions of 30%–50%—while strictly preserving the intrinsic geometric structure of the data. This work establishes a provably correct and scalable paradigm for distributed signal processing and machine learning in privacy-sensitive and resource-constrained environments.
📝 Abstract
Linear algebraic operations are ubiquitous in engineering applications, and arise often in a variety of fields including statistical signal processing and machine learning. With contemporary large datasets, to perform linear algebraic methods and regression tasks, it is necessary to resort to both distributed computations as well as data compression. In this paper, we study extit{distributed} $ell_2$-subspace embeddings, a common technique used to efficiently perform linear regression. In our setting, data is distributed across multiple computing nodes and a goal is to minimize communication between the nodes and the coordinator in the distributed centralized network, while maintaining the geometry of the dataset. Furthermore, there is also the concern of keeping the data private and secure from potential adversaries. In this work, we address these issues through randomized sketching, where the key idea is to apply distinct sketching matrices on the local datasets. A novelty of this work is that we also consider extit{hybrid sketching}, extit{i.e.} a second sketch is applied on the aggregated locally sketched datasets, for enhanced embedding results. One of the main takeaways of this work is that by hybrid sketching, we can interpolate between the trade-offs that arise in off-the-shelf sketching matrices. That is, we can obtain gains in terms of embedding dimension or multiplication time. Our embedding arguments are also justified numerically.