Fast, Space-Optimal Streaming Algorithms for Clustering and Subspace Embeddings

📅 2025-04-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates asymptotically optimal streaming algorithms for ((k,z))-clustering and subspace embedding. To overcome the fundamental bottlenecks of classical methods—where memory and update time scale with input size (n) and data range (Delta)—we propose the first (n)- and (Delta)-independent streaming framework, integrating core-set compression, hierarchical sampling, random projection, and adaptive reweighting, along with novel (p)-norm sensitivity analysis and dynamic maintenance mechanisms. Our theoretical contributions are: (i) for ((k,z))-clustering, space complexity improves to ( ilde{O}(dk / min{varepsilon^4, varepsilon^{z+2}})) and update time to (d cdot log k cdot mathrm{polylog}(log(nDelta))); (ii) for subspace embedding, we achieve (O(d)) update time and ( ilde{O}(d^2/varepsilon^2)) space—the first streaming algorithm matching the time and space efficiency of offline counterparts.

Technology Category

Machine Learning: ClusteringData Mining & Knowledge Management: Data Stream MiningSearch and Optimization: Distributed Search

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsWeb Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSystems and Infrastructure for Web, Mobile and WoT: Data management and stream processing for Web, mobile and wireless applications
📝 Abstract
We show that both clustering and subspace embeddings can be performed in the streaming model with the same asymptotic efficiency as in the central/offline setting. For $(k, z)$-clustering in the streaming model, we achieve a number of words of memory which is independent of the number $n$ of input points and the aspect ratio $Delta$, yielding an optimal bound of $ ilde{mathcal{O}}left(frac{dk}{min(varepsilon^4,varepsilon^{z+2})} ight)$ words for accuracy parameter $varepsilon$ on $d$-dimensional points. Additionally, we obtain amortized update time of $d,log(k)cdot ext{polylog}(log(nDelta))$, which is an exponential improvement over the previous $d, ext{poly}(k,log(nDelta))$. Our method also gives the fastest runtime for $(k,z)$-clustering even in the offline setting. For subspace embeddings in the streaming model, we achieve $mathcal{O}(d)$ update time and space-optimal constructions, using $ ilde{mathcal{O}}left(frac{d^2}{varepsilon^2} ight)$ words for $ple 2$ and $ ilde{mathcal{O}}left(frac{d^{p/2+1}}{varepsilon^2} ight)$ words for $p>2$, showing that streaming algorithms can match offline algorithms in both space and time complexity.
Problem

Research questions and friction points this paper is trying to address.

Optimize streaming clustering with minimal memory usage
Achieve space-optimal subspace embeddings in streaming
Match offline algorithm efficiency in streaming models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimal memory usage for streaming clustering
Exponential improvement in update time
Space-time optimal streaming subspace embeddings
🔎 Similar Papers