Score
Designs and implements streaming sketch algorithms and high-performance mixed-precision kernels (e.g., FP16 accumulation and mixed-precision sparse sketching) that compute compact linear projections of data. Builds and analyzes rounding schemes (deterministic, stochastic, dithered), accumulation and atomic-update strategies, and memory/compute optimizations to trade off numerical accuracy and throughput while measuring sketch/embedding quality.
This work investigates the implementation of Sparse Unaware Subspace Embeddings (SparseStack) under FP16 mixed-precision arithmetic on GPUs, where performance is typically constrained by memory bandwidth and atomic operations, and reduced precision may degrade embedding quality. The study systematically evaluates the impact of deterministic rounding, stochastic rounding, and dithered quantization on subspace distortion and least-squares accuracy. Empirical results demonstrate that all FP16 SparseStack variants preserve nearly identical embedding fidelity across coherent, incoherent, and adversarial problem instances, indicating that the sketching distribution—not the rounding strategy—primarily governs numerical accuracy. Among the approaches examined, deterministic rounding incurs the lowest computational overhead, establishing it as the optimal choice. These findings validate an efficient and practical pathway for deploying FP16 SparseStack in large-scale applications.
This work addresses the inefficiency of sparse sketching on GPUs, where irregular memory accesses caused by random sparsity severely limit bandwidth utilization and computational throughput. To overcome this, the authors propose a co-designed approach combining a novel BlockPerm-SJLT sparse structure with a customized FlashSketch CUDA kernel, yielding the first GPU-optimized sparse sketching method that preserves the theoretical guarantees of Oblivious Subspace Embedding. The design introduces tunable parameters to explicitly balance efficiency and accuracy. Experimental results demonstrate that the proposed method achieves a 1.7× geometric mean speedup over existing GPU sketching schemes on RandNLA benchmarks and GraSS data attribution tasks, while significantly advancing the Pareto frontier between speed and accuracy.
This work addresses the fundamental limitations of conventional linear sketching methods in data stream scenarios, where achieving a balance among reconstruction accuracy, computational efficiency, and real-time performance remains challenging due to information gaps caused by loss of orthogonal components. To overcome this, we propose FLORE—the first unsupervised generative sketching framework that incorporates a generative prior, enabling high-fidelity signal recovery without requiring ground-truth training data. By synergistically integrating generative modeling, linear sketching, and a lightweight recovery algorithm, FLORE achieves high-quality reconstruction while reducing recovery error by up to three orders of magnitude and accelerating computation by up to 100× compared to existing learning-based approaches.
This paper addresses approximate nearest neighbor (ANN) search and approximate kernel density estimation (A-KDE) over large-scale dynamic data streams. We propose a dynamic sketching framework based on locality-sensitive hashing (LSH), supporting both sliding-window and Turnstile stream models. Our method is the first to achieve sublinear space complexity for ANN in streaming settings while attaining near-optimal trade-offs between estimation error and query time. It also provides the first sublinear-space, theoretically guaranteed summary for A-KDE under sliding windows. By integrating batched query optimization and efficient window maintenance mechanisms, the algorithm ensures low-latency queries and high-throughput updates. Experiments on multiple real-world datasets demonstrate that our sketch achieves significantly higher accuracy and faster query times than state-of-the-art baselines, using substantially less memory—far below linear space—while providing rigorous theoretical guarantees.
Neural sketches suffer from poor generalization across data domains, inflexible memory allocation, and a lack of theoretical error guarantees. To address these issues, this paper proposes Lego Sketch—a modular, scalable neural sketch architecture—whose core innovation is the decoupling of memory-augmented neural networks (MANNs) into composable “memory bricks,” enabling dynamic adaptation to varying memory budgets and data stream distributions. We establish, for the first time, a provable upper bound on estimation error for neural sketches. Lego Sketch achieves unified high-accuracy frequency estimation across domains and memory constraints. Extensive experiments on diverse real-world data streams demonstrate that Lego Sketch significantly outperforms classical sketches (e.g., Count-Min, CMS) and state-of-the-art neural sketches: under identical memory budgets, it reduces average relative error by 32%–57%, thereby achieving superior trade-offs among time, space, and accuracy.
This work addresses the challenge of efficiently processing high-throughput matrix streams under stringent resource constraints, where existing approaches suffer from poor update efficiency due to frequent cubic-time matrix decompositions under tight error bounds. The authors propose AeroSketch, a framework that leverages randomized numerical linear algebra (RandNLA) to construct compact matrix sketches suitable for persistent, sliding-window, and distributed streaming settings. AeroSketch is the first method to simultaneously achieve optimal communication and space complexity while reducing the per-update time complexity from cubic to quadratic, thereby attaining near-optimal (within logarithmic factors) update performance. Experimental results on both synthetic and real-world datasets demonstrate that AeroSketch significantly improves throughput while maintaining comparable approximation accuracy and optimal resource consumption.
该论文提出了一种名为Stuffed IBLT的线性草图方法,用于精确恢复原始向量。此方法在保持信息理论最优空间使用的同时,支持高效的更新和解码操作,解决了多集合协调问题。
This work resolves the long-standing open question of whether efficient turnstile streaming algorithms for polynomial-length streams are essentially equivalent to linear sketching. By introducing tools from Fourier analysis and additive combinatorics, we establish this equivalence in the practical turnstile model for the first time: any turnstile algorithm using space $S$ can be simulated by a linear sketch requiring only $O(S)$ linear measurements, with total space $O(S \log S)$. Our approach abandons the traditional transition-graph machinery, enabling efficient reconstruction of the final vector via a linear sketch. This yields new lower bounds and, under a natural smoothness assumption, leads to a compact sketch with bounded entries and total space merely $O(S)$.
This work addresses the trade-off between memory usage and estimation accuracy in frequency estimation over data streams. Under a stochastic stationary stream model, the authors analyze the Elastic Sketch structure, which combines a heavy block for exact counting of high-frequency items and a light block based on Count-Min Sketch for aggregating the remaining traffic. They derive, for the first time, closed-form expressions for the limiting distribution of counters and the expected estimation error, revealing the theoretical structure of the optimal eviction threshold and substantially narrowing the parameter tuning space. Leveraging probabilistic modeling, asymptotic analysis, and hashing techniques, they formulate an efficiently computable error expression that enables near-optimal allocation of memory and threshold settings. The theoretical findings are validated empirically on Zipf-distributed streams.
This work addresses a key limitation in traditional linear sketching methods for streaming data statistics, which rely on the strong assumption that each hash bucket contains only a single key—thereby constraining space efficiency. To overcome this, the authors propose a novel approach that stores randomized linear combinations of multiple keys within each bucket and reconstructs key-value pairs during recovery by solving a sparse linear system. This design effectively relaxes the single-key-per-bucket constraint, achieving substantially improved space efficiency with only a modest increase in computational overhead. Experimental results demonstrate that the proposed method significantly reduces memory consumption while markedly enhancing space utilization for streaming data statistics.