Score
Designs and analyzes linear embedding and sketching operators (randomized or structured) and randomized low-rank projection procedures that preserve task-relevant structure while minimizing worst-case or statistical loss; proves optimality statements (e.g., minimax, rotation- or distribution-invariant) and prescribes classes of embeddings tailored to particular downstream algorithms.
This study investigates the optimality of random embeddings in sketch-and-solve least squares and randomized SVD for low-rank approximation. By integrating tools from random matrix theory, minimax analysis, and rotational invariance principles, the authors establish that random orthogonal matrices achieve minimax optimality within the sketch-and-solve framework, while any rotationally invariant embedding is optimal for randomized SVD. Building on these theoretical insights, they derive the tightest error bounds to date and corroborate their findings through numerical experiments, demonstrating that a variety of commonly used embeddings closely approach the theoretically optimal performance in practice—thereby revealing a striking universality phenomenon across different embedding schemes.
This paper addresses the problem of jointly estimating the low-rank and diagonal components of a linear operator with low-rank-plus-diagonal (LoRD) structure, using only a small number of matrix-vector products. To this end, we propose SKETCHLORD—a novel method that formulates joint estimation as a scalable convex optimization problem, thereby avoiding error accumulation inherent in sequential estimation approaches. Leveraging randomized sketching and matrix decomposition principles, SKETCHLORD achieves structured approximation solely via black-box matrix-vector multiplication. Theoretically, it provides robustness guarantees under noise. Empirically, on both synthetic benchmarks and large-scale operators—including Hessian matrices from deep learning models—SKETCHLORD significantly outperforms existing methods: it achieves higher recovery accuracy and superior computational efficiency at the same query cost. By unifying estimation and optimization within a sketching-based framework, SKETCHLORD establishes a new paradigm for high-dimensional structured operator approximation.
This work addresses the problem of randomized dimensionality reduction for real algebraic varieties and images of polynomial maps—such as low-rank tensors and tensor networks. We develop a unified sketching theory framework for such structured algebraic sets. Our method introduces the *median sketch*, the first approach achieving norm-preserving embedding with only $ ilde{O}(dim V)$ measurements—sharply improving upon the classical $ ilde{O}((dim V)^2)$ bound. The median sketch is compatible with diverse sketching operators, including sub-Gaussian, fast Johnson–Lindenstrauss (JL), and tensor-structured sketches. Leveraging generalized set theory, we provide a unified characterization of their embedding guarantees. Theoretical results ensure efficient compression and accelerated computation for high-dimensional tensors and polynomial-structured data. This yields a tight, scalable foundation for large-scale tensor learning, enabling provably accurate low-dimensional representations while preserving essential algebraic geometry properties.
To address high communication overhead, weak privacy preservation, and low geometric fidelity in distributed large-scale linear regression, this paper proposes the first distributed hybrid sketching framework: local random projections are applied at each node, followed by a secondary sketching step at the central node to construct an efficient and robust $ell_2$ subspace embedding. The method jointly optimizes embedding dimension and computational time, overcoming the inherent trade-off limitations of conventional single-layer sketching. Under rigorous theoretical guarantees on embedding accuracy, it significantly reduces both communication cost and target dimensionality—experiments demonstrate reductions of 30%–50%—while strictly preserving the intrinsic geometric structure of the data. This work establishes a provably correct and scalable paradigm for distributed signal processing and machine learning in privacy-sensitive and resource-constrained environments.
This paper addresses the challenge of uncertainty quantification arising from the inherent non-determinism of randomized algorithms—such as random projection and stochastic optimization—by proposing the first asymptotic statistical inference framework that requires no prior distributional assumptions. Methodologically, it establishes the first set of verifiable conditions for asymptotic normality of randomized outputs and introduces three novel strategies: sub-randomization, multi-run plug-in, and multi-run aggregation—integrating multi-run resampling, Polyak–Ruppert averaging, momentum SGD, and randomized sketching. The key contributions are: (i) enabling reliable confidence interval construction for high-dimensional and large-scale stochastic optimization and randomized least squares; and (ii) achieving negligible computational and communication overhead. Extensive simulations demonstrate robustness and practical efficacy on high-dimensional sparse and ultra-large-scale datasets.
本文解决了$\ell_2$回归中坐标精度的问题,提出了一种结合随机Hadamard展平、随机排列和高斯池化的新方法,保证了在$O(\epsilon^{-2}d\log d)$行数下获得$\ell_\infty$误差界。
本文解决了Khatri-Rao结构矩阵在子空间嵌入问题中的性能分析,通过证明仅需m = Õ(k/ε^2)维度即可达到(1±ε)误差的嵌入效果,改善了之前的研究结果。
本文针对高维度线性bandits问题,提出了一种基于聚类的CS-LB算法,通过保留每个簇内的完整协方差信息并调整草图大小来减少计算成本,同时保证了较低的遗憾。
本文探讨了随机投影在保持几何结构方面的局限性,通过分析Johnson-Lindenstrauss引理,并采用线性草图与奇异值分解方法来恢复距离特征。
本文解决了大规模矩阵迹估计问题,提出了一种基于递归TensorSketch的方法,减少了所需的随机比特数,并保证了估计的无偏性和方差的多项式增长。