design optimal embeddings

Designs and analyzes linear embedding and sketching operators (randomized or structured) and randomized low-rank projection procedures that preserve task-relevant structure while minimizing worst-case or statistical loss; proves optimality statements (e.g., minimax, rotation- or distribution-invariant) and prescribes classes of embeddings tailored to particular downstream algorithms.

designoptimalembeddings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the optimality of random embeddings in sketch-and-solve least squares and randomized SVD for low-rank approximation. By integrating tools from random matrix theory, minimax analysis, and rotational invariance principles, the authors establish that random orthogonal matrices achieve minimax optimality within the sketch-and-solve framework, while any rotationally invariant embedding is optimal for randomized SVD. Building on these theoretical insights, they derive the tightest error bounds to date and corroborate their findings through numerical experiments, demonstrating that a variety of commonly used embeddings closely approach the theoretically optimal performance in practice—thereby revealing a striking universality phenomenon across different embedding schemes.

error boundrandom embeddingrandomized algorithms

Sketching Low-Rank Plus Diagonal Matrices

Sep 27, 2025
AF
Andres Fernandez
🏛️ University of Tübingen | Vector Institute

This paper addresses the problem of jointly estimating the low-rank and diagonal components of a linear operator with low-rank-plus-diagonal (LoRD) structure, using only a small number of matrix-vector products. To this end, we propose SKETCHLORD—a novel method that formulates joint estimation as a scalable convex optimization problem, thereby avoiding error accumulation inherent in sequential estimation approaches. Leveraging randomized sketching and matrix decomposition principles, SKETCHLORD achieves structured approximation solely via black-box matrix-vector multiplication. Theoretically, it provides robustness guarantees under noise. Empirically, on both synthetic benchmarks and large-scale operators—including Hessian matrices from deep learning models—SKETCHLORD significantly outperforms existing methods: it achieves higher recovery accuracy and superior computational efficiency at the same query cost. By unifying estimation and optimization within a sketching-based framework, SKETCHLORD establishes a new paradigm for high-dimensional structured operator approximation.

Improves accuracy over sequential diagonal and low-rank approximationsProvides scalable convex optimization for large-scale linear operatorsSimultaneously estimates low-rank and diagonal matrix components

Norming Sets for Tensor and Polynomial Sketching

Jun 05, 2025
YZ
Yifan Zhang
🏛️ University of Texas at Austin

This work addresses the problem of randomized dimensionality reduction for real algebraic varieties and images of polynomial maps—such as low-rank tensors and tensor networks. We develop a unified sketching theory framework for such structured algebraic sets. Our method introduces the *median sketch*, the first approach achieving norm-preserving embedding with only $ ilde{O}(dim V)$ measurements—sharply improving upon the classical $ ilde{O}((dim V)^2)$ bound. The median sketch is compatible with diverse sketching operators, including sub-Gaussian, fast Johnson–Lindenstrauss (JL), and tensor-structured sketches. Leveraging generalized set theory, we provide a unified characterization of their embedding guarantees. Theoretical results ensure efficient compression and accelerated computation for high-dimensional tensors and polynomial-structured data. This yields a tight, scalable foundation for large-scale tensor learning, enabling provably accurate low-dimensional representations while preserving essential algebraic geometry properties.

Control sketching dimension for various sketch operatorsDevelop sketching theory for real algebraic varietiesPropose median sketch method with efficient measurements

Distributed Hybrid Sketching for $ell_2$-Embeddings

Dec 29, 2024
NC
Neophytos Charalambides
🏛️ University of California, San Diego

To address high communication overhead, weak privacy preservation, and low geometric fidelity in distributed large-scale linear regression, this paper proposes the first distributed hybrid sketching framework: local random projections are applied at each node, followed by a secondary sketching step at the central node to construct an efficient and robust $ell_2$ subspace embedding. The method jointly optimizes embedding dimension and computational time, overcoming the inherent trade-off limitations of conventional single-layer sketching. Under rigorous theoretical guarantees on embedding accuracy, it significantly reduces both communication cost and target dimensionality—experiments demonstrate reductions of 30%–50%—while strictly preserving the intrinsic geometric structure of the data. This work establishes a provably correct and scalable paradigm for distributed signal processing and machine learning in privacy-sensitive and resource-constrained environments.

Data integrityLarge-scale datasetsLinear regression

A Framework for Statistical Inference via Randomized Algorithms

Jul 20, 2023
ZZ
Zhixiang Zhang
🏛️ University of Macau | Columbia University | University of Pennsylvania

This paper addresses the challenge of uncertainty quantification arising from the inherent non-determinism of randomized algorithms—such as random projection and stochastic optimization—by proposing the first asymptotic statistical inference framework that requires no prior distributional assumptions. Methodologically, it establishes the first set of verifiable conditions for asymptotic normality of randomized outputs and introduces three novel strategies: sub-randomization, multi-run plug-in, and multi-run aggregation—integrating multi-run resampling, Polyak–Ruppert averaging, momentum SGD, and randomized sketching. The key contributions are: (i) enabling reliable confidence interval construction for high-dimensional and large-scale stochastic optimization and randomized least squares; and (ii) achieving negligible computational and communication overhead. Extensive simulations demonstrate robustness and practical efficacy on high-dimensional sparse and ultra-large-scale datasets.

Address computational challenges in large dataset analysisDevelop inference methods for non-deterministic algorithm resultsQuantify uncertainty in randomized algorithm outputs

Latest Papers

What's happening recently
View more

本文针对高维度线性bandits问题,提出了一种基于聚类的CS-LB算法,通过保留每个簇内的完整协方差信息并调整草图大小来减少计算成本,同时保证了较低的遗憾。

Computational EfficiencyHigh-dimensional SettingsLinear Bandits

Hot Scholars

YS

Yiyang Sun

Department of Mechanical and Aerospace Engineering, Syracuse University
unsteady aerodynamicsflow controlmodal analysisand reduced-order modeling
YS

Yulei Sui

University of New South Wales (UNSW Sydney)
Static Program AnalysisSecure Software EngineeringAI4SESE4AI
HP

Hammond Pearce

Senior Lecturer (a.k.a. Assistant Prof), UNSW School of Computer Science and Engineering
CybersecurityEmbedded SystemsHardware designLarge Language Models
GC

Grzegorz Chrupała

Associate Professor at Tilburg University
Computational LinguisticsSpoken Language ProcessingVision and LanguageInterpretability
CR

Cynthia Rudin

Professor of Computer Science, ECE, Statistics, and Biostatistics & Bioinformatics, Duke University
machine learninginterpretabilitydata science