sparse random projection

Designs and implements sparse randomized linear mappings that project high‑dimensional feature vectors into lower‑dimensional spaces using sparse projection matrices (SRP) and other randomized projection techniques; analyzes their mathematical properties, such as approximate distance and inner‑product preservation and related Johnson–Lindenstrauss bounds, as well as computational and storage benefits. Builds preprocessing pipelines and evaluates how these projections affect downstream model training efficiency, preservation of predictive signal, and transferability across datasets.

sparserandomprojection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Faster Linear Algebra Algorithms with Structured Random Matrices

Aug 28, 2025
CC
Chris Camaño
🏛️ California Institute of Technology

Structured random matrices lack a unified theoretical analysis and a general design framework in randomized linear algebra. Method: This paper introduces the “Oblivious Subspace Injection” (OSI) property, establishing the first decoupled abstract analytical framework that separates correctness proofs of algorithms from instantiation-specific verification. Contribution/Results: We prove that sparse random matrices, random triangular transforms, and tensor-product-structured matrices all satisfy OSI, thereby unifying their dimensionality-reduction fidelity guarantees for tasks such as low-rank approximation and least-squares regression. Leveraging this framework, we design accelerated algorithms with near-optimal time complexity. Empirical evaluation on synthetic datasets and scientific computing benchmarks confirms both efficiency and practical utility.

Analyzing algorithms using the Oblivious Subspace Injection (OSI) propertyDesigning faster randomized linear algebra algorithms with structured matricesIdentifying practical OSI examples for efficient low-rank approximation

Explicit Group Sparse Projection with Applications to Deep Learning and NMF

Dec 09, 2019
RO
Riyasat Ohib
🏛️ TReNDS Center | Georgia Institute of Technology | University of Mons | J.P. Morgan AI Research

This work addresses the challenge of explicitly controlling the average sparsity—measured by the Hoyer metric—in sparse projections of vector sets. We propose the first group-level explicit sparsity projection method: it directly specifies a target average sparsity level and jointly optimizes the sparsity patterns of all vectors, eliminating per-vector processing or reliance on implicit regularization parameters. Our approach generalizes the weighted ℓ₁ norm, enabling flexible sparsity modeling with linear-time computational complexity. The key innovation is the first formulation that imposes interpretable, tunable average sparsity constraints over an entire vector group. Experiments demonstrate substantial improvements in the accuracy–sparsity trade-off for ResNet50 pruning, and competitive reconstruction error and classification performance on CIFAR-10, ImageNet, and non-negative matrix factorization tasks.

Applies the method to deep learning pruning and nonnegative matrix factorization tasksDesigns a sparse projection method for vector groups with explicit average sparsity controlProposes a computationally efficient projection operator linear in problem size

This paper investigates the structural properties of the family ℱₘ,α of empirical distributions induced by all *m*-dimensional linear projections of high-dimensional Gaussian data, where *n*, *d* → ∞ and *n*/*d* → α. Using tools from random matrix theory, optimal transport, and information theory, we (i) precisely characterize the Wasserstein radius of ℱₘ,α—fully solving the *m* = 1 case—and derive tight upper and lower bounds on its KL divergence and Rényi information dimension; (ii) extend classical unsupervised projection analysis to supervised learning settings, establishing sharp Wasserstein radius bounds; and (iii) derive a tight upper bound on the interpolation threshold for two-layer neural networks with *m* hidden neurons. Collectively, these results provide a unified theoretical framework and fundamental benchmarks for high-dimensional projection statistics, representation learning, and the analysis of overparameterized models.

Characterize asymptotic distributions of Gaussian projections in high dimensionsDetermine Wasserstein radius bounds for low-dimensional projection setsEstablish interpolation threshold bounds for two-layer neural networks

Weighted least-squares approximation with determinantal point processes and generalized volume sampling

Dec 21, 2023
AN
Anthony Nouy
🏛️ Nantes Université | Centrale Nantes | CNRS UMR 6629

This work studies weighted least-squares function approximation in $L^2$ space based on random sampling: given an $m$-dimensional subspace $V_m$, how to achieve near-optimal $L^2$ approximation error with minimal sampling cost. We propose a generalized volume resampling framework that, for the first time, achieves expected near-optimal $L^2$ error—i.e., bounded by a constant multiple of the best approximation error—using only $O(m log m)$ samples. Furthermore, in embedding normed spaces, we establish almost-sure error control in the $H$-norm. Our method integrates projection determinantal point processes (DPPs), generalized volume sampling, and independent repeated DPP sampling, significantly enhancing sample diversity and feature selection efficiency. Numerical experiments demonstrate that our approach attains accuracy comparable to i.i.d. or classical volume sampling—but with substantially fewer samples.

Approximating L2 functions using m-dimensional space V_mPromoting feature diversity via DPP and volume samplingReducing sample count while maintaining error bounds

Sparse Linear Regression and Lattice Problems

Feb 22, 2024
AG
Aparna Gupte
🏛️ MIT

This work investigates the average-case computational complexity of sparse linear regression (SLR), focusing on whether polynomial-time algorithms exist for ill-conditioned design matrices—e.g., those with low rank or high correlation. The authors establish the first rigorous, instance-level reduction from classical worst-case lattice problems—specifically Bounded Distance Decoding (BDD)—to SLR. Their framework directly links the condition number of the lattice problem to the restricted eigenvalue condition of the SLR design matrix. This reduction holds in both identifiable and unidentifiable regimes. Leveraging worst-case-to-average-case hardness amplification, they prove that if BDD is hard in the worst case, then SLR remains computationally intractable on average for all polynomial-time algorithms. The result bridges a fundamental gap at the intersection of high-dimensional sparse statistics and computational complexity theory, providing the first evidence of average-case hardness for SLR under realistic design matrix conditions.

Average-case hardness of sparse linear regressionHardness in unidentifiable regime for SLRReduction from lattice problems to SLR

Latest Papers

What's happening recently
View more

This study addresses the theoretical foundations of randomized dimensionality reduction by establishing sharp spectral norm bounds for the product of sparse random matrices and low-dimensional subspace embeddings. Departing from conventional approaches based on trace methods or Gaussian comparison inequalities, the work introduces a novel entropy-based analysis of vector level sets to achieve a refined understanding of the spectral properties of sparse random embeddings. By integrating models of negatively associated random variables with subspace isometry theory, the paper proves that, under the conditions \(k \geq C r(\log\log r)^2\) and \(p \geq (\log k)/k\), the spectral norm satisfies \(\|\Pi U_V\| \leq C\sqrt{kp}\) with high probability. This result is further extended to a broader class of negatively associated random matrix ensembles.

dimension reductionlevel-set entropyrandom matrix

This work addresses the challenge of efficiently preserving continuous curve distances—such as the Fréchet distance—under dimensionality reduction for high-dimensional polygonal curves. The authors propose a randomized projection method based on sparse oblivious subspace embeddings that simultaneously approximates multiple curve dissimilarity measures, including Fréchet, q-DTW, and Hausdorff distances, within a relative error of (1±ε) using a target dimension of O(ε⁻² log(nm)). By constructing a unified framework for generalized curve distance metrics, the approach extends dimensionality reduction theory to piecewise linear surfaces, substantially simplifying existing analyses and broadening applicability across diverse curve comparison tasks.

curve dissimilaritydimension reductionFréchet distance

This work addresses the computational and memory bottlenecks of traditional Grassmannian kernel methods, which require constructing full Gram matrices and thus struggle with high-dimensional subspace data. To overcome these limitations, the authors propose a scalable kernel approximation framework based on random rank-one projections combined with bounded nonlinear transformations—either periodic or binary—that yield compact one-bit subspace feature representations. This approach enables continuous interpolation between the inverse Binet–Cauchy kernel and Gaussian-like kernels while effectively preserving the intrinsic geometry of subspaces. The method substantially reduces computational, memory, and storage costs. Experimental results on synthetic data and the ETH-80 classification benchmark demonstrate that the proposed technique accurately maintains Grassmannian geometric relationships with high fidelity, confirming its efficiency and practical utility.

Grassmannian kernelsrandom featuresrank-one projections

Hot Scholars

HZ

Heming Zou

Tsinghua University
Machine Learning
SS

Solmaz S. Kia

Professor, University of California Irvine
Control theoryDistributed algorithm design for cooperative networked systemsData fusion
YZ

Yunliang Zang

Brandeis University
Computational NeuroscienceBrain-inspired ComputingSystems Biology
TB

Thomas Bäck

Professor of Computer Science, Leiden University; Chief Scientist, NORCE Research Centre, Norway
Evolutionary ComputationEvolutionary AlgorithmsMachine LearningIndustry 4.0