subspace dimension estimation

Designs and analyzes methods that recover the linear subspace(s) underlying a set of vectors or feature representations and produce estimates of their dimensionality (and optionally the subspace basis). Also derives and validates theoretical bounds or confidence statements on the true subspace dimension and on estimation error or robustness to noise/sparsity.

subspacedimensionestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Subspace Recovery in Winsorized PCA: Insights into Accuracy and Robustness

Feb 23, 2025
SH
Sangil Han
🏛️ Seoul National University

This paper investigates the theoretical properties of Winsorized Principal Component Analysis (WPCA) for subspace recovery in high-dimensional data contaminated with outliers. Addressing the lack of rigorous characterization of subspace consistency and robustness in existing work, we establish the first asymptotic consistency theory for WPCA. We introduce a generalized breakdown point for subspace estimators and derive a tight lower bound, thereby revealing the inherent trade-off between estimation accuracy and robustness induced by Winsorization. By integrating data truncation, random matrix perturbation analysis, and projection distance metrics, we prove that the WPCA-estimated subspace converges almost surely to the true subspace as the sample size grows and the outlier proportion vanishes, while achieving an optimal-order perturbation bound. Both theoretical analysis and numerical experiments confirm that WPCA simultaneously attains strong robustness against outliers and high estimation accuracy.

Perturbation bounds for contaminated dataRobustness to outliers analysisSubspace recovery in Winsorized PCA

Theoretical Guarantees for the Subspace-Constrained Tyler's Estimator

Mar 27, 2024
GL
Gilad Lerman
🏛️ University of Minnesota | University of Central Florida

Robust low-dimensional subspace recovery under highly contaminated data—where inliers constitute an extremely small fraction—is challenging for existing methods. Method: We propose the Subspace-Constrained Tyler Estimator (STE), which embeds the Tyler M-estimator into a subspace-constrained optimization framework and solves it via iterative reweighted least squares. Contributions: First, we establish the first convergence guarantee for STE. Second, under a weak inlier–outlier model, we prove its exact subspace recovery property. Third, we significantly lower the feasible inlier rate bound—enabling exact recovery even when inliers drop far below the failure threshold of classical Tyler estimation, i.e., in extreme sparsity regimes. Experiments on the generalized haystack model demonstrate that STE, initialized with the Tyler Mean Estimator (TME), successfully recovers subspaces with as little as 0.1% inlier fraction—substantially expanding the practical applicability of robust subspace learning.

Establish recovery guarantees for subspace-constrained Tyler's estimatorHandle inlier fractions below typical robust recovery thresholdsRecover low-dimensional subspace from outlier-corrupted data

This paper addresses the fundamental problem of determining the minimum number of linear measurements required for unique signal recovery from algebraic varieties. We propose a unified analytical framework grounded in algebraic geometry, which— for the first time—systematically leverages dimension theory and projection variety properties to characterize sampling lower bounds under algebraic priors, yielding necessary and sufficient conditions for unique reconstruction and tight bounds on the minimal measurement count. The framework unifies modeling of two canonical inverse problems: phase retrieval and low-rank matrix recovery, and rigorously verifies the tightness of the derived bounds in both settings. Our core innovation lies in deeply integrating algebraic geometric tools into the analysis of linear inverse problems, transcending traditional reliance on specific structural assumptions (e.g., sparsity or low-rankness). This provides a general theoretical foundation for optimal sampling design of signals admitting algebraic structure.

Algebraic geometry tools applied to phase retrievalLow-rank matrix recovery using linear samplesMinimum measurements for unique signal recovery on varieties

Dimension-free bounds in high-dimensional linear regression via error-in-operator approach

Feb 21, 2025
FN
Fedor Noskov
🏛️ HSE University | Humboldt University | WIAS Berlin

This paper addresses prediction risk control in high-dimensional linear regression under random design: how to achieve dimension-free, non-asymptotic upper bounds on prediction error without explicitly estimating the high-dimensional covariance matrix. To this end, we propose a novel “error-in-operator” paradigm, which implicitly incorporates the design covariance structure into the empirical risk minimization objective—bypassing explicit covariance estimation. Theoretically, we establish the first dimension-free, non-asymptotic upper bound on prediction error; rigorously prove that auxiliary variables do not inflate the effective dimension; and attain statistically optimal convergence rates. Computationally, the method eliminates dependence on covariance estimation, substantially reducing both statistical and computational complexity. Numerical experiments demonstrate its robustness and efficiency in high-dimensional sparse settings.

Dimension-free risk bounds derivationError-in-operator empirical risk minimizationHigh-dimensional linear regression analysis

Stochastic tensor space feature theory with applications to robust machine learning

Oct 04, 2021
JC
J. Castrillón-Candás
🏛️ Boston University

To address weak class discriminability and insufficient feature robustness in machine learning, this paper proposes a Multi-level Orthogonal Subspace (MOS) Karhunen–Loève feature theory within a random tensor space. Training data are modeled as stochastic processes in a Bochner space, and hierarchical KL expansions explicitly decouple dominant class structures from inter-class anomalous signals, enabling class-wise subspace disentanglement and interpretable projection features. This work establishes, for the first time, a MOS feature construction paradigm under the random tensor framework—uniquely integrating statistical modeling rigor with geometric interpretability. Evaluated on the ADNI plasma dataset, the method significantly outperforms gradient boosting, RUS Boost, random forests, and CNNs, achieving substantial gains in classification accuracy. These results validate its robust discriminative capability for high-noise biomedical data.

Constructs Multilevel Orthogonal Subspace for anomaly detection in data.Develops robust machine learning features using stochastic tensor spaces.Improves classification accuracy on Alzheimer's Disease dataset significantly.

Latest Papers

What's happening recently
View more

This work proposes a novel architecture based on adaptive feature fusion and dynamic inference to address the limited generalization of existing methods in complex scenarios. By incorporating a multi-scale context-aware module and a learnable routing strategy, the approach effectively integrates local details with global semantic information and dynamically adjusts its computational pathway during inference according to input content. Experimental results demonstrate that the model significantly outperforms state-of-the-art methods across multiple benchmark datasets while maintaining low computational overhead. The primary contribution lies in introducing the first dynamic feature fusion framework that jointly optimizes accuracy and efficiency, offering a new perspective for efficient visual understanding.

anti-concentrationcomputational tractabilitylow-degree polynomial method

This work addresses the lack of theoretical characterization in existing kernel methods for machine learning regarding the residual structure and energy stability of multichannel signals in complex systems. The authors propose an analytical framework grounded in operator defect identities, introducing the novel concept of “telescopic energy residuals.” By integrating iterative products with a λₙ-relaxed Kaczmarz scheme, they establish admissibility conditions for residuals and derive prior energy bounds. For the first time, this framework incorporates operator defect theory into kernel methods and kernel principal component analysis (KPCA), rigorously proving explicit convergence of generalized algorithms, a residual energy decomposition theorem, and stability criteria under noise. The approach significantly extends infinite-dimensional Kaczmarz theory to broader applications in machine learning.

kernel methodsoperator defect identitiesresidual analysis

Learning Confidence Ellipsoids and Applications to Robust Subspace Recovery

Dec 18, 2025
CG
Chao Gao
🏛️ University of Chicago | Toyota Technological Institute at Chicago | Northwestern University

This paper addresses the efficient learning of minimum-volume confidence ellipsoids for arbitrary high-dimensional distributions: given i.i.d. samples and a confidence level α, find the ellipsoid E of minimal volume satisfying Pr_D[E] ≥ 1−α. The problem is NP-hard when the ellipsoid’s condition number β is unbounded; thus, we focus on the β-bounded regime. We propose the first polynomial-time algorithm achieving an O(β^{γd}) volume approximation ratio—where γ > 0 is a small constant—and prove that this exponential dependence on d is computationally nearly tight. Our approach integrates the primal-dual structure of the minimum-volume enclosing ellipsoid problem with geometric Brascamp–Lieb inequalities to formulate a robust optimization framework. This yields the first polynomial-time robust subspace recovery algorithm with worst-case theoretical guarantees. The method significantly enhances stability and practicality in high-dimensional anomaly detection and dimensionality reduction.

Address NP-hardness in high dimensions for arbitrary distributionsEfficiently find confidence ellipsoids with volume approximation guaranteesProvide polynomial-time algorithm for robust subspace recovery problem

Hot Scholars

WL

Weilin Li

City University of New York
Applied harmonic analysisapproximation theorysignal processingdata science
JL

Jingyang Li

PhD Student, National University of Singapore
optimizationdeep learning
ID

Ilias Diakonikolas

University of Wisconsin-Madison
theoretical computer sciencealgorithmic statisticsmachine learningprobability theory
GI

Giannis Iakovidis

University of Wisconsin-Madison
Theoretical Computer ScienceMachine LearningOptimization
HJ

Han-Jia Ye

Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning