Score
Designs and analyzes methods that recover the linear subspace(s) underlying a set of vectors or feature representations and produce estimates of their dimensionality (and optionally the subspace basis). Also derives and validates theoretical bounds or confidence statements on the true subspace dimension and on estimation error or robustness to noise/sparsity.
This paper addresses the lack of rigorous methodological guidance for intrinsic dimension estimation in high-dimensional data. We systematically survey and categorize mainstream algorithms grounded in local affine structure, parametric distributional assumptions, and topological invariance. Within a unified analytical framework and through extensive numerical experiments, we conduct the first comprehensive comparative evaluation of maximum likelihood estimation (MLE), PCA-based tangent space estimation, manifold neighborhood methods, and statistical fitting techniques across varying curvature, noise levels, and sample sizes. Results reveal that most methods exhibit high sensitivity to hyperparameters, suffer from overfitting, and experience sharp declines in accuracy, robustness, and generalizability in high dimensions. Our core contribution lies in rigorously characterizing the applicability boundaries of existing approaches, identifying nonlinear geometric structure and finite-sample effects as primary determinants of estimation reliability. This work provides both empirical evidence and theoretical guidance for principled algorithm selection and future methodological improvements in intrinsic dimension estimation.
This paper investigates the theoretical properties of Winsorized Principal Component Analysis (WPCA) for subspace recovery in high-dimensional data contaminated with outliers. Addressing the lack of rigorous characterization of subspace consistency and robustness in existing work, we establish the first asymptotic consistency theory for WPCA. We introduce a generalized breakdown point for subspace estimators and derive a tight lower bound, thereby revealing the inherent trade-off between estimation accuracy and robustness induced by Winsorization. By integrating data truncation, random matrix perturbation analysis, and projection distance metrics, we prove that the WPCA-estimated subspace converges almost surely to the true subspace as the sample size grows and the outlier proportion vanishes, while achieving an optimal-order perturbation bound. Both theoretical analysis and numerical experiments confirm that WPCA simultaneously attains strong robustness against outliers and high estimation accuracy.
Robust low-dimensional subspace recovery under highly contaminated data—where inliers constitute an extremely small fraction—is challenging for existing methods. Method: We propose the Subspace-Constrained Tyler Estimator (STE), which embeds the Tyler M-estimator into a subspace-constrained optimization framework and solves it via iterative reweighted least squares. Contributions: First, we establish the first convergence guarantee for STE. Second, under a weak inlier–outlier model, we prove its exact subspace recovery property. Third, we significantly lower the feasible inlier rate bound—enabling exact recovery even when inliers drop far below the failure threshold of classical Tyler estimation, i.e., in extreme sparsity regimes. Experiments on the generalized haystack model demonstrate that STE, initialized with the Tyler Mean Estimator (TME), successfully recovers subspaces with as little as 0.1% inlier fraction—substantially expanding the practical applicability of robust subspace learning.
This paper addresses the fundamental problem of determining the minimum number of linear measurements required for unique signal recovery from algebraic varieties. We propose a unified analytical framework grounded in algebraic geometry, which— for the first time—systematically leverages dimension theory and projection variety properties to characterize sampling lower bounds under algebraic priors, yielding necessary and sufficient conditions for unique reconstruction and tight bounds on the minimal measurement count. The framework unifies modeling of two canonical inverse problems: phase retrieval and low-rank matrix recovery, and rigorously verifies the tightness of the derived bounds in both settings. Our core innovation lies in deeply integrating algebraic geometric tools into the analysis of linear inverse problems, transcending traditional reliance on specific structural assumptions (e.g., sparsity or low-rankness). This provides a general theoretical foundation for optimal sampling design of signals admitting algebraic structure.
This paper addresses prediction risk control in high-dimensional linear regression under random design: how to achieve dimension-free, non-asymptotic upper bounds on prediction error without explicitly estimating the high-dimensional covariance matrix. To this end, we propose a novel “error-in-operator” paradigm, which implicitly incorporates the design covariance structure into the empirical risk minimization objective—bypassing explicit covariance estimation. Theoretically, we establish the first dimension-free, non-asymptotic upper bound on prediction error; rigorously prove that auxiliary variables do not inflate the effective dimension; and attain statistically optimal convergence rates. Computationally, the method eliminates dependence on covariance estimation, substantially reducing both statistical and computational complexity. Numerical experiments demonstrate its robustness and efficiency in high-dimensional sparse settings.
To address weak class discriminability and insufficient feature robustness in machine learning, this paper proposes a Multi-level Orthogonal Subspace (MOS) Karhunen–Loève feature theory within a random tensor space. Training data are modeled as stochastic processes in a Bochner space, and hierarchical KL expansions explicitly decouple dominant class structures from inter-class anomalous signals, enabling class-wise subspace disentanglement and interpretable projection features. This work establishes, for the first time, a MOS feature construction paradigm under the random tensor framework—uniquely integrating statistical modeling rigor with geometric interpretability. Evaluated on the ADNI plasma dataset, the method significantly outperforms gradient boosting, RUS Boost, random forests, and CNNs, achieving substantial gains in classification accuracy. These results validate its robust discriminative capability for high-noise biomedical data.
This work proposes a novel architecture based on adaptive feature fusion and dynamic inference to address the limited generalization of existing methods in complex scenarios. By incorporating a multi-scale context-aware module and a learnable routing strategy, the approach effectively integrates local details with global semantic information and dynamically adjusts its computational pathway during inference according to input content. Experimental results demonstrate that the model significantly outperforms state-of-the-art methods across multiple benchmark datasets while maintaining low computational overhead. The primary contribution lies in introducing the first dynamic feature fusion framework that jointly optimizes accuracy and efficiency, offering a new perspective for efficient visual understanding.
This work addresses the lack of theoretical characterization in existing kernel methods for machine learning regarding the residual structure and energy stability of multichannel signals in complex systems. The authors propose an analytical framework grounded in operator defect identities, introducing the novel concept of “telescopic energy residuals.” By integrating iterative products with a λₙ-relaxed Kaczmarz scheme, they establish admissibility conditions for residuals and derive prior energy bounds. For the first time, this framework incorporates operator defect theory into kernel methods and kernel principal component analysis (KPCA), rigorously proving explicit convergence of generalized algorithms, a residual energy decomposition theorem, and stability criteria under noise. The approach significantly extends infinite-dimensional Kaczmarz theory to broader applications in machine learning.
This paper addresses the efficient learning of minimum-volume confidence ellipsoids for arbitrary high-dimensional distributions: given i.i.d. samples and a confidence level α, find the ellipsoid E of minimal volume satisfying Pr_D[E] ≥ 1−α. The problem is NP-hard when the ellipsoid’s condition number β is unbounded; thus, we focus on the β-bounded regime. We propose the first polynomial-time algorithm achieving an O(β^{γd}) volume approximation ratio—where γ > 0 is a small constant—and prove that this exponential dependence on d is computationally nearly tight. Our approach integrates the primal-dual structure of the minimum-volume enclosing ellipsoid problem with geometric Brascamp–Lieb inequalities to formulate a robust optimization framework. This yields the first polynomial-time robust subspace recovery algorithm with worst-case theoretical guarantees. The method significantly enhances stability and practicality in high-dimensional anomaly detection and dimensionality reduction.
该研究解决了高维因子模型中主成分估计误差问题,通过将误差分解为可解释的两部分,并在金融经济、基因组学等领域应用。
本文针对高维数据的主成分分析问题,提出了一种基于子空间分析的新算法SuperPCA,通过利用样本协方差矩阵的部分特征空间来更高效准确地估计主要信号。