estimate intrinsic dimension

Designs, implements, and analyzes methods that estimate the intrinsic (manifold or fractal) dimensionality of a dataset or probability distribution, using approaches such as local-neighborhood statistics (e.g., twoNN), correlation/fractal dimensions, PCA-slope methods, and nonlinear embedding techniques (MDS, UMAP) to produce a numeric dimension estimate. Evaluates estimator behavior (bias, variance, robustness to noise and sampling, sample complexity), compares recovered dimension to analytical or empirical ground truth, and uses the estimates to inform dimensionality reduction, model capacity decisions, or geometric interpretation of representations.

estimateintrinsicdimension

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Manifold Dimension Estimation: An Empirical Study

Sep 18, 2025
ZB
Zelong Bi
🏛️ University of New South Wales

High-dimensional data often reside on low-dimensional manifolds, yet existing manifold dimension estimation algorithms lack systematic evaluation and reproducible benchmarks. Method: We conduct a comprehensive empirical assessment of eight representative methods across synthetic and real-world datasets, quantifying the effects of noise, curvature, and sample size on estimation accuracy. We introduce a dataset-aware hyperparameter tuning principle and establish a controlled-variable experimental framework integrating local linear embedding, nearest-neighbor statistics, and multiscale geometric analysis. Contribution/Results: Contrary to the “complexity implies superiority” assumption, simple methods—such as the nearest-neighbor distance ratio and PCA-based gradient estimation—consistently achieve higher accuracy and robustness across most scenarios. This work delivers the first open-source, fully reproducible benchmark for manifold dimension estimation, accompanied by practical guidelines. It provides both theoretical insight and empirical evidence to inform unsupervised learning and dimensionality reduction methodology selection.

Compares estimators on synthetic and real-world datasetsEmpirically studies manifold dimension estimation methodsEvaluates performance under noise, curvature, and sample size variations

A Novel Approach for Intrinsic Dimension Estimation

Mar 12, 2025
KO
Kadir Ozccoban
🏛️ Middle East Technical University | Kadir Has University

This paper addresses the “curse of dimensionality” and computational bottlenecks in intrinsic dimension (ID) estimation for high-dimensional nonlinear data. We propose an efficient, robust ID estimation algorithm that avoids both eigen-decomposition and neighborhood search. Our method constructs a matrix-vector multiplication framework via random projection and power iteration, integrated with gradient-sensitive local manifold curvature estimation—significantly reducing time and space complexity. Evaluated on diverse real-world and synthetic datasets, it achieves 3–12× speedup over state-of-the-art methods, reduces ID estimation error by 37%, and cuts memory consumption by 85%. To our knowledge, this is the first ID estimator relying solely on matrix-vector products, achieving high accuracy, low computational overhead, and strong scalability. The approach establishes a practical, scalable paradigm for large-scale nonlinear data analysis.

Addresses challenges of non-linear data structures.Estimates intrinsic dimension for dimensionality reduction.Improves efficiency in big data dimensionality reduction.

Beyond the noise: intrinsic dimension estimation with optimal neighbourhood identification

May 24, 2024
AD
Antonio Di Noia
🏛️ ETH Zurich | SISSA Trieste | Banca d’Italia | University of Insubria

Estimating intrinsic dimensionality (ID) from real-world data is highly sensitive to neighborhood scale: small scales overestimate ID due to noise, while large scales introduce bias from manifold curvature and topology. This work proposes a self-consistent scale selection protocol that identifies the optimal “sweet spot” for ID estimation by enforcing local density constancy. Our key contribution is the first formal coupling of ID estimation and scale selection, resolved via iterative optimization that yields a theoretically guaranteed robust decoupling—effectively suppressing both noise and curvature effects. The method integrates local neighborhood graph construction, asymptotic statistical analysis, and rigorous error-bound derivation. Evaluated on diverse synthetic and real-world datasets, it reduces ID estimation error by over 30% compared to state-of-the-art methods, while significantly improving stability and noise robustness.

Addressing scale-dependent ID variations due to noise and curvatureDeveloping automatic protocol for meaningful scale selection in datasetsEstimating intrinsic dimension by identifying optimal neighborhood scales

This paper addresses the limited accuracy of intrinsic dimension estimation on manifolds caused by neglecting local curvature and geometric structure. To overcome this, we propose a unified framework integrating principal component analysis (PCA) with regression modeling. Methodologically, we introduce two novel estimators—quadratic embedding and total least squares—that explicitly encode local curvature and graph-topological structure, thereby mitigating bias inherent in conventional PCA-based approaches, especially in highly curved regions. Theoretically grounded and empirically validated, our framework exhibits both robustness and interpretability. Experiments on synthetic and real-world datasets demonstrate that it matches or surpasses state-of-the-art methods in estimation accuracy, with particularly pronounced improvements in strongly curved manifold regions.

Developing regression-based dimension estimators that outperform existing methodsEstimating intrinsic manifold dimension using local graph structureImproving PCA by accounting for manifold curvature effects

HUMAP: Hierarchical Uniform Manifold Approximation and Projection

Jun 14, 2021
WE
Wilson E. Marc'ilio-Jr
🏛️ São Paulo State University (UNESP) | Nomic AI | Eindhoven University of Technology | Linnaeus University

Existing hierarchical dimensionality reduction methods struggle to simultaneously preserve both local and global structures across multiple granularity levels while maintaining users’ mental map continuity. To address this, we propose HUMAP—the first hierarchical dimensionality reduction framework built upon UMAP. HUMAP achieves dual optimization of structural fidelity and mental map consistency through three core techniques: cross-scale neighborhood graph propagation, multi-granularity similarity modeling, and incremental embedding alignment. Compared to state-of-the-art methods, HUMAP significantly improves hierarchical structure preservation across multiple benchmark datasets. Furthermore, it demonstrates strong interpretability and interactive utility in real-world data labeling tasks. By unifying geometric structure preservation with cognitive consistency, HUMAP establishes a novel paradigm for multi-granularity visual analytics, enabling scalable, intuitive, and semantically grounded exploration of high-dimensional hierarchical data.

Develops hierarchical dimensionality reduction for multi-granularity data analysisImproves mental map consistency compared to existing hierarchical approachesPreserves both local and global structures during hierarchical exploration

Latest Papers

What's happening recently
View more

Traditional linear dimensionality reduction methods often fail to effectively uncover the intrinsic low-dimensional manifold structure embedded in high-dimensional data. This work systematically traces the historical development of manifold fitting and, for the first time, categorizes it into three distinct phases: nonparametric statistics, mathematically inspired analysis, and modern practical statistics. It clarifies manifold fitting’s role as an independent geometric data analysis tool and delineates its conceptual boundaries from related techniques such as manifold embedding and denoising. By integrating nonparametric methods, differential geometry, and contemporary statistical learning approaches, the paper explores cutting-edge applications of manifold fitting in neural networks and bioinformatics, offering a comprehensive reference framework that elucidates both its theoretical limits and practical utility.

dimension reductionhigh-dimensional datalatent geometric structure

This work addresses the high variance in local intrinsic dimensionality (LID) estimation caused by data sparsity in small neighborhoods, which undermines estimation accuracy. To mitigate this issue, the study introduces subbagging—a variant of bootstrap aggregation—into LID estimation for the first time, constructing an ensemble method that substantially reduces estimation variance while preserving the local distribution of nearest-neighbor distances. Through a systematic analysis of the interplay among sampling rate, neighborhood size (k), and ensemble scale, combined with a neighborhood smoothing strategy, the proposed approach effectively enhances estimation precision. Experimental results demonstrate that the method consistently achieves significantly lower mean squared error across a broad range of hyperparameters, while enabling controlled bias adjustment, thereby outperforming existing state-of-the-art techniques in overall performance.

bagginghyper-parameter selectionLocal Intrinsic Dimensionality

This study systematically evaluates the effectiveness of UMAP and its supervised variants in leveraging label information for dimensionality reduction in both regression and classification tasks, presenting the first in-depth analysis of supervised UMAP in regression settings. Through comprehensive comparisons across synthetic and real-world datasets—including UMAP, supervised UMAP, PCA, Kernel PCA, Sliced Inverse Regression (SIR), Kernel SIR, and t-SNE—the quality of low-dimensional embeddings is assessed by their predictive performance. The findings reveal that while supervised UMAP excels in classification tasks, it struggles to effectively incorporate response variable information in regression scenarios, leading to substantially degraded performance. This limitation underscores a critical methodological gap in current supervised UMAP formulations for regression and provides clear guidance for future algorithmic improvements.

dimensionality reductionmanifold learningregression

Estimating the intrinsic dimensionality of high-dimensional data is crucial for machine learning and computer vision, yet existing methods often fail due to reliance on specific geometric or distributional assumptions. This work proposes a nonparametric estimator based on nearest-neighbor distance ratios that requires no prior assumptions about the underlying data manifold or distribution. For the first time, it is theoretically proven that this estimator consistently converges to the true intrinsic dimension under arbitrary data distributions. Extensive experiments demonstrate that the method achieves state-of-the-art performance on both synthetic manifolds and real-world datasets, exhibiting high accuracy, strong robustness, and broad applicability.

computer visiondimensionality estimationintrinsic dimensionality

This work addresses the limitations of traditional Euclidean dimensionality reduction methods in effectively handling data intrinsically residing on nonlinear Riemannian manifolds—such as hyperspheres or the manifold of symmetric positive-definite matrices. By extending classical techniques like principal component analysis and discriminant analysis into a Riemannian geometric framework, the study proposes geometry-aware nonlinear dimensionality reduction approaches grounded in geodesic distances, tangent space mappings, and intrinsic statistical measures. These include Principal Geodesic Analysis (PGA) and manifold-based discriminant analysis. Experimental results demonstrate that the proposed methods significantly outperform their Euclidean counterparts on benchmark datasets embedded in curved spaces, achieving superior preservation of intrinsic manifold structure, enhanced quality of low-dimensional embeddings, and improved downstream classification performance.

Dimensionality ReductionGeometric Data AnalysisManifold Structure

Hot Scholars

IM

Iuri Macocco

PostDoc, UPF
unsupervised learningmanifold learningLLM
AM

Antonietta Mira

Professore di Statistica, Università della Svizzera italiana, Lugano
computational statisticsMarkov chain Monte Carlo methods
MB

Marco Baroni

ICREA Professor (Barcelona)
Computational LinguisticsLinguisticsAICognitive Science
EC

Emily Cheng

Universitat Pompeu Fabra
computational linguisticsmachine learningNLP