Score
Designs, implements, and analyzes methods that estimate the intrinsic (manifold or fractal) dimensionality of a dataset or probability distribution, using approaches such as local-neighborhood statistics (e.g., twoNN), correlation/fractal dimensions, PCA-slope methods, and nonlinear embedding techniques (MDS, UMAP) to produce a numeric dimension estimate. Evaluates estimator behavior (bias, variance, robustness to noise and sampling, sample complexity), compares recovered dimension to analytical or empirical ground truth, and uses the estimates to inform dimensionality reduction, model capacity decisions, or geometric interpretation of representations.
This paper addresses the lack of rigorous methodological guidance for intrinsic dimension estimation in high-dimensional data. We systematically survey and categorize mainstream algorithms grounded in local affine structure, parametric distributional assumptions, and topological invariance. Within a unified analytical framework and through extensive numerical experiments, we conduct the first comprehensive comparative evaluation of maximum likelihood estimation (MLE), PCA-based tangent space estimation, manifold neighborhood methods, and statistical fitting techniques across varying curvature, noise levels, and sample sizes. Results reveal that most methods exhibit high sensitivity to hyperparameters, suffer from overfitting, and experience sharp declines in accuracy, robustness, and generalizability in high dimensions. Our core contribution lies in rigorously characterizing the applicability boundaries of existing approaches, identifying nonlinear geometric structure and finite-sample effects as primary determinants of estimation reliability. This work provides both empirical evidence and theoretical guidance for principled algorithm selection and future methodological improvements in intrinsic dimension estimation.
High-dimensional data often reside on low-dimensional manifolds, yet existing manifold dimension estimation algorithms lack systematic evaluation and reproducible benchmarks. Method: We conduct a comprehensive empirical assessment of eight representative methods across synthetic and real-world datasets, quantifying the effects of noise, curvature, and sample size on estimation accuracy. We introduce a dataset-aware hyperparameter tuning principle and establish a controlled-variable experimental framework integrating local linear embedding, nearest-neighbor statistics, and multiscale geometric analysis. Contribution/Results: Contrary to the “complexity implies superiority” assumption, simple methods—such as the nearest-neighbor distance ratio and PCA-based gradient estimation—consistently achieve higher accuracy and robustness across most scenarios. This work delivers the first open-source, fully reproducible benchmark for manifold dimension estimation, accompanied by practical guidelines. It provides both theoretical insight and empirical evidence to inform unsupervised learning and dimensionality reduction methodology selection.
This paper addresses the “curse of dimensionality” and computational bottlenecks in intrinsic dimension (ID) estimation for high-dimensional nonlinear data. We propose an efficient, robust ID estimation algorithm that avoids both eigen-decomposition and neighborhood search. Our method constructs a matrix-vector multiplication framework via random projection and power iteration, integrated with gradient-sensitive local manifold curvature estimation—significantly reducing time and space complexity. Evaluated on diverse real-world and synthetic datasets, it achieves 3–12× speedup over state-of-the-art methods, reduces ID estimation error by 37%, and cuts memory consumption by 85%. To our knowledge, this is the first ID estimator relying solely on matrix-vector products, achieving high accuracy, low computational overhead, and strong scalability. The approach establishes a practical, scalable paradigm for large-scale nonlinear data analysis.
Estimating intrinsic dimensionality (ID) from real-world data is highly sensitive to neighborhood scale: small scales overestimate ID due to noise, while large scales introduce bias from manifold curvature and topology. This work proposes a self-consistent scale selection protocol that identifies the optimal “sweet spot” for ID estimation by enforcing local density constancy. Our key contribution is the first formal coupling of ID estimation and scale selection, resolved via iterative optimization that yields a theoretically guaranteed robust decoupling—effectively suppressing both noise and curvature effects. The method integrates local neighborhood graph construction, asymptotic statistical analysis, and rigorous error-bound derivation. Evaluated on diverse synthetic and real-world datasets, it reduces ID estimation error by over 30% compared to state-of-the-art methods, while significantly improving stability and noise robustness.
This paper addresses the limited accuracy of intrinsic dimension estimation on manifolds caused by neglecting local curvature and geometric structure. To overcome this, we propose a unified framework integrating principal component analysis (PCA) with regression modeling. Methodologically, we introduce two novel estimators—quadratic embedding and total least squares—that explicitly encode local curvature and graph-topological structure, thereby mitigating bias inherent in conventional PCA-based approaches, especially in highly curved regions. Theoretically grounded and empirically validated, our framework exhibits both robustness and interpretability. Experiments on synthetic and real-world datasets demonstrate that it matches or surpasses state-of-the-art methods in estimation accuracy, with particularly pronounced improvements in strongly curved manifold regions.
Existing hierarchical dimensionality reduction methods struggle to simultaneously preserve both local and global structures across multiple granularity levels while maintaining users’ mental map continuity. To address this, we propose HUMAP—the first hierarchical dimensionality reduction framework built upon UMAP. HUMAP achieves dual optimization of structural fidelity and mental map consistency through three core techniques: cross-scale neighborhood graph propagation, multi-granularity similarity modeling, and incremental embedding alignment. Compared to state-of-the-art methods, HUMAP significantly improves hierarchical structure preservation across multiple benchmark datasets. Furthermore, it demonstrates strong interpretability and interactive utility in real-world data labeling tasks. By unifying geometric structure preservation with cognitive consistency, HUMAP establishes a novel paradigm for multi-granularity visual analytics, enabling scalable, intuitive, and semantically grounded exploration of high-dimensional hierarchical data.
Traditional linear dimensionality reduction methods often fail to effectively uncover the intrinsic low-dimensional manifold structure embedded in high-dimensional data. This work systematically traces the historical development of manifold fitting and, for the first time, categorizes it into three distinct phases: nonparametric statistics, mathematically inspired analysis, and modern practical statistics. It clarifies manifold fitting’s role as an independent geometric data analysis tool and delineates its conceptual boundaries from related techniques such as manifold embedding and denoising. By integrating nonparametric methods, differential geometry, and contemporary statistical learning approaches, the paper explores cutting-edge applications of manifold fitting in neural networks and bioinformatics, offering a comprehensive reference framework that elucidates both its theoretical limits and practical utility.
This work addresses the high variance in local intrinsic dimensionality (LID) estimation caused by data sparsity in small neighborhoods, which undermines estimation accuracy. To mitigate this issue, the study introduces subbagging—a variant of bootstrap aggregation—into LID estimation for the first time, constructing an ensemble method that substantially reduces estimation variance while preserving the local distribution of nearest-neighbor distances. Through a systematic analysis of the interplay among sampling rate, neighborhood size (k), and ensemble scale, combined with a neighborhood smoothing strategy, the proposed approach effectively enhances estimation precision. Experimental results demonstrate that the method consistently achieves significantly lower mean squared error across a broad range of hyperparameters, while enabling controlled bias adjustment, thereby outperforming existing state-of-the-art techniques in overall performance.
This study systematically evaluates the effectiveness of UMAP and its supervised variants in leveraging label information for dimensionality reduction in both regression and classification tasks, presenting the first in-depth analysis of supervised UMAP in regression settings. Through comprehensive comparisons across synthetic and real-world datasets—including UMAP, supervised UMAP, PCA, Kernel PCA, Sliced Inverse Regression (SIR), Kernel SIR, and t-SNE—the quality of low-dimensional embeddings is assessed by their predictive performance. The findings reveal that while supervised UMAP excels in classification tasks, it struggles to effectively incorporate response variable information in regression scenarios, leading to substantially degraded performance. This limitation underscores a critical methodological gap in current supervised UMAP formulations for regression and provides clear guidance for future algorithmic improvements.
Estimating the intrinsic dimensionality of high-dimensional data is crucial for machine learning and computer vision, yet existing methods often fail due to reliance on specific geometric or distributional assumptions. This work proposes a nonparametric estimator based on nearest-neighbor distance ratios that requires no prior assumptions about the underlying data manifold or distribution. For the first time, it is theoretically proven that this estimator consistently converges to the true intrinsic dimension under arbitrary data distributions. Extensive experiments demonstrate that the method achieves state-of-the-art performance on both synthetic manifolds and real-world datasets, exhibiting high accuracy, strong robustness, and broad applicability.
This work addresses the limitations of traditional Euclidean dimensionality reduction methods in effectively handling data intrinsically residing on nonlinear Riemannian manifolds—such as hyperspheres or the manifold of symmetric positive-definite matrices. By extending classical techniques like principal component analysis and discriminant analysis into a Riemannian geometric framework, the study proposes geometry-aware nonlinear dimensionality reduction approaches grounded in geodesic distances, tangent space mappings, and intrinsic statistical measures. These include Principal Geodesic Analysis (PGA) and manifold-based discriminant analysis. Experimental results demonstrate that the proposed methods significantly outperform their Euclidean counterparts on benchmark datasets embedded in curved spaces, achieving superior preservation of intrinsic manifold structure, enhanced quality of low-dimensional embeddings, and improved downstream classification performance.