🤖 AI Summary
This paper addresses the “curse of dimensionality” and computational bottlenecks in intrinsic dimension (ID) estimation for high-dimensional nonlinear data. We propose an efficient, robust ID estimation algorithm that avoids both eigen-decomposition and neighborhood search. Our method constructs a matrix-vector multiplication framework via random projection and power iteration, integrated with gradient-sensitive local manifold curvature estimation—significantly reducing time and space complexity. Evaluated on diverse real-world and synthetic datasets, it achieves 3–12× speedup over state-of-the-art methods, reduces ID estimation error by 37%, and cuts memory consumption by 85%. To our knowledge, this is the first ID estimator relying solely on matrix-vector products, achieving high accuracy, low computational overhead, and strong scalability. The approach establishes a practical, scalable paradigm for large-scale nonlinear data analysis.
📝 Abstract
The real-life data have a complex and non-linear structure due to their nature. These non-linearities and the large number of features can usually cause problems such as the empty-space phenomenon and the well-known curse of dimensionality. Finding the nearly optimal representation of the dataset in a lower-dimensional space (i.e. dimensionality reduction) offers an applicable mechanism for improving the success of machine learning tasks. However, estimating the required data dimension for the nearly optimal representation (intrinsic dimension) can be very costly, particularly if one deals with big data. We propose a highly efficient and robust intrinsic dimension estimation approach that only relies on matrix-vector products for dimensionality reduction methods. An experimental study is also conducted to compare the performance of proposed method with state of the art approaches.