Score
Designs, implements, and evaluates models and algorithms that discover structure in unlabeled data, including clustering, density estimation, dimensionality reduction, manifold learning, generative modeling, and representation learning. Develops training objectives, architectures, and validation metrics to learn, visualize, sample from, or detect anomalies in data without labeled supervision.
This paper addresses the challenge of evaluating and selecting appropriate clustering algorithms—K-means, DBSCAN, and spectral clustering—for high-dimensional data. We propose a unified unsupervised evaluation framework that integrates multiple dimensionality reduction techniques (PCA, t-SNE, UMAP) with composite validity metrics (silhouette coefficient, Adjusted Rand Index, Normalized Mutual Information). For the first time, we systematically uncover synergistic interactions between dimensionality reduction methods and clustering algorithms: UMAP preprocessing significantly enhances spectral clustering performance on complex manifold-structured data (e.g., MNIST, Fashion-MNIST, UCI HAR); K-means maintains superior computational efficiency; and DBSCAN excels at detecting irregularly shaped clusters. The framework provides reproducible, interpretable, and empirically grounded guidance for algorithm selection in high-dimensional clustering tasks.
In unsupervised anomaly detection on high-dimensional data, dimensionality reduction often obscures the geometric distribution of anomalies, hindering discriminative capability. Method: We propose a “manifold-in/manifold-out” binary anomaly classification framework. Innovatively adopting a manifold-aware paradigm, we characterize the intrinsic geometric nature of anomalies on the learned low-dimensional manifold and design a manifold-aligned multi-method fusion strategy that synergistically integrates models such as Isolation Forest. Contribution/Results: Our approach maintains precision while significantly improving recall—achieving a 16% gain over the best-performing single model (Isolation Forest) on MNIST. Extensive experiments demonstrate strong generalization and robustness in real-world high-dimensional scenarios. By explicitly leveraging manifold geometry, our work establishes a novel geometric-structure-driven paradigm for anomaly detection.
Traditional linear dimensionality reduction methods often fail to effectively uncover the intrinsic low-dimensional manifold structure embedded in high-dimensional data. This work systematically traces the historical development of manifold fitting and, for the first time, categorizes it into three distinct phases: nonparametric statistics, mathematically inspired analysis, and modern practical statistics. It clarifies manifold fitting’s role as an independent geometric data analysis tool and delineates its conceptual boundaries from related techniques such as manifold embedding and denoising. By integrating nonparametric methods, differential geometry, and contemporary statistical learning approaches, the paper explores cutting-edge applications of manifold fitting in neural networks and bioinformatics, offering a comprehensive reference framework that elucidates both its theoretical limits and practical utility.
This work addresses the lack of a universal statistical interpretation for the manifold hypothesis—that high-dimensional data approximately reside on low-dimensional manifolds. We propose the Latent Metric Model (LMM), a generative framework grounded in fundamental statistical concepts: latent variables, variable dependence, and stationarity—providing the first unified statistical justification for the manifold assumption. Methodologically, LMM integrates neighborhood graph construction, spectral analysis, and an interpretable inference framework to enable unsupervised manifold discovery and geometric structure recovery under weak priors. Experiments demonstrate that complex manifold geometries naturally emerge from minimal statistical mechanisms; LMM significantly reduces reliance on hand-crafted priors on both synthetic and real-world datasets, while enabling interpretable reconstruction of manifold dimensionality, curvature, and coordinate systems.
This work addresses anomaly detection on low-dimensional manifold-structured data. We propose an efficient isolation-based method that embeds data into a high-dimensional semantic-enhanced preference space and employs Locality-Sensitive Hashing (LSH) to accelerate sparse neighborhood estimation, thereby identifying the most isolated samples as anomalies. To our knowledge, this is the first approach to integrate LSH into a preference-space isolation framework, achieving both theoretical soundness and computational efficiency. Extensive experiments on multiple benchmark datasets demonstrate state-of-the-art detection performance, with inference speed improved by 3–5× over existing methods, alongside substantial reductions in time and memory overhead. The source code is publicly available.
Existing anomaly detection methods typically assume that normal data occupy a non-zero volume in the ambient space, overlooking their intrinsic geometric structure as lying on a low-dimensional manifold, which limits performance. This work proposes a novel manifold projection paradigm: learning a projection operator that maps inputs onto the manifold of normal samples and using the projection residual as the anomaly criterion. By avoiding explicit modeling of the degenerate data distribution, the approach prevents misclassifying rare yet normal instances and provides a unified explanation for both the effectiveness and failure modes of reconstruction-based methods. Extensive experiments demonstrate that the proposed framework significantly outperforms conventional boundary-learning approaches and achieves state-of-the-art results across multiple benchmarks compared to existing reconstruction-based models.
Social science research urgently requires interpretable and reproducible exploratory discovery from unstructured text—without presupposing measurement constructs. This paper proposes an end-to-end framework addressing this challenge. First, it constructs a high-dimensional, semantically transparent, and interpretable concept dictionary via sparse coding and semantic modeling. Second, it introduces a novel high-dimensional multiple testing procedure that rigorously controls the k-familywise error rate (k-FWER) under arbitrary variable dependence, substantially reducing researcher degrees of freedom. Third, it integrates machine learning interpretability techniques with selective inference to ensure statistical validity. The method is empirically validated in economic analyses—both causal and descriptive—and is accompanied by an open-source Jupyter toolkit, enabling low-cost, fully reproducible empirical workflows.
While contemporary self-supervised and masked/denoising autoencoder methods effectively learn strong representations from massive unlabeled data, their representational nature, cross-task generalization capability, and emergence mechanisms remain theoretically unexplained. Method: This project integrates statistical inference and nonconvex optimization theory to establish a unified analytical framework for unsupervised representation learning. Contribution/Results: It provides the first mathematical characterization of how self-supervised objectives—such as contrastive learning and reconstruction losses—induce structured latent spaces, and quantitatively links representation linear separability, invariance, and downstream generalization. The work identifies key theoretical conditions under which pretrained models achieve zero-shot transfer and task emergence in vision foundation models. Crucially, it delivers the first theoretical foundation for large-scale pretraining that is both statistically interpretable and optimization-traceable—bridging statistical guarantees with practical training dynamics.
Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.
This work addresses the challenge that existing unsupervised methods struggle to reliably detect subtle and noisy anomalies in complex time series, often being misled by noise in normal samples and missing near-normal anomalies. To overcome this limitation, we propose a novel unsupervised anomaly detection framework that integrates active learning: it enhances temporal dependency modeling through a masked time series reconstruction feedback mechanism and employs a minimax optimization strategy to differentially treat normal and anomalous samples, thereby improving robustness against noise and weak anomalies. Extensive experiments across four multivariate time series datasets and seven backbone models demonstrate that our method achieves an average AUC improvement of 12.39%, significantly outperforming current unsupervised approaches.