Score
Designs and implements algorithms and scoring functions that assign continuous anomaly or outlier scores to individual data items using distance-based measures (neighborhoods, nearest neighbors, distance-to-centroid) and robust statistical estimators; builds and calibrates distance metrics, neighborhood models, robust scaling procedures, and detection thresholds. Analyzes and evaluates score reliability and discrimination, addressing issues such as contamination, masking, high-dimensional scaling, computational efficiency, and robustness to distributional assumptions.
High-dimensional anomaly detection suffers from the “curse of dimensionality,” rendering conventional methods ineffective; existing approaches often compromise interpretability or computational efficiency. This paper introduces a novel outlierness statistic based on “distance-to-distance,” leveraging the asymptotic concentration of pairwise distances and inner products in high-dimensional spaces—transforming the curse of dimensionality into a discriminative advantage that ensures asymptotic separability between anomalies and inliers. We theoretically establish the existence of a non-vanishing separation boundary for this statistic in high dimensions. Furthermore, we propose a distribution-free random rotation testing framework, requiring no parametric assumptions and exhibiting strong robustness. Experiments on synthetic data and diverse real-world high-dimensional datasets—including gene expression profiles and image features—demonstrate that our method significantly outperforms state-of-the-art baselines: it achieves substantially higher recall while maintaining low false positive rates, combining statistical rigor, computational feasibility, and result interpretability.
This work addresses the robustness of mean estimation in statistical learning under three concurrent challenges: adversarial data contamination, heavy-tailed distributions, and differential privacy constraints. Methodologically, it unifies robust statistics, high-dimensional geometry, stochastic optimization, and differential privacy theory to establish the first conceptual and algorithmic bridge across distinct robustness paradigms. Key technical abstractions—including iterative filtering, covariance trimming, and fractional gradient descent—are identified as common algorithmic primitives. The paper proposes a suite of computationally efficient estimators achieving statistically optimal convergence rates; each attains the information-theoretic lower bound under all three constraint classes simultaneously. By reconciling theoretical tightness with practical efficiency, this framework advances robust mean estimation from ad hoc heuristics toward a principled, unified design paradigm.
This paper addresses anomaly detection in structured data by proposing Preference-based Isolation Forest (PIF), a novel method that maps raw data into a preference-driven high-dimensional embedding space and constructs a PI-Forest tree structure for efficient anomaly scoring. Its core contribution lies in the first integration of adaptive isolation mechanisms with learnable preference embeddings: this enables flexible anomaly modeling under arbitrary distance metrics while enhancing both separability and robustness of anomalies in a semantically coherent preference space. Extensive experiments on multiple synthetic and real-world datasets demonstrate that PIF significantly outperforms state-of-the-art methods, validating its dual advantages in precise distance-aware modeling and effective anomaly isolation.
To address the limitations of existing scan statistics for sparse anomaly detection under multi-source heterogeneous coordinate systems—namely, their reliance on strong distributional assumptions and difficulty in calibration—this paper proposes a rank-based high-criticism scoring method. The approach leverages only the relative ordering among independent observations, avoiding parametric modeling entirely, and introduces a novel rank-based high-criticism framework to nonparametrically characterize detectability conditions. We theoretically establish that detection power is uniquely determined by the probability that an anomalous observation exceeds a typical one, and prove asymptotic optimality under both exponential families and convolution models. The method achieves strong robustness and theoretical interpretability: it successfully identifies process anomalies in pharmaceutical quality control data, and simulations demonstrate performance approaching that of the oracle test.
This work addresses the problem of k-means clustering in the presence of outliers by proposing a robust method based on k-nearest neighbor (KNN) distances. Given a budget to remove z outliers, the approach simply discards the z points with the largest KNN distances and then applies standard k-means clustering to the remaining data. Under a mild assumption on the minimum optimal cluster size, the method is the first to theoretically guarantee a constant-factor approximation—comparable to existing algorithms—without requiring additional cluster centers or excessive point removal, thereby establishing a formal connection between outlier detection and robust clustering. Empirical evaluations on multiple real-world datasets demonstrate that the proposed method is both efficient and practical, achieving clustering costs and running times that match or outperform several more complex baseline algorithms.
This study addresses the lack of a unified anomaly detection framework for mixed-type data comprising both continuous and ordinal categorical variables. The authors propose a robust approach based on a latent Gaussian variable model, wherein non-anomalous observations are assumed to follow a multivariate Gaussian distribution, and ordinal variables are represented through underlying latent Gaussian variables. Parameter estimation is performed using the Minimum Covariance Determinant (MCD) estimator, explicitly accounting for potential incompleteness in the observed ordinal information. Theoretical analysis demonstrates that the method effectively identifies extreme outliers and maintains robustness even under data contamination. Empirical evaluations on synthetic datasets show high detection rates coupled with low false alarm rates, and the method’s practical utility is further validated through application to real-world Airbnb listing data.
This work addresses the high computational cost of leave-one-out density refitting in unsupervised anomaly detection by proposing a density-based leave-one-out influence score. The score quantifies local perturbation by measuring the difference in density estimates at a fixed grid and bandwidth before and after removing a given observation. For the first time, we derive a closed-form update formula for this score under the linear bin frequency polygon (LBFP) density estimator, significantly improving computational efficiency while preserving interpretability of the underlying density. Theoretically, we establish an asymptotic order separation between normal and anomalous points under this score. Empirical results demonstrate that the proposed method matches or outperforms state-of-the-art baselines across various contamination models, with low computational overhead, and achieves strong detection performance on a 29-dimensional real-world credit card fraud dataset.
This work proposes a novel approach that integrates conditional anomaly detection with instance-based learning to identify context-dependent anomalous behaviors in clinical settings, such as inappropriate hospital admission decisions or unwarranted ordering of HPF4 tests for heparin-induced thrombocytopenia. By incorporating conditional attributes to constrain the scope of anomaly assessment and leveraging multiple distance metrics enhanced through metric learning, the method accurately pinpoints anomaly instances most sensitive to their contextual conditions. Experimental results on two real-world clinical tasks—admission decisions for community-acquired pneumonia and HPF4 test ordering—demonstrate that the proposed framework substantially improves both the performance and clinical interpretability of conditional anomaly detection.
This work addresses the challenge of effectively identifying anomalous patterns that depend on a subset of attributes when other attributes are given as conditions. To this end, we propose a metric learning approach tailored for conditional anomaly detection. The method adaptively learns a distance metric that explicitly models conditional dependencies among attributes and integrates instance-level neighborhood analysis to enable precise local anomaly discrimination. In contrast to existing instance-based approaches that rely on fixed distance metrics, our approach significantly improves detection accuracy and demonstrates enhanced capability in recognizing critical anomalous instances under dynamic conditional scenarios.