anomaly scoring methods (distance-based and robust outlier scoring)

Designs and implements algorithms and scoring functions that assign continuous anomaly or outlier scores to individual data items using distance-based measures (neighborhoods, nearest neighbors, distance-to-centroid) and robust statistical estimators; builds and calibrates distance metrics, neighborhood models, robust scaling procedures, and detection thresholds. Analyzes and evaluates score reliability and discrimination, addressing issues such as contamination, masking, high-dimensional scaling, computational efficiency, and robustness to distributional assumptions.

anomalyscoringmethods(distance-based

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.71
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$184K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

DOD: Detection of outliers in high dimensional data with distance of distances

Nov 04, 2025
SL
Seong-ho Lee
🏛️ University of Seoul | Yonsei University

High-dimensional anomaly detection suffers from the “curse of dimensionality,” rendering conventional methods ineffective; existing approaches often compromise interpretability or computational efficiency. This paper introduces a novel outlierness statistic based on “distance-to-distance,” leveraging the asymptotic concentration of pairwise distances and inner products in high-dimensional spaces—transforming the curse of dimensionality into a discriminative advantage that ensures asymptotic separability between anomalies and inliers. We theoretically establish the existence of a non-vanishing separation boundary for this statistic in high dimensions. Furthermore, we propose a distribution-free random rotation testing framework, requiring no parametric assumptions and exhibiting strong robustness. Experiments on synthetic data and diverse real-world high-dimensional datasets—including gene expression profiles and image features—demonstrate that our method significantly outperforms state-of-the-art baselines: it achieves substantially higher recall while maintaining low false positive rates, combining statistical rigor, computational feasibility, and result interpretability.

Detecting outliers in high-dimensional data using geometric distance relationshipsDeveloping computationally efficient outlier detection with theoretical separation guaranteesOvercoming traditional methods' breakdown in high-dimensional asymptotic settings

The Broader Landscape of Robustness in Algorithmic Statistics

Dec 03, 2024
GK
Gautam Kamath
🏛️ University of Waterloo | Vector Institute

This work addresses the robustness of mean estimation in statistical learning under three concurrent challenges: adversarial data contamination, heavy-tailed distributions, and differential privacy constraints. Methodologically, it unifies robust statistics, high-dimensional geometry, stochastic optimization, and differential privacy theory to establish the first conceptual and algorithmic bridge across distinct robustness paradigms. Key technical abstractions—including iterative filtering, covariance trimming, and fractional gradient descent—are identified as common algorithmic primitives. The paper proposes a suite of computationally efficient estimators achieving statistically optimal convergence rates; each attains the information-theoretic lower bound under all three constraint classes simultaneously. By reconciling theoretical tightness with practical efficiency, this framework advances robust mean estimation from ad hoc heuristics toward a principled, unified design paradigm.

Addresses mean estimation with heavy-tailed dataDevelops robust estimators for contaminated datasetsEnsures privacy preservation in statistical methods

PIF: Anomaly detection via preference embedding

Jan 10, 2021
FL
Filippo Leveni
🏛️ Politecnico di Milano | Università della Svizzera italiana

This paper addresses anomaly detection in structured data by proposing Preference-based Isolation Forest (PIF), a novel method that maps raw data into a preference-driven high-dimensional embedding space and constructs a PI-Forest tree structure for efficient anomaly scoring. Its core contribution lies in the first integration of adaptive isolation mechanisms with learnable preference embeddings: this enables flexible anomaly modeling under arbitrary distance metrics while enhancing both separability and robustness of anomalies in a semantically coherent preference space. Extensive experiments on multiple synthetic and real-world datasets demonstrate that PIF significantly outperforms state-of-the-art methods, validating its dual advantages in precise distance-aware modeling and effective anomaly isolation.

Combining adaptive isolation with preference embeddingDetecting anomalies in structured patternsMeasuring arbitrary distances in preference space

To address the limitations of existing scan statistics for sparse anomaly detection under multi-source heterogeneous coordinate systems—namely, their reliance on strong distributional assumptions and difficulty in calibration—this paper proposes a rank-based high-criticism scoring method. The approach leverages only the relative ordering among independent observations, avoiding parametric modeling entirely, and introduces a novel rank-based high-criticism framework to nonparametrically characterize detectability conditions. We theoretically establish that detection power is uniquely determined by the probability that an anomalous observation exceeds a typical one, and prove asymptotic optimality under both exponential families and convolution models. The method achieves strong robustness and theoretical interpretability: it successfully identifies process anomalies in pharmaceutical quality control data, and simulations demonstrate performance approaching that of the oracle test.

Analyzes conditions for anomaly detection in non-parametric, robust mannerDetects anomalies in large datasets from multiple independent sourcesUses rank-based higher criticism without stringent modeling assumptions

Latest Papers

What's happening recently
View more

This work addresses the problem of k-means clustering in the presence of outliers by proposing a robust method based on k-nearest neighbor (KNN) distances. Given a budget to remove z outliers, the approach simply discards the z points with the largest KNN distances and then applies standard k-means clustering to the remaining data. Under a mild assumption on the minimum optimal cluster size, the method is the first to theoretically guarantee a constant-factor approximation—comparable to existing algorithms—without requiring additional cluster centers or excessive point removal, thereby establishing a formal connection between outlier detection and robust clustering. Empirical evaluations on multiple real-world datasets demonstrate that the proposed method is both efficient and practical, achieving clustering costs and running times that match or outperform several more complex baseline algorithms.

clusteringKNN-based heuristicoutlier detection

This study addresses the lack of a unified anomaly detection framework for mixed-type data comprising both continuous and ordinal categorical variables. The authors propose a robust approach based on a latent Gaussian variable model, wherein non-anomalous observations are assumed to follow a multivariate Gaussian distribution, and ordinal variables are represented through underlying latent Gaussian variables. Parameter estimation is performed using the Minimum Covariance Determinant (MCD) estimator, explicitly accounting for potential incompleteness in the observed ordinal information. Theoretical analysis demonstrates that the method effectively identifies extreme outliers and maintains robustness even under data contamination. Empirical evaluations on synthetic datasets show high detection rates coupled with low false alarm rates, and the method’s practical utility is further validated through application to real-world Airbnb listing data.

anomaly detectioncontinuous variablesmixed-type data

This work addresses the high computational cost of leave-one-out density refitting in unsupervised anomaly detection by proposing a density-based leave-one-out influence score. The score quantifies local perturbation by measuring the difference in density estimates at a fixed grid and bandwidth before and after removing a given observation. For the first time, we derive a closed-form update formula for this score under the linear bin frequency polygon (LBFP) density estimator, significantly improving computational efficiency while preserving interpretability of the underlying density. Theoretically, we establish an asymptotic order separation between normal and anomalous points under this score. Empirical results demonstrate that the proposed method matches or outperforms state-of-the-art baselines across various contamination models, with low computational overhead, and achieves strong detection performance on a 29-dimensional real-world credit card fraud dataset.

computational efficiencydensity estimationleave-one-out

This work proposes a novel approach that integrates conditional anomaly detection with instance-based learning to identify context-dependent anomalous behaviors in clinical settings, such as inappropriate hospital admission decisions or unwarranted ordering of HPF4 tests for heparin-induced thrombocytopenia. By incorporating conditional attributes to constrain the scope of anomaly assessment and leveraging multiple distance metrics enhanced through metric learning, the method accurately pinpoints anomaly instances most sensitive to their contextual conditions. Experimental results on two real-world clinical tasks—admission decisions for community-acquired pneumonia and HPF4 test ordering—demonstrate that the proposed framework substantially improves both the performance and clinical interpretability of conditional anomaly detection.

conditional anomaly detectionHPF4 test ordersinstance-based methods

This work addresses the challenge of effectively identifying anomalous patterns that depend on a subset of attributes when other attributes are given as conditions. To this end, we propose a metric learning approach tailored for conditional anomaly detection. The method adaptively learns a distance metric that explicitly models conditional dependencies among attributes and integrates instance-level neighborhood analysis to enable precise local anomaly discrimination. In contrast to existing instance-based approaches that rely on fixed distance metrics, our approach significantly improves detection accuracy and demonstrates enhanced capability in recognizing critical anomalous instances under dynamic conditional scenarios.

anomaly detectionconditional anomaly detectiondistance metric learning

Hot Scholars

JC

Jianfei Chen

Associate Professor, Tsinghua University
Machine Learning
ZY

Zhihang Yuan

Bytedance
Efficient AIModel CompressionMLLM
MF

Minghong Fang

University of Louisville
SecurityPrivacyAI SafetyMachine Learning
MV

Michal Valko

Chief Models Officer @ Stealth Startup, Inria & MVA - Ex: Llama at Meta; Gemini and BYOL @ Deepmind
large language modelsreasoningfine-tuningtest-time computation
MH

Mia Hubert

Professor of Statistics, KU Leuven
Robust statisticsOutlier detectionDepth