mahalanobis scoring

Designs and implements scoring methods that compute Mahalanobis-distance-based outlier or anomaly scores for multivariate observations—including interval-valued data—by estimating means and covariances (including class-conditional covariances) and producing distances or squared distances for ranking and thresholding. This includes building robust variants (e.g., using MCD) and interval-aware Mahalanobis estimators to produce analytic scores used for detection, attribution, and alerting.

mahalanobisscoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the severe distortion of conventional mean and covariance estimates caused by outliers in interval-valued data. To this end, it introduces the first extension of the Minimum Covariance Determinant (MCD) estimator to the interval-valued setting, integrating robust location and scale estimates based on the Mallows distance to construct a robust interval Mahalanobis distance. An adaptive thresholding mechanism is further incorporated to enable effective outlier detection. The proposed method demonstrates superior performance over classical approaches across varying contamination levels, achieving notably higher accuracy in both covariance estimation and outlier identification. Its efficacy is validated through experiments on real-world datasets, confirming its practical applicability and robustness.

anomalous observationscovariance estimationinterval-valued data

Existing methods for interval-valued anomaly detection suffer from limited interpretability and struggle to pinpoint the root causes of anomalies. To address this, this work proposes a novel approach that decomposes robust interval Mahalanobis distances using Shapley values, grounded in the minimum covariance determinant estimator for interval data. The method derives, for the first time, a closed-form solution for Shapley values, enabling efficient quantification of each variable’s contribution across center, range, and cross-term components. Furthermore, it incorporates Shapley interaction indices to capture synergistic effects among variables that jointly drive anomalous behavior. Experimental results on two real-world datasets demonstrate that the proposed framework effectively identifies cell-level anomalies invisible at the multivariate level, substantially enhancing interpretability while maintaining robust detection performance.

Cellwise OutliersExplainable Outlier DetectionInterpretability

This work addresses the challenge of robust covariance estimation in online settings where data volume grows continuously and contamination rates increase, rendering traditional estimators vulnerable to bias and masking effects. The authors propose a novel method that simultaneously estimates the geometric median and the median-based covariance matrix in a streaming fashion—integrating these two robust statistics for the first time in an online framework. By computing Mahalanobis distances in real time using these estimates, the approach enables effective outlier detection while mitigating masking effects. The method maintains computational efficiency suitable for real-time applications and significantly enhances the robustness of covariance estimation. Experimental results on synthetic data demonstrate its accuracy in recovering true covariance structures and reliably identifying anomalies under contamination.

covariance matrixgeometric medianmasking effect

Model-Based Clustering with Sequential Outlier Identification using the Distribution of Mahalanobis Distances

May 16, 2025
UP
Ult'an P. Doherty
🏛️ Trinity College Dublin | McMaster University

Outliers severely impair clustering structure identification, yet existing methods often rely on prespecified outlier proportions or strong distributional assumptions. This paper proposes outlierMBC—a model-based, iterative framework that jointly performs clustering and outlier detection without requiring prior knowledge of outlier prevalence. Its core innovation lies in fitting a Gaussian mixture model (GMM), computing scaled squared Mahalanobis distances, and empirically estimating their distribution; it then automatically determines the optimal number of outliers to remove by minimizing the discrepancy between this empirical distribution and a Beta reference distribution. outlierMBC integrates model-driven clustering, density-ordered sequential outlier removal, and statistical distributional validation. Extensive experiments on synthetic and real-world datasets demonstrate that outlierMBC significantly improves clustering accuracy and adaptive outlier identification, consistently outperforming state-of-the-art baseline methods.

Combines Gaussian mixture modeling with sequential outlier detectionIdentifies outliers in clustering without pre-specifying their numberUses iterative Mahalanobis distance analysis for outlier removal

This work addresses the challenges of anomaly detection in high-dimensional data, where traditional methods often suffer from performance degradation and sensitivity to parameter settings. The authors propose a novel Multidimensional Outlier Detection (MDOD) algorithm that enhances discriminative power by extending the original data with an additional dimension filled with zeros and constructing vectors rooted at a designated observation point. Anomalies are identified through cosine similarity measures between these vectors. By innovatively integrating dimensionality expansion with an observation-point-based mechanism, MDOD significantly improves detection accuracy in high-dimensional settings. Empirical evaluations across multiple datasets demonstrate the method’s efficiency and effectiveness, and the implementation has been made publicly available as the open-source Python package “mdod” on PyPI.

anomaly identificationcosine similaritymulti-dimensional data

Latest Papers

What's happening recently
View more

This study addresses the lack of a unified anomaly detection framework for mixed-type data comprising both continuous and ordinal categorical variables. The authors propose a robust approach based on a latent Gaussian variable model, wherein non-anomalous observations are assumed to follow a multivariate Gaussian distribution, and ordinal variables are represented through underlying latent Gaussian variables. Parameter estimation is performed using the Minimum Covariance Determinant (MCD) estimator, explicitly accounting for potential incompleteness in the observed ordinal information. Theoretical analysis demonstrates that the method effectively identifies extreme outliers and maintains robustness even under data contamination. Empirical evaluations on synthetic datasets show high detection rates coupled with low false alarm rates, and the method’s practical utility is further validated through application to real-world Airbnb listing data.

anomaly detectioncontinuous variablesmixed-type data

This work proposes a robust and interpretable anomaly detection method for multivariate functional data with separable covariance structure. By establishing a connection between separable covariance stochastic processes and matrix-variate distributions under basis representations, the approach integrates matrix-variate minimum covariance determinant (MMCD) estimation with a truncated multivariate functional Mahalanobis semi-distance to achieve robust mean and covariance estimation. The study innovatively extends Shapley values to functional anomaly interpretation, reducing computational complexity from exponential to linear while preserving their axiomatic properties, thereby decomposing global anomaly scores into localized contributions over the time domain. Theoretical analysis, simulation studies, and real-data applications demonstrate the method’s superior performance in both robust estimation and interpretable anomaly detection.

explainable AImultivariate functional dataoutlier detection

This study addresses the challenge of identifying abnormal operating conditions in district heating systems by proposing an integrated multivariate anomaly detection approach that combines principal component analysis (PCA), Isolation Forest, and Hotelling’s T² test. The method explicitly accounts for practical operational constraints—such as data invalidity during zero-energy-consumption periods—by incorporating outdoor temperature and heat transfer data. Detected anomalies are further refined through domain expert evaluation to ensure relevance and validity. Experimental results demonstrate that the proposed ensemble framework accurately identifies non-normal operating states, offering reliable detection performance while supporting operational decisions aimed at reducing natural gas consumption and carbon emissions.

anomaly detectionCO2 emissionDistrict Heating System

This study addresses the lack of robust anomaly detection methods for interval-valued functional data (IVFD) by proposing the first projection-based robust anomaly detection framework. The approach decomposes interval-valued functions into center and log-radius components, derives a joint low-dimensional representation via interval-valued functional principal component analysis (IFPCA), and introduces the Interval-valued Least Trimmed Function Score (ILTFS) to identify a robust reference subset. Anomaly decisions are made by computing empirical p-values based on projection distances and controlling the false discovery rate (FDR) using the Benjamini–Hochberg procedure. Theoretically, the finite-sample breakdown point of ILTFS is established, and algorithmically, a concentration step iteration with descent properties is developed. Experiments demonstrate that the proposed ILTFS-FDR method achieves superior detection performance and robustness on both simulated and high-frequency ETF datasets.

false discovery ratefunctional data analysisinterval-valued functional data

Hot Scholars

NA

Nilesh Ahuja

Intel
Probabilistic methods in Machine LearningAnomaly DetectionComputer VisionImage and Video Processing
SB

Samir Brahim Belhaouari

Hamad Bin Khalifa University
Mathematics of Machine LearningScientific Machine LearningComputational Science
TB

Timothy Baldwin

MBZUAI and The University of Melbourne
computational linguisticsnatural language processingartificial intelligence
NK

Nikolaus Kriegeskorte

Professor of Psychology and Neuroscience, Columbia University
visionneural networksfMRIneuronal recordings