perform robust outlier rejection

Design, implement, and evaluate algorithms and pipelines that detect, remove, downweight, or otherwise mitigate anomalous data points and long-tailed measurements across datasets and model internals, including mixed-type and mixed-precision inputs and activation channels. This work includes building outlier detectors and rejection methods, robust statistical estimators and thresholds, routing or re-encoding of outlier columns to higher precision, merging dual-branch or mixed-precision results, and performing theoretical and empirical robustness analyses and tests.

performrobustoutlierrejection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$180K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

In data stream regression, simultaneous occurrence and indistinguishability of outliers and concept drift—particularly under continuous output spaces—pose significant challenges. To address this, we propose a dual-channel joint detection framework: a fast channel employs residual analysis coupled with Exponentially Weighted Moving Absolute Deviation (EWMAD) for real-time point outlier filtering; a deep channel integrates dynamic threshold adaptation with an EWMAD-based Drift Type Decision Tree (EWMAD-DT) to distinguish abrupt from gradual concept drift online. This is the first approach enabling synchronous, fine-grained identification of both outliers and drift types, achieving both low latency and high accuracy. Extensive evaluation on multiple synthetic and real-world datasets demonstrates substantial improvements over state-of-the-art baselines, validating the method’s effectiveness and practical applicability.

Differentiates between abrupt and incremental concept driftsHandles continuous output spaces in regression tasksJointly detects outliers and concept drifts in data streams

This work addresses the performance degradation of machine learning interatomic potentials caused by noise from unconverged or inconsistent electronic structure calculations in training data. Existing denoising approaches rely on manual curation or iterative retraining, which are computationally expensive. To overcome this limitation, the authors propose an unsupervised online denoising method that dynamically tracks the loss distribution during a single training run using exponential moving averages, enabling real-time detection and automatic down-weighting of anomalous samples without requiring additional reference calculations or iterative retraining. The method successfully recovers accurate diffusion coefficients from unconverged liquid water data and reduces energy prediction errors by a factor of three on the SPICE dataset of organic molecules, demonstrating high efficiency, scalability, and robustness.

data qualitymachine learning interatomic potentialsnumerical noise

Approximations to worst-case data dropping: unmasking failure modes

Aug 16, 2024
JY
Jenny Y. Huang
🏛️ Massachusetts Institute of Technology | MIT

This paper addresses the challenge of robustness diagnostics in ordinary least squares (OLS) linear regression—specifically, whether removing a small number of data points induces substantial changes in statistical inference. We demonstrate that mainstream approximation methods systematically fail to detect influential subsets under realistic data configurations. To resolve this, we propose a recursive greedy algorithm and provide the first rigorous proof of its 100% detection success rate across multiple benchmark datasets. Comprehensive evaluations on both synthetic and real-world data show that our method achieves perfect detection accuracy (zero failure rate), outperforming existing approaches by up to several orders of magnitude in computational efficiency. Our core contribution lies in exposing the fundamental theoretical limitations of approximate influence assessment and delivering the first exact solution that simultaneously offers provable guarantees and practical reliability.

Assessing algorithm performance in OLS linear regressionDetecting non-robustness in data analysis conclusionsEvaluating approximations for worst-case data dropping

This work addresses the challenge of detecting both scattered and clustered anomalies in IoT data, where the latter—due to their high local density—are often misclassified as normal instances, thereby degrading detection performance. To tackle this issue, the authors propose an unsupervised graph-based anomaly detection method that constructs natural neighbor relationships among data points and introduces, for the first time, a hierarchical reference set mechanism. This mechanism enables coordinated anomaly assessment across multiple scales—both local and global—effectively distinguishing between the two types of anomalies while preventing mutual interference. Experimental results demonstrate that the proposed approach significantly outperforms existing methods, not only improving anomaly detection accuracy but also enhancing downstream clustering performance, all while exhibiting strong robustness to hyperparameter settings.

clustered outliersIoT dataoutlier detection

The Broader Landscape of Robustness in Algorithmic Statistics

Dec 03, 2024
GK
Gautam Kamath
🏛️ University of Waterloo | Vector Institute

This work addresses the robustness of mean estimation in statistical learning under three concurrent challenges: adversarial data contamination, heavy-tailed distributions, and differential privacy constraints. Methodologically, it unifies robust statistics, high-dimensional geometry, stochastic optimization, and differential privacy theory to establish the first conceptual and algorithmic bridge across distinct robustness paradigms. Key technical abstractions—including iterative filtering, covariance trimming, and fractional gradient descent—are identified as common algorithmic primitives. The paper proposes a suite of computationally efficient estimators achieving statistically optimal convergence rates; each attains the information-theoretic lower bound under all three constraint classes simultaneously. By reconciling theoretical tightness with practical efficiency, this framework advances robust mean estimation from ad hoc heuristics toward a principled, unified design paradigm.

Addresses mean estimation with heavy-tailed dataDevelops robust estimators for contaminated datasetsEnsures privacy preservation in statistical methods

Latest Papers

What's happening recently
View more

This study addresses the lack of a unified anomaly detection framework for mixed-type data comprising both continuous and ordinal categorical variables. The authors propose a robust approach based on a latent Gaussian variable model, wherein non-anomalous observations are assumed to follow a multivariate Gaussian distribution, and ordinal variables are represented through underlying latent Gaussian variables. Parameter estimation is performed using the Minimum Covariance Determinant (MCD) estimator, explicitly accounting for potential incompleteness in the observed ordinal information. Theoretical analysis demonstrates that the method effectively identifies extreme outliers and maintains robustness even under data contamination. Empirical evaluations on synthetic datasets show high detection rates coupled with low false alarm rates, and the method’s practical utility is further validated through application to real-world Airbnb listing data.

anomaly detectioncontinuous variablesmixed-type data

This study addresses the challenge of distinguishing between mild and severe outliers in circular data by proposing a three-component Bayesian mixture model. The model employs a symmetric unimodal circular distribution—such as the von Mises or wrapped normal—as a reference component, incorporates a uniform distribution to capture severe outliers, and introduces a low-concentration component sharing the same mean to represent mild anomalies. This dual-contamination framework uniquely enables automatic identification and quantification of both outlier types within a unified probabilistic structure, without requiring predefined thresholds. It further yields interpretable estimates of outlier proportions and dispersion inflation. Simulation studies and real-data analyses—including applications to animal movement and wind direction—demonstrate that the proposed approach substantially enhances model robustness and effectively uncovers latent structures in directional data.

anomaly detectioncircular datagross anomalies

Risk valuation systems are susceptible to undetected errors caused by data failures, misconfigurations, or anomalies, potentially leading to significant operational losses. This work proposes EQAF, a hierarchical unsupervised ensemble framework for anomaly detection that uniquely integrates domain-specific deterministic rules with multiple complementary statistical outlier detection methods to enable real-time integrity monitoring of risk computation outputs. EQAF effectively identifies subtle anomalies—such as “frozen values”—that are often missed by conventional purely statistical approaches. Experimental evaluation on four real-world risk datasets demonstrates that EQAF achieves F1 scores between 61% and 79% and improves AUC-ROC by 4–6 percentage points over the best individual baseline method, substantiating its robustness and effectiveness.

anomaly detectiondata qualityoperational risk

Hot Scholars

SK

Simon Klüttermann

Phd Student, Computer science, TU Dortmund
anomaly detectionensemble learningmachine learning
MH

Mia Hubert

Professor of Statistics, KU Leuven
Robust statisticsOutlier detectionDepth
OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
LF

Long Feng

Professor of Nankai University
High Dimensional DataHigh Frequency Data