out-of-distribution detection

Designs and implements methods and diagnostics to detect inputs that fall outside a model’s training operating range by measuring distances, densities, likelihoods, or other discrepancy scores in feature space and deriving operational thresholds. Builds analyses, visualizations, acceptance/gating rules, and alerting to flag anomalous or out-of-distribution samples for cautious handling, retraining, or human review.

out-of-distributiondetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Jul 31, 2024
AM
Atsuyuki Miyai
🏛️ The University of Tokyo | Nanyang Technological University | Duke University | Salesforce AI Research | LY Corporation | Tokyo University of Science | University of Wisconsin-Madison | NTU

Amid the rise of vision-language models (VLMs/LVLMs), out-of-distribution (OOD) detection and related tasks—such as anomaly detection, novelty detection, open-set recognition, and outlier detection—suffer from conceptual ambiguity and paradigmatic fragmentation. This paper proposes “Generalized OOD Detection v2”, a unified framework that systematically clarifies semantic boundaries and evolutionary relationships among these tasks, identifies OOD detection and anomaly detection as the central challenges, and formalizes novel evaluation paradigms and problem settings introduced by LVLMs (e.g., GPT-4V). Leveraging CLIP-based semantic alignment analysis, task taxonomy modeling, benchmark evolution comparison, and cross-task method review, the work synthesizes over 100 studies. The resulting VLM-driven conceptual framework redefines foundational assumptions, pinpoints critical technical challenges—including semantic misalignment, evaluation inconsistency, and LVLM-specific failure modes—and charts concrete directions for future research, establishing itself as the definitive survey in this rapidly evolving domain.

Clarifies OOD detection and anomaly detection challenges in VLM eraGeneralized OOD detection framework unifies related problems in VLMsReviews methodology and benchmarks for OOD detection in LVLM era

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of unreliable sensor readings in industrial inspection robots caused by occlusions, limited viewpoints, or environmental anomalies, which hinder real-time task status assessment. The authors propose a hybrid framework that integrates supervised fault classification with unsupervised anomaly detection, uniquely combining conformal prediction and world models to enable policy-agnostic, distribution-free early discrimination among three states—success, known faults, and out-of-distribution anomalies—using compressed video inputs. The approach facilitates training data quality evaluation and model feedback, achieving over 90% recognition accuracy on both office and industrial instrument inspection datasets. It outperforms human observers in decision speed and has been successfully deployed on a Boston Dynamics Spot robot for real-time operation.

Anomaly DetectionAutonomous InspectionFailure Classification

Out-of-Distribution Detection Methods Answer the Wrong Questions

Jul 02, 2025
YL
Yucen Lily Li
🏛️ New York University | Capital One

Current mainstream out-of-distribution (OOD) detection methods suffer from a fundamental objective misalignment: supervised models trained exclusively on in-distribution (ID) data erroneously equate high predictive uncertainty or large feature-space distances with OOD samples, leading to irreducible detection errors. Method: The authors systematically analyze major paradigms—including uncertainty estimation, feature-logit hybrid models, density estimation, and generative approaches—and rigorously prove their theoretical limitations under common distribution shifts. Contribution/Results: They introduce the “objective misalignment” framework, establishing that OOD detection is inherently a distribution discrimination task—not an uncertainty or distance regression problem. Crucially, they demonstrate that prevalent mitigation strategies—such as anomaly exposure, model architecture expansion, or cognitive uncertainty modeling—cannot rectify this foundational flaw. This work provides critical theoretical grounding for paradigmatic reformulation of OOD detection.

Classifier trained on in-distribution data fails to detect OODCurrent OOD detection methods misidentify distribution shiftsUncertainty and feature-based methods conflate wrong OOD indicators

Introducing 'Inside' Out of Distribution

Jul 05, 2024
TL
Teddy Lazebnik
🏛️ Ariel University | University College London

Existing out-of-distribution (OOD) research predominantly focuses on extrapolative (“outside”) anomalies while overlooking interpolative (“inside”) in-distribution anomalies—i.e., samples that reside within the support of the training distribution yet are semantically anomalous. Method: This work formally defines and empirically validates the “inside OOD” concept, proposing a two-dimensional analytical framework to distinguish inside from outside OOD. Leveraging statistical distribution analysis, geometric modeling in feature space, and multi-model robustness evaluation, we systematically characterize their co-occurrence patterns and differential impacts on model behavior. Results: We demonstrate that inside OOD triggers latent, progressive performance degradation, whereas outside OOD induces sharp confidence collapse. Our framework bridges a critical gap in conventional OOD detection, providing both theoretical grounding and empirical evidence for designing targeted defense mechanisms against distinct OOD categories.

Addressing neglect of interpolatory inside OOD in current studiesDistinguishing inside versus outside out-of-distribution detection casesExamining unique impacts of inside-outside OOD profiles on ML performance

Industrial anomaly detection is often hindered by the extreme scarcity of fault samples, leading to suboptimal model performance and limited generalization. This work addresses this challenge by constructing a problem-agnostic hyperspherical synthetic dataset to systematically evaluate 14 anomaly detection algorithms—including kNN, LOF, XGBOD, SVM, and CatBoost—under rigorously controlled conditions across varying fault rates (0.05%–20%) and training set sizes. The study quantitatively demonstrates for the first time that unsupervised methods achieve optimal performance when fewer than 20 fault samples are available; semi-supervised and supervised approaches significantly outperform others with 30–50 fault samples; and further increasing normal samples yields diminishing returns. Additionally, feature dimensionality is found to critically influence the efficacy of semi-supervised methods, thereby clarifying the operational boundaries of each algorithmic category.

anomaly detectionclass imbalancefaulty data scarcity

This work addresses the performance gap between academic benchmarks and real-world deployment in unsupervised anomaly detection, where existing methods often exhibit instability, sensitivity to preprocessing, and inconsistent behavior in industrial settings. The authors conduct a systematic evaluation of 19 models on BowTie, a complex manufacturing dataset, revealing significant discrepancies between benchmark results and practical efficacy. To bridge this gap, they propose a human-in-the-loop unified detection framework that integrates SAM-generated refined candidate regions, heatmap-guided inspection, mask-based evaluation, and interactive verification. This framework enables quality inspectors to efficiently confirm defects, refine boundaries, and trace historical cases. Preliminary deployment demonstrates that the system substantially enhances both reliability and efficiency in industrial visual inspection.

anomaly detectionbenchmark gapindustrial deployment

Latest Papers

What's happening recently
View more

Risk valuation systems are susceptible to undetected errors caused by data failures, misconfigurations, or anomalies, potentially leading to significant operational losses. This work proposes EQAF, a hierarchical unsupervised ensemble framework for anomaly detection that uniquely integrates domain-specific deterministic rules with multiple complementary statistical outlier detection methods to enable real-time integrity monitoring of risk computation outputs. EQAF effectively identifies subtle anomalies—such as “frozen values”—that are often missed by conventional purely statistical approaches. Experimental evaluation on four real-world risk datasets demonstrates that EQAF achieves F1 scores between 61% and 79% and improves AUC-ROC by 4–6 percentage points over the best individual baseline method, substantiating its robustness and effectiveness.

anomaly detectiondata qualityoperational risk

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work proposes a novel paradigm for anomaly detection, termed On-Model AD, which leverages the intrinsic knowledge of the primary model—specifically, the normal output ranges of its neurons—to detect anomalies without requiring a separate, dedicated detection model. Traditional approaches often deploy independent anomaly detectors, overlooking the rich distributional information already embedded within the main model, thereby introducing redundancy and inefficiency. In contrast, the proposed method eliminates the need for additional training or deployment overhead. Building upon this paradigm, we introduce RangeAD, an algorithm that achieves strong detection performance in high-dimensional tasks while substantially reducing inference costs, effectively balancing accuracy and computational efficiency.

anomaly detectiondistributional shiftmachine learning

Hot Scholars

SG

Stephan Günnemann

Professor of Computer Science, Technical University of Munich
Machine LearningGraphsGraph Neural NetworksRobustness
TL

Tongliang Liu

Director, Sydney AI Centre, University of Sydney & Mohamed bin Zayed University of AI
Machine LearningLearning with Noisy LabelsTrustworthy Machine Learning
JH

Jungong Han

Chair Professor in Computer Vision, University of Sheffield, UK, FIAPR, FAAIA
Computer VisionVideo AnalyticsMachine Learning
AK

Anthony Kobanda

PhD @ Inria Scool & Ubisoft
Deep LearningReinforcement LearningContinual Learning