Score
Designs and implements diagnostic evaluations, scoring methods, and metrics that detect and quantify out-of-distribution (OOD) inputs and related failure modes, including unified SAE-style OOD scores, temporal OOD scoring for drift and consistency, and per-instance or personalized OOD detectors. Builds held-out OOD testbeds, evaluation protocols, and analysis routines to compare robustness across random seeds and interventions, combine semantic and visual anomaly signals, and report standardized OOD evaluation metrics.
Amid the rise of vision-language models (VLMs/LVLMs), out-of-distribution (OOD) detection and related tasks—such as anomaly detection, novelty detection, open-set recognition, and outlier detection—suffer from conceptual ambiguity and paradigmatic fragmentation. This paper proposes “Generalized OOD Detection v2”, a unified framework that systematically clarifies semantic boundaries and evolutionary relationships among these tasks, identifies OOD detection and anomaly detection as the central challenges, and formalizes novel evaluation paradigms and problem settings introduced by LVLMs (e.g., GPT-4V). Leveraging CLIP-based semantic alignment analysis, task taxonomy modeling, benchmark evolution comparison, and cross-task method review, the work synthesizes over 100 studies. The resulting VLM-driven conceptual framework redefines foundational assumptions, pinpoints critical technical challenges—including semantic misalignment, evaluation inconsistency, and LVLM-specific failure modes—and charts concrete directions for future research, establishing itself as the definitive survey in this rapidly evolving domain.
Detecting out-of-distribution (OOD) samples at test time remains challenging: ID-only methods suffer from limited discriminative capacity, while leveraging external anomaly data introduces privacy risks and task misalignment. Method: We propose AUTO, the first framework for *test-time adaptive OOD detection*, which requires no predefined anomaly data. Instead, it dynamically leverages unlabeled, real-world OOD samples from the incoming test stream to continuously refine the detector online. Contributions/Results: AUTO introduces three key components: (i) an in-out-aware filter for safe in-distribution sample selection; (ii) a dynamic memory module enabling robust replay of historical OOD patterns; and (iii) a prediction alignment objective preserving model stability. Guided by pseudo-labels, online gradient calibration, and test-time model adaptation, AUTO significantly outperforms state-of-the-art methods across standard, multi-OOD, and temporal OOD benchmarks—achieving superior detection accuracy and generalization robustness.
The out-of-distribution (OOD) detection field lacks a systematic, scenario-aware taxonomy, hindering principled comparison and advancement. Method: We propose the first task-oriented, unified taxonomy—categorizing OOD methods into *training-driven*, *training-agnostic*, and *foundation-model-based* paradigms, grounded in problem scenarios and model access constraints; we formally establish foundation-model-enabled OOD detection as an independent paradigm and systematically survey its adaptation mechanisms, including test-time adaptation and multimodal extensions. Contribution/Results: Our work establishes an extensible classification framework, clarifies evaluation challenges and deployment bottlenecks, and releases a high-quality, curated literature repository on GitHub. This advances OOD detection from ad hoc method enumeration toward structured evolution and practical deployment.
Existing out-of-distribution (OOD) detection methods are often confined to single technical paradigms or specific OOD categories, limiting their generalizability and robustness. To address this, we propose the Multi-Method Ensemble (MME) scoring framework—a unified, extensible OOD detection paradigm that systematically integrates feature truncation, multiple scoring functions (e.g., energy score, Mahalanobis distance), and logit-layer output fusion. Theoretical analysis and empirical evaluation demonstrate synergistic gains among components, substantially improving discrimination between near- and far-OOD samples. Evaluated on over ten benchmarks—including ImageNet-1K—using pre-trained models such as BiT, MME achieves state-of-the-art performance: an average false positive rate at 95% true positive rate (FPR95) of 27.57% on ImageNet-1K, outperforming the best prior method by 6 percentage points. This advancement establishes a more robust foundation for safety-critical open-world applications.
Existing out-of-distribution (OOD) detection evaluation suffers from unreliability and susceptibility to data split bias. To address this, we propose Dual-CV, a dual cross-validation evaluation framework. It applies standard k-fold cross-validation on in-distribution (ID) data and employs class-aware leave-one-class-out cross-validation on OOD data—grouped by semantic categories and aligned with hierarchical class structure to ensure fair, semantically meaningful splits. Dual-CV supports unified evaluation of OOD detectors both with and without anomaly exposure. Experiments demonstrate that Dual-CV significantly improves evaluation stability and cross-scenario consistency, accelerates convergence to true model performance, and exhibits robustness and generalizability across multiple benchmarks.
In open-world settings where ground-truth labels are unavailable, automatically selecting the optimal out-of-distribution (OOD) detection model remains an unsolved challenge. Method: This paper proposes the first zero-shot, unsupervised meta-learning framework for OOD model selection. It innovatively leverages large language models to generate task-level OOD characteristic embeddings, enabling cross-dataset task similarity modeling; it then integrates meta-learning with nonparametric statistical significance testing (Wilcoxon signed-rank test, *p* < 0.01) to perform label-free model ranking. Results: Evaluated across 24 dataset pairs and 11 OOD detectors, our method consistently outperforms all baselines by significant margins while incurring negligible inference overhead. It delivers a highly reliable, low-dependency solution for adaptive OOD model selection—critical for safety-sensitive real-time applications including online transaction monitoring, autonomous driving, and clinical decision support.
This work addresses the common practice of treating out-of-distribution (OOD) detection and in-distribution (ID) misclassification prediction as separate tasks, despite their intrinsic connection in building reliable classifiers. To bridge this gap, the authors propose SURE+, a unified framework that jointly models both tasks through a dual-scoring mechanism. The study further introduces novel joint evaluation metrics—DS-F1 and DS-AURC—to holistically assess performance across OOD and ID failure detection. Comprehensive experiments on the OpenOOD benchmark demonstrate that SURE+ significantly outperforms conventional single-score approaches, with particularly pronounced gains in scenarios involving easy or far-OOD samples. This work thus establishes a new paradigm and benchmark for trustworthy classification by explicitly integrating OOD detection and ID error prediction into a cohesive framework.
This study addresses the lack of systematic evaluation of existing test selection metrics under multi-objective settings, distribution shifts, and multimodal data—challenges that hinder practical metric selection. To bridge this gap, the authors construct the first unified benchmark encompassing three testing objectives (fault detection, performance estimation, and retraining guidance), five types of distribution shifts, three data modalities (images, text, and Android packages), and 13 deep learning models. Through a large-scale empirical study involving 1,640 experimental scenarios, they conduct rigorous statistical analyses to comprehensively compare the performance of 15 widely used metrics, elucidate their respective applicability boundaries, and provide reliable guidance and actionable recommendations for test selection in safety-critical systems.
Deep generative models often fail in out-of-distribution (OOD) detection due to unreliable likelihood estimates. This work proposes SITN, a method that leverages the diffeomorphic nature and mass-conservation property of continuous normalizing flows to perform an unsupervised goodness-of-fit test on individual samples within a factorized latent space. Requiring no OOD data and incurring low computational overhead, SITN rigorously controls the false positive rate and effectively mitigates the complexity bias inherent in conventional likelihood-based approaches. Empirical evaluations demonstrate that SITN achieves superior OOD detection performance on both standard benchmarks and synthetically perturbed datasets, significantly avoiding likelihood misinterpretation while enabling precise false positive control.
This work addresses the challenge of detecting out-of-distribution (OOD) inputs in deployed machine learning models, which often suffer from performance degradation under distributional shift. The authors propose TV-OOD, a novel OOD detection method that, for the first time, leverages Total Variation (TV) as a core principle. By employing a Total Variation network estimator to quantify each input’s contribution to the overall total variation, TV-OOD constructs a detection criterion that requires neither additional training nor generative models. Extensive experiments across multiple image classification architectures and standard benchmark datasets demonstrate that TV-OOD achieves performance comparable to or better than current state-of-the-art methods under common OOD evaluation metrics, thereby confirming its effectiveness and broad applicability.
This work addresses the vulnerability of large language models to alignment failures under out-of-distribution (OOD) conditions, a risk inadequately captured by existing monitoring mechanisms. To systematically evaluate OOD alignment failure detection, the authors introduce MOOD, a benchmark comprising a constrained training set and seven diverse OOD test sets. They propose a novel hybrid monitoring paradigm that integrates safety classifiers with OOD detectors based on Mahalanobis distance and perplexity. Evaluated across multiple model scales, this approach improves recall from 39% to 45% and exhibits positive scaling with model size, significantly outperforming a pure guard model even when the latter’s parameter count is increased twentyfold.