holdout evaluation design

Designs and implements held-out evaluation protocols that partition training and evaluation data or tasks into disjoint splits and define in-distribution vs out-of-distribution (OOD) manifolds so that model training cannot contaminate test measurements. Builds and analyzes validation experiments and metrics to quantify generalization and robustness across out-of-sample conditions and distribution shifts, diagnose spurious robustness, and compare performance between held-in and held-out scenarios.

holdoutevaluationdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.79
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Recent Advances in OOD Detection: Problems and Approaches

Sep 18, 2024
SL
Shuo Lu
🏛️ NLPR & MAIS | Institute of Automation | Chinese Academy of Sciences | Anhui University | University of Science and Technology of China | Meituan

The out-of-distribution (OOD) detection field lacks a systematic, scenario-aware taxonomy, hindering principled comparison and advancement. Method: We propose the first task-oriented, unified taxonomy—categorizing OOD methods into *training-driven*, *training-agnostic*, and *foundation-model-based* paradigms, grounded in problem scenarios and model access constraints; we formally establish foundation-model-enabled OOD detection as an independent paradigm and systematically survey its adaptation mechanisms, including test-time adaptation and multimodal extensions. Contribution/Results: Our work establishes an extensible classification framework, clarifies evaluation challenges and deployment bottlenecks, and releases a high-quality, curated literature repository on GitHub. This advances OOD detection from ad hoc method enumeration toward structured evolution and practical deployment.

Classifying methods as training-driven or training-agnosticDetecting test samples outside training categories for reliable MLSurveying OOD detection from a task-oriented perspective

DCV-ROOD Evaluation Framework: Dual Cross-Validation for Robust Out-of-Distribution Detection

Sep 06, 2025
AU
Arantxa Urrea-Castaño
🏛️ University of Granada | Andalusian Institute of Data Science and Computational Intelligence (DaSCI) | Department of Computer Science and Artificial Intelligence (DECSAI) | Department of Software Engineering (LSI)

Existing out-of-distribution (OOD) detection evaluation suffers from unreliability and susceptibility to data split bias. To address this, we propose Dual-CV, a dual cross-validation evaluation framework. It applies standard k-fold cross-validation on in-distribution (ID) data and employs class-aware leave-one-class-out cross-validation on OOD data—grouped by semantic categories and aligned with hierarchical class structure to ensure fair, semantically meaningful splits. Dual-CV supports unified evaluation of OOD detectors both with and without anomaly exposure. Experiments demonstrate that Dual-CV significantly improves evaluation stability and cross-scenario consistency, accelerates convergence to true model performance, and exhibits robustness and generalizability across multiple benchmarks.

Enables reliable assessment of OOD detection methods performanceIntegrates in-distribution and out-of-distribution data characteristicsProposes dual cross-validation for robust OOD detection evaluation

ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction

Oct 31, 2025
HY
Han Yu
🏛️ Tsinghua University | Renmin University of China

Existing out-of-distribution (OOD) performance prediction research suffers from inconsistent evaluation protocols and insufficient coverage of real-world OOD datasets and distribution shift types. Method: We introduce OOD-PPB—the first systematic OOD performance prediction benchmark—integrating 12 real-world datasets, 6 canonical distribution shift categories (e.g., semantic, compound), and state-of-the-art prediction algorithms, with standardized evaluation pipelines and pre-trained models to eliminate redundant training overhead. Crucially, OOD-PPB enables performance prediction evaluation under zero-shot, unlabeled OOD settings, supporting risk-sensitive deployment. Contribution/Results: Through extensive cross-dataset and cross-shift experiments, we systematically characterize the capabilities and limitations of existing methods for the first time, revealing significant failures under semantic and compound shifts. We publicly release code, models, and evaluation tools, establishing a reproducible, extensible, and authoritative testbed for future research.

Addressing limited real-world datasets and distribution shift typesBenchmarking inconsistent OOD performance prediction evaluation protocolsProviding fair algorithm comparisons with trained model testbench

Introducing 'Inside' Out of Distribution

Jul 05, 2024
TL
Teddy Lazebnik
🏛️ Ariel University | University College London

Existing out-of-distribution (OOD) research predominantly focuses on extrapolative (“outside”) anomalies while overlooking interpolative (“inside”) in-distribution anomalies—i.e., samples that reside within the support of the training distribution yet are semantically anomalous. Method: This work formally defines and empirically validates the “inside OOD” concept, proposing a two-dimensional analytical framework to distinguish inside from outside OOD. Leveraging statistical distribution analysis, geometric modeling in feature space, and multi-model robustness evaluation, we systematically characterize their co-occurrence patterns and differential impacts on model behavior. Results: We demonstrate that inside OOD triggers latent, progressive performance degradation, whereas outside OOD induces sharp confidence collapse. Our framework bridges a critical gap in conventional OOD detection, providing both theoretical grounding and empirical evidence for designing targeted defense mechanisms against distinct OOD categories.

Addressing neglect of interpolatory inside OOD in current studiesDistinguishing inside versus outside out-of-distribution detection casesExamining unique impacts of inside-outside OOD profiles on ML performance

AUTO: Adaptive Outlier Optimization for Test-Time OOD Detection

Mar 22, 2023
PY
Puning Yang
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences

Detecting out-of-distribution (OOD) samples at test time remains challenging: ID-only methods suffer from limited discriminative capacity, while leveraging external anomaly data introduces privacy risks and task misalignment. Method: We propose AUTO, the first framework for *test-time adaptive OOD detection*, which requires no predefined anomaly data. Instead, it dynamically leverages unlabeled, real-world OOD samples from the incoming test stream to continuously refine the detector online. Contributions/Results: AUTO introduces three key components: (i) an in-out-aware filter for safe in-distribution sample selection; (ii) a dynamic memory module enabling robust replay of historical OOD patterns; and (iii) a prediction alignment objective preserving model stability. Guided by pseudo-labels, online gradient calibration, and test-time model adaptation, AUTO significantly outperforms state-of-the-art methods across standard, multi-OOD, and temporal OOD benchmarks—achieving superior detection accuracy and generalization robustness.

Adapts OOD detector using real unlabeled test dataDetects test samples outside training distribution classesImproves OOD detection over state-of-the-art methods

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic evaluation of existing test selection metrics under multi-objective settings, distribution shifts, and multimodal data—challenges that hinder practical metric selection. To bridge this gap, the authors construct the first unified benchmark encompassing three testing objectives (fault detection, performance estimation, and retraining guidance), five types of distribution shifts, three data modalities (images, text, and Android packages), and 13 deep learning models. Through a large-scale empirical study involving 1,640 experimental scenarios, they conduct rigorous statistical analyses to comprehensively compare the performance of 15 widely used metrics, elucidate their respective applicability boundaries, and provide reliable guidance and actionable recommendations for test selection in safety-critical systems.

data modalitiesdistribution shiftsempirical evaluation

This study addresses the inflated out-of-distribution (OOD) detection performance in existing methods, which stems from models fitting dataset identity rather than genuine novelty. To resolve this, we introduce a novel whole-dataset hold-out protocol that decouples identity bias from novelty bias. By integrating posterior detection, feature-space directional fitting, and closed-form mathematical derivations, we quantify the inflation of reported predictive gains. Our analysis reveals that conventionally reported improvements are largely spurious and that model capacity is not the primary contributing factor. Furthermore, we establish that a single constant baseline serves as an upper bound for genuine performance gains. This baseline remains robust on a predefined validation set, effectively recalibrating the evaluation standards within the field.

Dataset identityEvaluation protocolNovelty detection

This study systematically investigates the impact of different training objectives on out-of-distribution (OOD) detection performance in image classification, with the aim of enhancing model robustness in safety-critical applications. Under the unified OpenOOD evaluation protocol, it presents the first comprehensive assessment of cross-entropy loss, prototype loss, triplet loss, and mean average precision loss across both near- and far-OOD detection settings on CIFAR-10/100 and ImageNet-200. The findings reveal that cross-entropy loss consistently achieves the most robust and reliable OOD detection performance while maintaining high in-distribution accuracy. Nevertheless, alternative objectives also demonstrate competitive results under specific configurations, offering empirical guidance for selecting training objectives tailored to OOD detection tasks.

Image classificationOOD performanceOpenOOD

This work identifies and theoretically analyzes domain-sensitive collapse (DSC)—a phenomenon wherein models trained on a single domain exhibit feature collapse into low-rank class subspaces, leading to failure in out-of-distribution (OOD) detection. To mitigate this, the authors propose Teacher-Guided Training (TGT), which leverages a frozen multi-domain pretrained teacher model (DINOv2) during training. TGT employs an auxiliary head to distill residual structures suppressed by class supervision, thereby recovering directions sensitive to domain shifts. Notably, TGT introduces no additional inference overhead and achieves substantial improvements across eight single-domain benchmarks, reducing far-OOD false positive rates by over 10 percentage points on average (measured by FPR@95 for ResNet-50), while maintaining or slightly enhancing in-domain OOD detection performance and classification accuracy.

domain shiftDomain-Sensitivity Collapsefeature subspace

Although learning models with noisy labels perform well on closed-set tasks, they suffer from “uncertainty collapse,” wherein misclassified in-distribution samples and out-of-distribution (OOD) samples become indistinguishable due to overlapping features and confidence scores. This work is the first to identify this phenomenon and introduces ACC-OOD, a learner-agnostic benchmark that uniformly evaluates a model’s ability to detect both near- and far-OOD samples. To mitigate uncertainty collapse, the authors propose Virtual Margin Regularization (VMR), a lightweight method that enhances OOD separability without compromising closed-set accuracy. Extensive experiments demonstrate that VMR significantly improves OOD detection performance while maintaining high in-distribution classification accuracy, underscoring the necessity of jointly evaluating noise robustness and open-world reliability.

closed-set accuracynoisy label learningout-of-distribution detection

Hot Scholars

SK

Subbarao Kambhampati

Arizona State University
Artificial IntelligenceAutomated planningLLM ReasoningHuman-AI Interaction
BQ

Bing Qin

Professor in Harbin Institute of Technology
Natural Language ProcessingInformation ExtractionSentiment Analysis
VP

Vardhan Palod

Student of Artificial Intelligence
Large Language ModelsReinforcement Learning
KV

Karthik Valmeekam

Student, Arizona State University
Large Language ModelsAutomated PlanningLLM ReasoningHuman Aware AI
KS

Kaya Stechly

Yale
Artificial IntelligenceCognitive ScienceLinguistics