group robustness evaluation

Designing and applying evaluation metrics and protocols to quantify model robustness to distribution shifts and spurious correlations (e.g., worst-group accuracy), and comparing how methods (such as different distillation strategies) affect fine-grained classification performance under those shifts.

grouprobustnessevaluation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study systematically evaluates the robustness of various knowledge distillation methods under limited or non-representative data exhibiting spurious correlations, with a particular focus on SubDistill. To the best of our knowledge, this is the first comprehensive assessment of how such methods perform as the strength of spurious correlations varies. Experimental results demonstrate that the performance gap between advanced distillation approaches and conventional baselines widens significantly with increasing spurious correlation strength. Notably, SubDistill maintains robust performance even under strong spurious interference, whereas several baseline methods degrade to near-random accuracy. These findings highlight critical challenges in applying knowledge distillation to real-world scenarios where training data may be biased or unrepresentative, and provide empirical insights for designing more robust distillation strategies.

foundation modelsknowledge distillationrobustness

Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?

Mar 31, 2025
OS
Olawale Salaudeen
🏛️ Massachusetts Institute of Technology | Stanford University

Current domain generalization (DG) benchmarks—e.g., ColoredMNIST and Waterbirds—exhibit a fundamental flaw in evaluating model robustness to spurious correlations: their distribution shifts fail to meaningfully alter the spurious associations governing out-of-distribution (OOD) generalization, leading to the “accuracy-on-a-line” phenomenon and rendering them incapable of probing whether models truly disentangle spurious dependencies. Method: We introduce the novel concept of *benchmark misspecification*, establishing necessary conditions for robustness evaluation grounded in causal modeling and distribution shift analysis. We theoretically prove that mainstream DG benchmarks violate these conditions and derive verifiable criteria for *well-specified* benchmarks. Contribution/Results: Our work provides the first formal theoretical foundation for assessing spurious-correlation robustness, yielding both principled guidelines and practical criteria for designing credible, causally sound DG evaluation protocols.

Assessing misspecified benchmarks for spurious correlation robustnessIdentifying conditions for reliable spurious correlation shift evaluationRethinking domain generalization benchmark design for meaningful robustness

An Analysis of Model Robustness across Concurrent Distribution Shifts

Jan 08, 2025
MJ
Myeongho Jeon
🏛️ École Polytechnique Fédérale de Lausanne | Seoul National University | CRABs.ai | Samsung Research | Singapore-MIT Alliance for Research and Technology

This paper investigates the robustness degradation of machine learning models under concurrent distribution shifts—specifically, the co-occurrence of domain shift and spurious correlations. To this end, we establish a comprehensive benchmark spanning eight datasets, 168 source–target domain pairs, and 26 algorithms, involving over 100,000 model training and evaluation runs. We propose a multi-source–multi-target shift construction framework and a statistical attribution analysis methodology. Our large-scale empirical study is the first to systematically quantify the compounding effect of concurrent shifts; reveals positive cross-shift generalization transferability; and demonstrates that heuristic data augmentation consistently outperforms large-model zero-shot inference—achieving state-of-the-art average robustness on both synthetic and real-world benchmarks. Crucially, we identify a consistent cross-shift pattern in generalization improvement, providing both theoretical grounding and practical guidance for robust modeling in complex, realistic deployment scenarios.

Data VariabilityMachine Learning RobustnessPerformance Degradation

Quantifying classifier robustness under distributional shift and few-shot settings remains challenging due to the lack of reliable, sample-efficient reliability measures. Method: This paper proposes a prediction-level robustness assessment framework for generative probabilistic classifiers, introducing imprecise probability theory—previously unexplored in individual prediction reliability modeling—to jointly characterize generative uncertainty and robust decision-making. Unlike conventional uncertainty quantification paradigms, our approach does not rely on large-sample assumptions and yields verifiable confidence bounds even under scarce training data or test-time distribution shift. Contribution/Results: Empirical evaluation across multiple few-shot and distribution-shift benchmarks demonstrates significant improvements over standard uncertainty baselines. The method provides theoretically grounded, actionable reliability guarantees for high-stakes classification tasks, establishing a novel paradigm for trustworthy classification under limited data and non-stationary environments.

Assess prediction reliability in generative probabilistic classifiersMaintain performance with small training sets from shifted distributionsQuantify robustness for reliable classification under distribution shifts

Rethinking Robustness in Machine Learning: A Posterior Agreement Approach

Mar 20, 2025
JB
Joao Borges S. Carvalho
🏛️ ETH Zurich | Free University of Bozen-Bolzano | University of Genoa

This paper addresses the theoretical gap in evaluating machine learning model robustness under covariate shift. We propose the first unsupervised robustness metric framework grounded in Posterior Agreement (PA), diverging from conventional accuracy-based empirical measures. Our method systematically extends PA theory to the covariate shift setting, yielding a statistically principled, label-free robustness quantification paradigm. Through Bayesian model validation, adversarial perturbation analysis, and cross-domain generalization experiments, we demonstrate that the framework sensitively detects model vulnerabilities and maintains stable, reliable assessment performance—even under minimal distributional shifts. The core contribution is the establishment of the first theoretically grounded, computationally tractable, and label-free standard for robustness evaluation under covariate shift.

Assessing vulnerabilities in learning algorithms under distribution shifts.Evaluating robustness of machine learning algorithms against covariate shifts.Proposing a novel framework based on Posterior Agreement theory.

Latest Papers

What's happening recently
View more

This work addresses the problem of provenance shift—performance degradation under out-of-distribution scenarios caused by changes in the relationship between data sources and labels during deployment. It formally establishes, for the first time, the theoretical connection between provenance shift and counterfactual invariance within the framework of invariant learning, and proposes a learning objective tailored for robustness. The core contributions include the development of DeconDTN-Toolkit, the first open-source toolkit enabling simulation and mitigation of provenance shift; the introduction of a novel evaluation metric for out-of-distribution robustness; and systematic experiments that expose the fragility of empirical risk minimization approaches while demonstrating the effectiveness of the proposed strategy in enhancing model robustness.

counterfactual invariancedistribution shiftinvariant learning

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

Existing evaluation methods struggle to disentangle whether performance degradation under temporal distribution shifts stems from insufficient model adaptability or increased data difficulty. To address this, this work introduces a novel approach that decouples model adaptability from the inherent difficulty of temporal data for the first time. The authors propose three dynamic metrics based on performance trajectories, which capture the adaptation process through dynamic evaluation and comparative analysis. These new metrics uncover fine-grained adaptation patterns obscured by conventional assessment techniques, substantially enhancing the interpretability and depth of understanding of temporal robustness in machine learning models.

intrinsic difficultymodel adaptationperformance degradation

Current benchmarks for evaluating toxicity in large language models exhibit underappreciated systematic biases that may lead to the deployment of unsafe models. This work systematically investigates how variations in task formulation—such as text completion versus summarization—input data domains, and evaluated models interact with multiple toxicity metrics. It reveals, for the first time, that both task type and data domain significantly influence toxicity scores. Experiments demonstrate that existing benchmarks are prone to misclassifying content as harmful when tasks are altered and show inconsistent performance across domains, highlighting their fragility and dependence on specific model-task configurations. These findings underscore the urgent need for more robust and reliable toxicity evaluation frameworks.

benchmark robustnessevaluation biasLLM evaluation

Deep neural networks often rely on spurious correlations in high-stakes scenarios, compromising their reliability. This work presents the first systematic integration of distributionally robust optimization, invariant risk minimization, and shortcut learning frameworks to evaluate the debiasing efficacy of explainable AI (XAI) methods—particularly counterfactual knowledge distillation (CFKD)—against non-XAI baselines under conditions of data scarcity and severe subgroup imbalance. Experimental results demonstrate that XAI approaches generally outperform non-XAI alternatives, with CFKD exhibiting the most stable generalization performance. The study further reveals that the practical challenges of acquiring group labels and the sparsity of minority-group samples significantly undermine the reliability of model deployment in real-world settings.

Clever HansGroup-Distributional Non-robustnessModel Reliability

Hot Scholars

KS

Kyungwoo Song

Yonsei University
Machine LearningDeep LearningNeural Networks
WY

Wenhan Yang

P.hD. student of Computer Science, University of California, Los Angeles
Self-supervised LearningModel Robustness
CP

Charlie Pilgrim

University of Leeds
Collective IntelligenceCommunicationCooperation
JM

Jeffri Murrugarra-Llerena

Ph.D. Student, Stony Brook University
Computer visionNatural Language ProcessingMachine Learning