Score
Design and implement methods that adjust model outputs or training labels so their aggregate label frequencies match a specified target distribution, including estimating target priors and applying post-hoc scaling or reweighting to correct class imbalance or systematic biases. Build diagnostics and validation procedures to measure, compare, and stabilize differences between predicted and desired label distributions (e.g., to mitigate central tendency or other distributional bias).
This survey systematically addresses distribution shift between training and deployment in machine learning, focusing on two fundamental challenges: covariate shift (changes in input feature distributions) and concept shift (changes in semantic or class-conditional label distributions). We formalize and unify shift taxonomy, integrating techniques—including distribution shift detection, uncertainty estimation, domain adaptation, anomaly identification, causal inference, and invariant representation learning—within a cohesive framework bridging statistical learning and deep learning. Our key contributions include: (i) a novel robust modeling framework designed to handle heterogeneous shift types; (ii) the first systematic taxonomy covering out-of-distribution (OOD) scenarios; and (iii) a critical analysis revealing limitations of existing methods in jointly mitigating multiple concurrent shifts and generalizing to unseen classes. We establish principled evaluation criteria and outline future research directions—particularly addressing compound shifts and semantic evolution—thereby filling a critical gap in prior surveys, which largely overlook real-world deployment complexities involving intertwined and dynamically evolving shifts.
This paper addresses data imbalance in regression tasks—where target variables are continuous—a problem extensively studied in classification but lacking systematic investigation in regression. We propose the first taxonomy of resampling methods specifically designed for imbalanced regression. Our framework systematically evaluates oversampling, undersampling, and hybrid strategies across three dimensions: regression models (linear regression, tree-based models, neural networks), learning processes, and specialized evaluation metrics (uM, wMAE). Experimental results demonstrate that judicious resampling significantly improves predictive accuracy in sparse regions of the target space, and that model sensitivity to sampling strategies varies substantially. We uncover mechanistic insights into how resampling operates effectively in continuous output spaces. To foster reproducibility and further research, we publicly release all source code and benchmark datasets. This work establishes a rigorous, extensible analytical framework and practical guidelines for addressing imbalance in regression.
The impact of class imbalance correction on model discriminative performance and probability calibration in clinical risk prediction remains unclear. This study systematically evaluates the effects of SMOTE, random oversampling (ROS), and random undersampling (RUS) across ten real-world clinical datasets using a range of linear and nonlinear models. Comprehensive comparisons are conducted using metrics including ROC-AUC, Brier score, and calibration intercept/slope. Results indicate that none of the three resampling methods significantly improve discrimination, yet all consistently degrade probability calibration—evidenced by increased Brier scores (0.029–0.080) and substantial shifts in calibration parameters—revealing systematic distortion in predicted risk estimates. These findings challenge the conventional use of resampling techniques in clinical prediction modeling.
Whether class imbalance correction improves the performance of clinical prediction models remains controversial. This study leverages data from the GUSTO-I clinical trial to systematically evaluate the impact of various correction strategies—including algorithm-level rebalancing, oversampling, and hybrid sampling—on model discrimination (AUC), calibration (calibration plots and MAPE), and predictive stability (Classification Instability Index, CII) across varying sample sizes. Using penalized logistic regression with 200 bootstrap replications, we find that all correction methods fail to enhance discriminative performance and instead introduce greater calibration bias, risk overestimation, and increased prediction instability. These results challenge the common practice of routinely applying class imbalance corrections in clinical modeling and, for the first time in large-scale simulations, reveal their potential harms.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
This work addresses the limitation of class-level evaluation metrics, which often obscure performance disparities among intra-class sub-concepts—particularly when classes are imbalanced and sub-concept distributions are skewed, leading to biased assessments. To mitigate this issue without requiring ground-truth sub-concept labels, the authors propose a utility-weighted evaluation framework that constructs uncertainty-aware soft weights from the posterior probabilities of a multi-class sub-concept model and introduces the prediction-weighted balanced accuracy (pBA). This approach enables, for the first time, a stable and interpretable evaluation grounded solely in predicted probabilities. Empirical results across tabular, medical imaging, and textual datasets demonstrate that conventional unweighted metrics can be misleading under intra-class heterogeneity, whereas pBA provides a more reliable performance measure under non-pathological, imbalanced sub-concept distributions.
Traditional cross-validation in spatial prediction suffers from biased risk estimation due to distributional mismatches between validation and deployment tasks, including covariate shift and task difficulty shift. This work proposes Target-Weighted Cross-Validation (TWCV), which for the first time incorporates task distribution alignment into spatial prediction by calibrating weights to match the target-domain distribution and enhancing task difficulty diversity through spatial buffering-based resampling. By integrating importance-weighted risk estimation with task descriptor modeling, TWCV substantially reduces estimation bias. Empirical evaluations on both synthetic data and real-world environmental pollution mapping demonstrate that the method yields more accurate and nearly unbiased estimates of deployment risk compared to existing spatial cross-validation approaches.
This work addresses the instability and performance degradation commonly observed during fine-tuning of pre-trained models, which often stems from gradient cancellation leading to optimization collapse. To mitigate this issue, the paper introduces, for the first time in the context of fine-tuning, a dynamic gradient scaling mechanism, proposing the Dynamic Scaled Gradient Descent (DSGD) algorithm. DSGD adaptively attenuates the gradient magnitudes of correctly classified samples, thereby effectively alleviating gradient cancellation. The method substantially enhances fine-tuning stability and robustness, consistently reducing performance variance and achieving higher accuracy than existing approaches across multiple benchmark datasets and large-scale models.