Score
Design and evaluate training objectives and loss functions that explicitly account for label sparsity and severe class imbalance—e.g., class-imbalance losses, reweighting, focal/margin-aware terms or sampling-aware penalties—and implement these to improve learning of rare classes, tighten decision boundaries, and stabilize model calibration and optimization when positive examples are scarce.
This paper addresses the lack of systematic guidance for loss function selection and design in deep learning. We propose the first cross-task loss taxonomy—covering vision, time-series, and tabular data—and unifying discriminative and generative paradigms. Through comprehensive survey analysis, rigorous mathematical modeling, and multi-scenario empirical evaluation, we characterize the applicability boundaries and failure modes of 12 mainstream losses, identifying three fundamental challenges: low computational efficiency, gradient instability, and poor adaptability to real-world constraints. Building on these insights, we formulate next-generation loss design principles centered on robustness, interpretability, and adaptivity. Furthermore, we deliver a practical, industry-deployment-oriented loss selection guide—grounded in empirical evidence and operational feasibility—to bridge the gap between theoretical design and practical application.
To address performance degradation caused by label noise in training data, this paper proposes two novel robust loss functions. Methodologically, the approach adapts the sample reweighting mechanism of Focal Loss but introduces a more refined strategy for identifying and suppressing hard-to-classify instances—thereby reducing overfitting to potentially mislabeled samples. The proposed losses dynamically down-weight gradient contributions from such instances in a self-adaptive manner, without requiring auxiliary modules or prior knowledge of noise rates. Extensive experiments on benchmark datasets with synthetically injected label noise demonstrate that the new losses significantly improve label noise detection accuracy, achieving average F1-score gains of +3.2–5.8 percentage points over standard cross-entropy and Focal Loss. The methods are both computationally lightweight and empirically effective, offering a simple yet powerful alternative for learning under label corruption.
Empirical Risk Minimization (ERM) suffers from degraded generalization under long-tailed class distributions. Method: This paper proposes a data-dependent shrinkage technique and establishes the first fine-grained, class-aware unified generalization upper bound. Unlike conventional coarse-grained analyses relying on global statistics, our bound explicitly quantifies how class-specific terms influence generalization error. Contribution/Results: The bound provides the first systematic theoretical explanation of the intrinsic mechanisms underlying reweighting and logit adjustment—resolving several counterintuitive empirical observations. Leveraging this theory, we design a principled learning algorithm that significantly improves minority-class accuracy on standard long-tailed benchmarks—including CIFAR-10-LT and ImageNet-LT—outperforming state-of-the-art methods.
Linear classifiers (e.g., SVM) suffer from degraded generalization performance on high-dimensional imbalanced data. Method: We establish a high-dimensional asymptotic theoretical framework and, for the first time, rigorously derive analytical expressions for the generalization error under undersampling and oversampling. Our approach integrates random matrix theory, high-dimensional statistical learning, and unsupervised probabilistic modeling–driven resampling. Contribution/Results: We quantify how resampling efficacy depends on the first- and second-order statistics of the data and the choice of evaluation metric. Crucially, we prove—and empirically verify—that hybrid sampling consistently outperforms either undersampling or oversampling alone. Extensive numerical experiments and evaluations on real-world datasets—including deep neural network features—demonstrate strong agreement between theoretical predictions and empirical results, with substantial improvements in minority-class classification accuracy. This work provides an interpretable, generalizable, and principle-based foundation for data rebalancing in high dimensions.
To bridge the gap between data-hungry deep models and human-like few-shot learning efficiency, this paper investigates the long-overlooked loss function component within meta-learning frameworks and proposes a novel dynamic adaptive loss learning paradigm. Methodologically, it introduces (1) EvoMAL—an interpretable symbolic loss evolution method that integrates symbolic regression with evolutionary algorithms to generate lightweight, task-adaptive, and interpretable loss functions; (2) Sparse Label Smoothing Regularization (SparseLSR), a new regularization technique for mitigating label noise in low-data regimes; and (3) NPBML—a unified framework jointly optimizing meta-initialization, meta-optimizer, and loss function. Experiments across multiple few-shot benchmarks demonstrate state-of-the-art performance: classification accuracy improves significantly, while memory overhead for loss learning is reduced by over 80%.
Conventional resampling methods for class-imbalanced classification suffer from inherent limitations—oversampling introduces noise and boundary ambiguity, while undersampling discards informative majority-class samples, leading to information loss and underfitting. Method: This paper proposes an intelligent majority-class sample selection mechanism guided by model loss improvement. Its core innovation is a novel gradient-driven, differentiable bilevel optimization framework: the upper-level objective maximizes generalization performance, while the lower-level optimizes a differentiable loss improvement metric, enabling end-to-end, deterministic undersampling. Contribution/Results: By directly selecting discriminative majority-class instances—without synthesizing noisy minority samples—the method preserves data fidelity and decision boundary clarity. Evaluated on multiple benchmark datasets, it achieves up to a 10% absolute improvement in F1-score over state-of-the-art methods, significantly enhancing minority-class detection while maintaining majority-class accuracy.
This work addresses the lack of a clear optimization objective in loss reweighting for long-tailed classification by formally casting it as an inverse problem. Guided by the equiangular tight frame (ETF) geometry from Neural Collapse (NC) theory, the study proposes equalizing the average per-class losses as an ideal target and dynamically infers class weights to approach this equilibrium. By modeling reweighting as an inverse problem and aligning feature geometry with the ETF structure, the method effectively reduces loss imbalance and encourages learned features to conform more closely to the geometric predictions of NC theory. Extensive experiments demonstrate consistent improvements over strong baseline methods across multiple long-tailed benchmarks.
This work addresses the challenge of accurately predicting deep, rare classes in hierarchical multi-label classification, where such categories suffer from intrinsic low frequency and further diminished prevalence due to hierarchical propagation. To this end, we propose a novel loss function that explicitly focuses on rare nodes—rather than rare samples—by integrating node-level class imbalance weighting with a focal weighting mechanism grounded in ensemble-based uncertainty quantification. This approach dynamically adjusts training emphasis based on model uncertainty and seamlessly integrates into mainstream neural architectures such as CNNs. Extensive experiments demonstrate substantial improvements, with recall gains up to fivefold and significantly higher F₁ scores compared to baseline methods. Notably, the proposed method maintains robust performance even under challenging conditions, including suboptimal encoders or scarce data regimes.
This work addresses the challenge of attributing misclassifications and evaluating robustness in black-box classifiers by proposing an explainability-aware optimization framework. The approach integrates L₀ sparsity regularization (XA-L₀) with a tolerance-region confusion matrix (TOR-Confusion Matrix) to generate minimal input perturbations that induce target predictions while preserving sparsity and semantic interpretability. This unified framework simultaneously enables the generation of highly interpretable counterfactual examples and fine-grained quantification of model robustness. Empirical evaluations on both image and tabular datasets demonstrate the method’s effectiveness, significantly outperforming existing black-box analysis techniques in terms of interpretability and robustness assessment fidelity.
This work addresses the limitations of conventional classification losses, which are prone to overfitting and exhibit poor robustness in scenarios involving ambiguous class boundaries or limited training samples. To this end, the authors propose a novel distributional loss function that models classification outputs as bimodal Gaussian distributions, thereby implicitly capturing class uncertainty and softening decision boundaries. This approach uniquely relaxes the classification target into a bimodal distribution without requiring additional annotations, necessitating only minimal modifications to standard training pipelines. End-to-end optimization is achieved through distribution matching. Extensive experiments demonstrate that the proposed method significantly enhances model robustness across multiple benchmark datasets, with particularly pronounced gains in low-data regimes.