Score
Designs, implements, and evaluates training objectives and algorithms that optimize minimax (worst-case) criteria to increase model robustness by explicitly accounting for adversarial or high‑loss perturbations. This includes constructing differential‑treatment loss functions or inner adversarial problems that treat classes or sample groups differently to enlarge separation between normal and near‑normal/abnormal examples and to reduce sensitivity to noisy normal samples.
Deep neural networks exhibit insufficient robustness against diverse perturbations—including ℓ₁, ℓ₂, and ℓ∞ adversarial noise as well as natural corruptions (e.g., adverse weather)—and existing adversarial training methods suffer from limited generalization across perturbation types. Method: This paper proposes the Mixture of Robust Experts (MoRE) framework, the first to formulate multi-perturbation robust learning as a mixture-of-experts mechanism. MoRE decouples robustness objectives along distinct ℓₚ-norm perturbation directions for joint optimization, incorporates dynamic gating for expert selection, enables robust feature sharing, and employs joint task training. Contribution/Results: Evaluated on CIFAR-10/100 and an ImageNet subset, MoRE significantly improves robust accuracy under mixed ℓₚ perturbations—achieving an average gain of +6.2% over unified adversarial training—while preserving clean-input accuracy. It overcomes the inflexibility of single-norm adversarial paradigms, enabling adaptive, cross-norm robustness without compromising standard performance.
This work investigates the root causes of adversarial examples and the mechanism by which adversarial training enhances model robustness. Standard training tends to learn dense yet non-robust features, compromising generalization under perturbations. Method: Under a structured data assumption, we propose a feature learning theoretical framework based on a two-layer smooth ReLU CNN; adversarial training (PGD-style) alternates gradient ascent to generate adversarial examples and gradient descent for optimization. Contribution/Results: We provide the first theoretical characterization of the learnability separation between robust and non-robust features, rigorously proving that PGD-based adversarial training promotes robust feature learning while suppressing non-robust feature learning—revealing an intrinsic link between feature robustness and adversarial perturbation directions. Our analysis combines feature disentanglement with generalization error bounds. Theoretically guaranteed robustness improvement is empirically validated on MNIST, CIFAR-10, and SVHN, confirming the predicted feature selection mechanism.
This work identifies function continuity—not an inherent trade-off—as the fundamental cause of the robustness-accuracy dilemma in machine learning. Theoretically, it provides the first explanation of adversarial examples through the lens of function continuity and uniform convergence. Methodologically, it introduces a novel paradigm of piecewise training with heterogeneous continuity assumptions, and constructs a learnable framework grounded in harmonic analysis and complex function theory. Experiments demonstrate that models adopting discontinuous hypotheses consistently outperform those assuming continuity, achieving superior robust accuracy (under adversarial attacks) and natural accuracy (on clean data). This work thus establishes a new theoretical foundation and practical pathway for designing classifiers that simultaneously attain high robustness and high accuracy.
This work addresses the fundamental trade-off between standard accuracy and adversarial robustness in supervised learning. Methodologically, it introduces the first architecture-level accuracy–robustness trade-off curve, quantifying the inverse relationship between these objectives across diverse neural network architectures; defines a sensitivity influence function to theoretically characterize the stability of optimal solutions under adversarial perturbations; and reveals—via theoretical analysis of overparameterized linear models—that adversarial training implicitly regularizes model dynamics, interpolating between L₁ (LASSO) and L₂ (ridge regression) behaviors. The approach integrates rigorous theoretical analysis, influence-function-based modeling, and extensive empirical evaluation across fully connected, deep, and varying-width networks. Results consistently validate the existence and structure of the trade-off, providing an interpretable, predictive theoretical foundation for principled neural architecture selection.
This work addresses the insufficient certified robustness of deep classifiers by proposing a novel framework that jointly optimizes the decision boundary margin in output space and the Lipschitz constant along vulnerable input directions. Methodologically, it introduces (1) a dynamic margin maximization mechanism that explicitly enlarges inter-class safety margins in the logit space; (2) a differentiable, tight upper bound estimator for the Lipschitz constant, incorporating both activation monotonicity and Lipschitz continuity constraints; and (3) a new differentiable activation layer designed specifically for robustness. Vulnerability-aware regularization enables end-to-end training. Extensive experiments on MNIST, CIFAR-10, and Tiny-ImageNet demonstrate substantial improvements in certified accuracy and generalization performance, consistently outperforming state-of-the-art certified robustness methods.
This work addresses the problem of deriving tight lower bounds on adversarial robustness for general multiclass loss functions—including cross-entropy, power-loss, and quadratic loss—overcoming the limitation of prior theories restricted to 0–1 loss. Methodologically, we propose a learner-agnostic dual characterization and barycentric reformulation of the robust risk minimization problem, establishing—for the first time—its intrinsic connection to α-fair binning and generalized barycenter problems. Leveraging an α-divergence framework regularized by KL divergence and Tsallis entropy, we integrate dual optimization with barycentric reconstruction to derive computationally tractable, tight lower bounds. Empirically, our approach significantly improves bound tightness under standard losses such as cross-entropy, yielding a more general and structurally insightful theoretical foundation for adversarial robustness in multiclass classification.
This work addresses the challenge of attributing misclassifications and evaluating robustness in black-box classifiers by proposing an explainability-aware optimization framework. The approach integrates L₀ sparsity regularization (XA-L₀) with a tolerance-region confusion matrix (TOR-Confusion Matrix) to generate minimal input perturbations that induce target predictions while preserving sparsity and semantic interpretability. This unified framework simultaneously enables the generation of highly interpretable counterfactual examples and fine-grained quantification of model robustness. Empirical evaluations on both image and tabular datasets demonstrate the method’s effectiveness, significantly outperforming existing black-box analysis techniques in terms of interpretability and robustness assessment fidelity.
This study addresses the computational intractability of directly evaluating probabilistic robustness and the inefficiency of existing adversarial example generation and defense mechanisms. Motivated by the intuition of distributional overlap, the authors derive a lower bound on the Kullback-Leibler divergence as a tractable surrogate objective and propose a novel probabilistic adversarial training algorithm. Furthermore, this work provides a probabilistic interpretation of conventional adversarial training, formally proving its equivalence to maximizing a lower bound on robustness. The proposed framework introduces a new theoretical perspective on adversarial robustness and significantly enhances the probabilistic robustness of deep learning models. Notably, the introduced scaling factor can be seamlessly integrated into existing non-probabilistic adversarial training methods to improve their performance without architectural modifications.
This work investigates the extent to which adversarial attacks reflect a model’s actual robustness under random noise of comparable magnitude, rather than merely characterizing worst-case scenarios. To this end, the authors propose a directional bias perturbation framework governed by a concentration parameter κ, which interpolates smoothly between isotropic noise and adversarial directions. They further introduce a novel attack strategy designed to better approximate realistic statistical noise. Through systematic evaluations on ImageNet and CIFAR-10, the study delineates the conditions under which common adversarial attacks effectively capture noise-induced failure risks, thereby offering both theoretical grounding and practical guidance for safety-oriented robustness evaluation of machine learning models.
This work addresses the instability in current adversarial robustness evaluations, which rely on fixed perturbation budgets and norm-specific attack ensembles, often failing to approximate worst-case performance. The authors propose a unified evaluation framework that constructs a minimal-norm attack pool spanning ℓ₀, ℓ₁, ℓ₂, and ℓ_∞ norms, transforming robustness assessment into the problem of approximating an attack frontier and a defense frontier under controllable query budgets via robustness–perturbation curves. They introduce the Defense Optimality Index—a novel metric that enables model ranking without predefining perturbation budgets—and devise an optimal attack subset selection strategy under query constraints. Experiments on CIFAR-10 and ImageNet demonstrate that this approach matches or surpasses AutoAttack across all query budgets, establishing a new paradigm for query-aware, curve-based robustness evaluation.