Robustness Emerges Early in Training Dynamics, but Is Not Preserved

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the degradation of robustness in deep neural networks during standard training, where although robust representations and flat loss landscapes emerge spontaneously in early epochs, these favorable properties deteriorate as training progresses, leading to insufficient robustness against natural perturbations. To counteract this issue, the paper proposes a training-dynamic intervention framework that requires no architectural modifications or additional parameters. The approach introduces two novel strategies—Early Phase Stabilization (EPS) and Asymmetric Weight Rollback (AWR)—to systematically preserve or restore the robust priors established in early training stages. Extensive experiments demonstrate that this method consistently enhances model robustness, transferability, and dynamic adaptability across diverse benchmarks and vision tasks.
📝 Abstract
Robustness to natural corruptions remains a fundamental challenge for deep neural networks. In this paper, we identify a robustness fading phenomenon where shallow layers spontaneously develop robust representations and flat loss landscapes in early training, yet these properties are not preserved during standard convergence. To address this, we propose a framework that performs strategic interventions on training dynamics to stabilize the empirically identified early-emergent robust priors. Our approach includes two parameter-free strategies: Early-Phase Stabilization~(EPS) and Asymmetric Weight Reversion~(AWR), which stabilize or recover robust shallow configurations without modifying the model architecture or introducing learnable parameters. Extensive experiments demonstrate the efficacy of our framework across various benchmarks and architectures, yielding significant gains in downstream transfer, dynamic adaptation, and diverse computer vision applications.
Problem

Research questions and friction points this paper is trying to address.

robustness
natural corruptions
training dynamics
loss landscape
deep neural networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

robustness fading
early-phase stabilization
asymmetric weight reversion
training dynamics
parameter-free intervention
🔎 Similar Papers
No similar papers found.
J
Jiangang Yang
Institute of Microelectronics, Chinese Academy of Sciences, Beijing, China
W
Wenhui Shi
Institute of Microelectronics, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
L
Lu Hu
Institute of Microelectronics, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Jing Xing
Jing Xing
Lingang Laboratory
Drug Discovery Data Mining
Jian Liu
Jian Liu
Beijing National Laboratory for Molecular Sciences, Peking University
Nonadiabatic FieldPhase Space Formulations of QMPath Integral QMQuantum Molecular Dynamics