🤖 AI Summary
This study addresses the challenge of data imbalance, where optimizers fail to distinguish gradient contributions from majority and minority classes, causing training to stagnate in majority-loss-dominated regions. By modeling small-step-size training dynamics via continuous-time dynamical systems, we characterize the geometric structure of these dominant regions. Combining theoretical analysis of optimizer geometry with experiments on AdamW and Muon, we investigate the escape mechanisms of different optimizers. Theoretically, we demonstrate that Sign, Spectral, and Newton descent exhibit weaker dependence on minority-class magnitudes compared to Euclidean gradient descent, thereby overcoming conventional optimization limitations. Empirically, across language, tabular, and vision tasks, we confirm that non-Euclidean geometric optimizers significantly outperform SGD, effectively mitigating learning biases induced by class imbalance.
📝 Abstract
Data imbalance is pervasive in machine learning, from rare words and anomalies to underrepresented patterns in heterogeneous or cross-tabulated data. We study idealized optimizers geometries in continuous time to model small-step training in deep learning. We assume that the source of imbalance is unobserved: the optimizer has only access to the aggregate training loss ignoring the exact contributions of the majority and minority groups. In this setting, we characterize a region where majority losses are optimized regardless of the admissible minority structure. We derive explicit equations of this zone and bounds on the time needed to leave it. These bounds exhibit a milder dependence on minority amplitude for sign, spectral, and Newton descent than for Euclidean gradient descent. Experiments with AdamW and Muon suggest similar advantages over SGD across language, tabular, and image tasks.