Score
Designs and applies analyses, visualizations, and metrics to detect, characterize, and explain recurring patterns and behaviors in objective-function values and landscapes during learning (e.g., plateaus, oscillations, sharp vs. flat minima, saddle points, gradient noise). Builds diagnostic procedures and controlled experiments to isolate root causes of loss behavior and to evaluate how optimization settings, model components, data issues, or regularization affect convergence and generalization.
This work addresses the limited interpretability of internal mechanisms in reinforcement learning—such as value estimation, policy optimization, and their interaction with temporal difference (TD) signals—by proposing the first systematic, multi-perspective framework for visualizing loss landscapes. The approach integrates the geometric structure of value functions, policy optimization trajectories, TD error dynamics, and state-driven regions through techniques including 3D loss surface reconstruction, policy landscape visualization under a frozen critic, joint trajectories of time–Bellman error–policy weights, and state–TD mappings. Applied to the ADHDP algorithm for spacecraft attitude control, the framework enables comparative analysis of multiple variants, revealing how training stabilizers and target update mechanisms reshape the optimization landscape and influence learning stability. This study establishes a new paradigm for interpretable reinforcement learning and provides actionable insights for algorithm design.
To address the limited generalization performance caused by rigid loss landscape topography, this paper proposes a dynamic loss function that introduces a time-varying periodic oscillation mechanism atop standard losses (e.g., cross-entropy or mean squared error), dynamically modulating per-class loss weights while preserving the global optimum. This is the first work to incorporate oscillatory control into supervised learning loss design, uncovering an intrinsic link between margin instability and optimization trajectory. Through loss surface evolution analysis and margin sensitivity probing, we empirically demonstrate that the dynamic loss guides optimizers toward wider, shallower minima with superior generalization. Experiments across multi-scale architectures yield significant improvements in validation accuracy, validating both the effectiveness and broad applicability of active loss landscape modulation.
This study addresses the high computational cost and poor safety controllability in training ultra-large AI models. Methodologically, it introduces a novel optimization paradigm grounded in mechanistic interpretability, pioneering the application of “circuit analysis” to model gradient descent trajectories. It structures the parameter space into functional subnetworks and designs a progressive curriculum learning strategy to dynamically regulate optimization paths within a controlled environment. Key contributions include: (1) establishing a formal mapping between gradient flow dynamics and circuit-like structural representations, enabling interpretable modeling of optimization behavior; and (2) leveraging structural priors to guide curriculum design, significantly accelerating convergence while suppressing the emergence of harmful behaviors. Experiments across multiple benchmark tasks demonstrate over 30% reduction in training cost alongside improved behavioral controllability, offering a principled pathway toward efficient and safe large-model training.
This work addresses the limitations of conventional Gaussian-noise perturbations by systematically uncovering previously overlooked one-dimensional (1D) and two-dimensional (2D) local geometric structures in deep neural network (DNN) loss landscapes. Methodologically, we introduce a progressive taxonomy of five types of 1D loss curves (e.g., *v-basin*, *vvv-basin*) and design a perturbation-direction mining algorithm that integrates low-dimensional subspace projection with Hessian spectral analysis to automatically extract and visualize complex geometric structures. Our contributions include: (i) the first empirical observation and visualization of canonical 2D loss structures—including saddle surfaces and “bottle-bottom” geometries—in real DNNs; (ii) a theoretical characterization linking the geometric properties of perturbation directions to the eigenvalue distribution of the Hessian; and (iii) a novel geometric perspective for understanding generalization behavior and optimization dynamics.
Current post-training of language models relies on abstract scalar rewards, which lack transparency regarding the instructional content of preference data and can lead models to learn spurious correlations, resulting in undesirable behaviors such as excessive stylization or sycophancy. This work proposes a data-centric post-training framework that, for the first time, leverages interpretability methods to explicitly model latent conceptual signals within preference data. By analyzing and identifying key features that distinguish preferred from non-preferred responses prior to optimization, the approach integrates interpretability protocols, statistical hypothesis testing, and fine-grained interventions at both feature and data levels. This enables effective diagnosis and suppression of harmful learning signals, significantly reducing off-target behaviors across multiple benchmarks while enhancing model safety and controllable personality.
This work addresses the challenge of effectively predicting out-of-distribution (OOD) generalization failure prior to model deployment. It proposes a biologically inspired, top-down approach that leverages the geometric structure of model representations on in-distribution (ID) data, introducing manifold dimensionality and task utility as system-level predictive indicators for the first time. Unlike methods relying on mechanistic interpretability, this framework provides high-level, generalizable diagnostic signals without requiring detailed knowledge of internal model mechanisms. Empirical results demonstrate that, in image classification tasks, geometric properties of ID data manifolds alone can reliably forecast OOD performance. Notably, on ImageNet transfer learning benchmarks, the proposed predictors significantly outperform conventional ID accuracy and exhibit strong generalization across diverse architectures and datasets.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.
This work addresses the instability commonly observed in online reinforcement learning for dynamic systems, which often stems from the opaque optimization dynamics of the critic network. To this end, the authors propose a critic-matching loss landscape visualization method that projects the critic’s parameter trajectory onto a low-dimensional linear subspace, enabling the construction of a three-dimensional loss surface and a two-dimensional optimization path. The study introduces, for the first time, the concept of a critic-matching loss landscape along with quantitative metrics, and integrates these with a normalized system performance index to enable joint qualitative and quantitative analysis of the training process. Experiments on inverted pendulum and spacecraft attitude control tasks demonstrate that the method effectively reveals distinct loss landscape characteristics associated with stable convergence versus unstable learning, offering a novel tool for understanding and diagnosing online reinforcement learning behavior.
This work investigates the mechanism underlying improved generalization of neural networks trained with large learning rates near the edge of stability. By modeling stochastic optimizers as random dynamical systems, the study reveals that optimization trajectories converge to fractal attractors of low intrinsic dimensionality. Building on the Lyapunov dimension, the authors introduce a novel metric—“sharpness dimension”—to characterize generalization performance. This measure uniquely incorporates the full spectral structure of the Hessian and its partial determinants, overcoming limitations of conventional sharpness measures that rely solely on the trace or spectral norm. The theoretical framework is validated across diverse architectures, including MLPs and Transformers, and offers a new perspective on the “grokking” phenomenon.