energy landscape analysis

Analyzing optimization landscapes and dynamical behaviour of models to identify modes, low-energy states, and phase transitions in training or circuit performance. Used to study mode collapse under guidance, how circuit connectivity affects reaching low-energy configurations, and weight-space phenomena like glassy memorization or grokking.

energylandscapeanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Understanding Mode Connectivity via Parameter Space Symmetry

May 29, 2025
BZ
Bo Zhao
🏛️ University of California, San Diego | IBM Research | Northeastern University

This paper investigates the ubiquitous “mode connectivity” phenomenon in neural network training—namely, the existence of low-loss paths connecting distinct global minima. To elucidate its theoretical origin, the work establishes, for the first time, a rigorous link between parameter-space symmetries—particularly continuous symmetry groups—and the topological connectivity of the loss landscape: it proves that symmetry-induced Lie group actions determine the connected components of minima; derives a strict theorem showing how “jump connections” reduce the number of such components; and obtains explicit symmetry-derived interpolation paths together with a curvature-based criterion for linear mode connectivity. Methodologically, the approach integrates Lie group theory, algebraic topology, and geometric modeling of loss surfaces. It yields exact characterization of the number of minimum-connected components in linear networks and proposes a computationally tractable path-construction algorithm—thereby providing a novel theoretical foundation for understanding generalization and model merging in deep learning.

Derives conditions for linear mode connectivityExamines when mode connectivity holds or failsExplores mode connectivity in neural networks via symmetry

Understanding Machine Unlearning Through the Lens of Mode Connectivity

Apr 08, 2025
JC
Jiali Cheng
🏛️ University of Massachusetts Lowell

Machine unlearning lacks a geometric understanding of how model parameter spaces evolve during forgetting. Method: This work pioneers the application of mode connectivity to machine unlearning, systematically conducting parameter interpolation, loss surface visualization, and multi-stage metric tracking across diverse unlearning methods, with and without curriculum learning, and under first-order (SGD/Adam) and second-order (K-FAC) optimizers. Contribution/Results: We discover pervasive smooth, low-loss connectivity paths between pre- and post-unlearning models; path connectivity strength correlates strongly with unlearning efficacy. We characterize distinctive oscillatory patterns of evaluation metrics along these paths, uncovering both mechanistic similarities and differences across unlearning strategies. Crucially, mode connectivity emerges as a novel, interpretable paradigm for evaluating and explaining unlearning behavior—validated across multiple benchmark datasets.

Compares mechanistic differences between unlearning methodsExplores loss landscapes in diverse unlearning scenariosInvestigates machine unlearning via mode connectivity analysis

Phase transitions reveal hierarchical structure in deep neural networks

Dec 05, 2025
IE
I. Ersoy
🏛️ University of Potsdam | Universidade Federal Fluminense

Apparent phenomena in deep neural network training—such as phase transitions, ubiquitous saddle points, and model connectivity—are often studied in isolation, yet they fundamentally arise from the intrinsic geometric structure of loss and error landscapes. Method: We propose a unified geometric framework that analytically links phase transitions to saddle-point distributions; design an L₂-regularized error landscape probing algorithm to uncover the geometric mechanism underlying inter-class transfer; and combine theoretical analysis with MNIST experiments, employing path optimization to connect global minima. Contribution/Results: Our approach empirically validates mode connectivity and reveals a hierarchical precision basin structure stratified by digit classes. The framework provides a novel geometric perspective on deep learning optimization dynamics and introduces computationally tractable tools for landscape analysis—advancing both theoretical understanding and practical optimization strategies.

Reveals hierarchical accuracy basins analogous to phases in physicsShows saddle points govern phase transitions in neural network learningUnifies phase transitions, saddle points, and mode connectivity in DNN training

This work investigates the high-dimensional geometric structure of the solution space achieving zero training error in neural networks—modeled by binary-weight perceptrons—and its evolution with training set size. Using statistical physics methods—including the replica method, Gardner capacity analysis, and geometric characterization of solution spaces—we uncover a phase transition from dense, clustered solutions to sparse, isolated ones. We introduce “linear modal connectivity” as a quantitative measure of the average shape of solution manifolds. Crucially, we identify that algorithmic hardness arises from the disappearance of distant solution clusters precisely at the critical data threshold. Our analysis quantitatively characterizes the SAT/UNSAT phase transition, scaling laws of solution cluster sizes, and local landscape ruggedness. Collectively, these results establish a unified geometric–statistical physical framework for understanding generalization and optimization difficulty in deep learning.

Analyze solution manifold in neural networksExplore SAT/UNSAT transition in storageStudy geometric arrangement of zero-error configurations

Cascade of phase transitions in the training of Energy-based models

May 23, 2024
DB
Dimitrios Bachtis
🏛️ ENS | Sorbonne Université | Université Paris Cité | Universidad Complutense | Université Paris-Saclay

This work investigates the dynamic evolution of feature encoding during Restricted Boltzmann Machine (RBM) training. Methodologically, it combines singular value decomposition to track weight matrix evolution, analytical mean-field analysis, numerical training experiments, and finite-size scaling theory. The key contribution is the first identification—within energy-based generative models (EBMs)—of a mean-field–like paramagnetic–ferromagnetic cascade phase transition sequence during training. Empirically, the transitions sharpen in the high-dimensional limit; the first-order transition belongs to the mean-field universality class; and learning proceeds sequentially: the model first captures the centroid of the empirical data distribution, then progressively extracts principal components. This study establishes a rigorous correspondence between high-dimensional learning dynamics and statistical-physics phase-transition theory, offering a novel paradigm for understanding the intrinsic learning mechanisms of deep generative models.

Investigates phase transitions in RBM trainingTracks weight matrix evolution via SVDValidates theory with real dataset experiments

Latest Papers

What's happening recently
View more

This work proposes a brain-inspired neural computing framework designed to unify learning, memory, control, and optimization within a single architecture that is scalable, robust, and energy-efficient. By integrating principles from energy landscapes, gradient flows, control theory, and neuroscience, the study introduces a novel paradigm that transcends conventional feedforward networks and backpropagation. Key mechanisms include continuous-time Hopfield networks, dense associative memory, oscillator-based dynamics, and proximal descent dynamics. The resulting architecture achieves markedly improved computational efficiency and biological plausibility, demonstrating superior performance in data-driven control, constrained reconstruction, and large-scale optimization tasks.

dynamical systemsenergy efficiencyneurocomputation

This work addresses the limitations of conventional synchrony stability analysis, which relies on scalar metrics and fails to capture node-level dynamic stability behaviors. To overcome this, the study introduces the concept of a “dynamic stability landscape” as an upstream learning task and pioneers a graph-to-image prediction paradigm that directly maps network topology to image-like stability representations for each node. The proposed method integrates graph neural networks (GNNs) to encode topological structure and convolutional neural networks (CNNs) to decode and generate stability images, enabling end-to-end training. Evaluated on two newly constructed datasets comprising tens of thousands of graphs, the model achieves high in-distribution accuracy and demonstrates strong generalization across varying graph scales and real-world power grid topologies, offering a richer, fine-grained tool for stability analysis across multiple domains.

dynamic stabilitygraph topologyscalar stability indices

This work challenges the reliability of static mechanistic localization in guiding parameter updates during post-training of large language models. By tracking the structural evolution of Transformer circuits throughout supervised fine-tuning, the study introduces three novel metrics—circuit distance, stability, and conflict—to reveal a “free evolution” phenomenon wherein mechanisms shift unpredictably over time, exposing the temporal lag inherent in static localization approaches. The analysis deconstructs the illusion of effectiveness in current methods, demonstrating empirically that static mechanisms fail to anticipate future model states. Consequently, the paper advocates for a forward-looking, dynamic mechanistic localization framework to better support interpretability-guided, efficient post-training strategies.

circuit evolutionmechanistic interpretabilityparameter updates

This work investigates the phenomenon of memory in generative model training—where models persistently output similar samples—and elucidates its underlying mechanism through the lens of dynamical systems theory. By integrating the two-timescale dynamics of stochastic gradient descent (SGD) with structural properties of the loss landscape, the study offers the first unified explanation linking memory effects, double descent, and mode collapse, emphasizing the pivotal role of training dynamics themselves. Building upon Austin’s (2016) loss modeling, Borkar’s (2025a, 2026) theories of collapse and double descent, and recent advances by Azizian et al. (2024) on constant-stepsize SGD, the authors construct a dynamical framework that reveals the fundamental causes of output stagnation during training.

dynamical systemsgenerative modelsmemorization phenomenon

This work proposes a sparse and unified theoretical framework to systematically uncover the core mechanisms underlying learning, optimization, and modeling. It conceptualizes learning as a multi-level process arising from the coupling of problem formulation, method selection, and optimization dynamics. By precisely defining “solvable problems” and “parameterized methods,” the framework reduces complex learning theory to a few fundamental concepts rooted in dynamical systems, differential geometry, and foundational physics. The approach yields a general convergence theorem and establishes a universal theoretical foundation for cross-domain modeling and algorithm design, substantially enhancing both the parsimony and explanatory power of learning theory.

learningoptimizationparametrised methods

Hot Scholars

SM

Sabita Maharjan

Professor
Wireless networksvehicular networksnetwork securityenergy informatics
SZ

Shiliang Zhang

Department of Computer Science, School of EECS, Peking University
Multimedia Information RetrievalMultimedia SystemsVisual Search
TB

Tom Brown

Oxford University Chemistry Department
Nucleic acid chemistrybiological and nanotechnology applications
CB

Christoph Becker

Professor of Information, University of Toronto
responsible computingsustainable computinghuman computer interactionpost-growth