replica method

Applying statistical mechanics techniques (replica symmetry, disorder averaging) to analyze Bayes-optimal inference limits, characterize learning versus memorization phases, and study stability of fixed points in high-dimensional models.

replicamethod

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Variational Gaussian Approximation in Replica Analysis of Parametric Models

Sep 15, 2025
TT
Takashi Takahashi
🏛️ University of Tokyo | RIKEN center for AIP

This work addresses parameter inference and learning in parametric models when the data-generating distribution is unknown or intractable. We propose a novel replica-theoretic framework based on variational Gaussian approximation. Within the grand canonical ensemble, we defer data averaging and replace the conventional quenched average with an empirical average; crucially, the trial Hamiltonian parameters are determined adaptively via a variational principle, eliminating reliance on idealized distributional assumptions. Our key contribution is the explicit incorporation of fluctuation effects into the analysis—establishing, for the first time within the replica method, a rigorous correspondence with information criteria such as AIC and BIC, thereby quantifying how statistical fluctuations govern generalization performance. Theoretical analysis yields exact learning curves for linear regression, and empirical validation confirms the framework’s efficacy in scenarios where standard replica methods break down.

Analyzing inference in models with unknown data distributionsDeriving learning curves for linear regression applicationsUsing variational Gaussian approximation for replica analysis

High-dimensional learning of narrow neural networks

Sep 20, 2024
HC
Hugo Cui
🏛️ École Polytechnique Fédérale de Lausanne (EPFL) | Harvard University

This project investigates the effective learning mechanisms of narrow-width neural networks—such as MLPs, autoencoders, and attention modules—in high-dimensional big-data regimes. Addressing the theoretical gap in unifying generalization analysis across finite-width architectures under diverse learning paradigms—including supervised/unsupervised learning, denoising, and contrastive learning—we propose a **sequential multi-metric asymptotic model**, enabling the first systematic and unified characterization of narrow networks across heterogeneous tasks. Methodologically, we integrate statistical physics (replica method), approximate message passing (AMP), and high-dimensional asymptotic analysis to rigorously derive exact performance limits in the joint large-dimension–large-sample limit. Our framework yields the first analytically tractable and universally applicable theory for both generalization behavior and optimization dynamics of narrow-width networks, thereby filling a critical void in modern neural network theory concerning width-constrained settings.

Big DataHigh-dimensional DataNeural Networks

This study addresses the degradation of posterior quality and predictive performance in practical Bayesian inference caused by data noise, high-dimensional strongly correlated parameters, and anomalous model outputs. To overcome these challenges, the authors embed Bayes’ theorem within a statistical mechanics framework and introduce a tempered posterior distribution governed by a temperature parameter τ. They propose, for the first time, the use of Wang–Landau sampling to estimate the posterior density of states. A single simulation run automatically identifies phase-transition-like signals to determine the optimal temperature, thereby eliminating the need for manual tuning or repeated computations. The method demonstrates remarkable efficacy in equation-of-state modeling in materials science, successfully handling high-dimensional correlated parameters and noisy, chaotic data while substantially improving inference accuracy and predictive capability.

Bayesian inferencematerials sciencephase transitions

This work addresses probabilistic inference for Ising models with hidden Markov structure in machine learning, focusing on the characterization and computation of translation-invariant Gibbs measures on Cayley trees. We propose a latent-variable Hamiltonian that jointly incorporates Ising interactions and data-observation coupling terms, yielding a generative framework amenable to hierarchical data modeling. Theoretically, we establish—rigorously and for the first time on Cayley trees—that at most three translation-invariant Gibbs measures exist for this model, providing a solid phase-transition foundation for multimodal inference. Algorithmically, we design an exact inference algorithm with provable convergence guarantees. Experiments demonstrate that our method significantly outperforms baselines on image denoising, weakly supervised learning, and anomaly detection tasks; moreover, all theoretical conditions are explicitly verifiable in practice.

Apply measures to denoising, weak learning, anomaly detectionExplore translation-invariant Gibbs measures on Cayley treesStudy Ising models with hidden Markov structure for inference

This work investigates the high-dimensional geometric structure of the solution space achieving zero training error in neural networks—modeled by binary-weight perceptrons—and its evolution with training set size. Using statistical physics methods—including the replica method, Gardner capacity analysis, and geometric characterization of solution spaces—we uncover a phase transition from dense, clustered solutions to sparse, isolated ones. We introduce “linear modal connectivity” as a quantitative measure of the average shape of solution manifolds. Crucially, we identify that algorithmic hardness arises from the disappearance of distant solution clusters precisely at the critical data threshold. Our analysis quantitatively characterizes the SAT/UNSAT phase transition, scaling laws of solution cluster sizes, and local landscape ruggedness. Collectively, these results establish a unified geometric–statistical physical framework for understanding generalization and optimization difficulty in deep learning.

Analyze solution manifold in neural networksExplore SAT/UNSAT transition in storageStudy geometric arrangement of zero-error configurations

Latest Papers

What's happening recently
View more

This paper addresses the challenge of reliably detecting symmetries in chaotic attractors under high noise and limited data. We propose a statistical framework integrating optimal transport and Bayesian inference: symmetry groups are selected via a Gibbs posterior constructed using the Wasserstein distance, inherently enforcing Occam’s razor, conjugate equivariance, and robustness to strong noise. Coupling Metropolis–Hastings sampling with group-action transformations enables probabilistic identification of symmetry structures. The method accurately recovers ground-truth symmetries even under severe noise and sparse observations. We validate it on human gait dynamics, revealing noise-robust, mechanically constrained symmetry evolution—uncovering how biomechanical constraints shape dynamical symmetry over time. This provides an interpretable, quantifiable paradigm for model reduction and mechanistic analysis of nonlinear dynamical systems.

Detecting symmetry in chaotic attractors from observed data trajectoriesIdentifying hierarchical symmetry structures in dynamical systemsQuantifying uncertainty in symmetry detection for noisy datasets

Traditional empirical risk minimization focuses solely on a single optimal solution, failing to capture the multistability and uncertainty inherent in data-driven learning. This work reframes the empirical loss function as an interaction potential and constructs an energy-based model grounded in Gibbs measures on a Cayley tree, thereby establishing—for the first time—a rigorous connection between loss landscapes and probabilistic inference on tree-structured graphs. By leveraging nonlinear integral fixed-point equations, data-dependent kernels inducing compact operators, and phase transition analysis, the study theoretically proves existence and uniqueness of solutions in the one-dimensional setting. Numerical experiments further demonstrate the coexistence of multiple solution branches under non-separable kernels, revealing that data can induce multiple learning states and associated phase transitions.

energy-based learningGibbs measureshierarchical structures

This work proposes a universal framework for identifying phase transitions without requiring order parameters or prior knowledge of the underlying model. Building on the hypothesis that infinitesimal parameter perturbations break statistical indistinguishability in the thermodynamic limit, the study redefines phase transitions as abrupt changes in distributional distinguishability. This perspective enables a model-agnostic, training-free detection method implemented via a distribution-free two-sample runs test. The approach unifies conventional criteria—including the Binder cumulant—and accurately locates the critical point in the two-dimensional Ising model, thereby demonstrating both its validity and broad applicability across different systems.

hypothesis testingorder-parameter-freephase transitions

This work aims to construct a conceptual bridge between statistical physics and deep learning for researchers without a physics background. By recasting statistical physics as a natural extension of probability theory, it systematically elucidates the Boltzmann–Gibbs distribution, Ising models, spin glasses, and phase transition theory, while uncovering their intrinsic connections to Hopfield networks and restricted Boltzmann machines (RBMs). The key insight lies in demonstrating the equivalence between integrating out hidden units in RBMs and the renormalization group transformation, thereby revealing a physically grounded mechanism underlying multilayer deep networks. This framework not only deepens the theoretical understanding of deep learning but also provides a clear physical interpretation of the developmental trajectory of large language models.

deep learningneural networksphase transitions

Hot Scholars

SS

Susmit Shannigrahi

Assistant Professor at Tennessee Tech University
Internet ProtocolsFuture Internet Architectures5G networksBig Data
DE

Daniel E. Lucani

Professor at the Department of Electronic and Computer Engineering, Aarhus University
Network Codingdata compressionDistributed StorageIoT
MG

Moncef Gabbouj

Professor, Tampere University
Machine learningArtificial intelligenceSignal processingimage processing
LQ

Lili Qiu

NAI Fellow, ACM Fellow, IEEE Fellow, Professor, Dept. of Computer Science, The University of Texas
Wireless NetworksWireless SensingMobile ComputingSystems