stochastic gradient dynamics analysis

Constructs and analyzes stochastic differential equation (SDE) models of stochastic gradient trajectories to characterize drift, diffusion, transient and steady-state behavior of optimization dynamics. Uses these SDE analyses to diagnose failure modes, predict convergence and generalization properties, and inform optimizer design and hyperparameter choices.

stochasticgradientdynamicsanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.85
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Noise estimation of SDE from a single data trajectory

Sep 29, 2025
MA
Munawar Ali
🏛️ Florida State University | King's College London | University of Washington | Purdue University

This work addresses the challenge of modeling a single nonstationary, non-ergodic stochastic differential equation (SDE) trajectory—a setting where conventional SDE identification methods fail due to their reliance on ergodicity or stationarity assumptions. Method: We propose Stochastic Sparse Identification of SDEs (SSISDE), the first data-driven algorithm capable of jointly estimating drift and diffusion functions while reconstructing Brownian increments from a single trajectory. SSISDE integrates stochastic Taylor expansions with the Girsanov transformation, constructing solvable estimators initialized from drift function approximations—bypassing the need for stationary or ergodic data. Contribution/Results: SSISDE establishes the first SDE modeling paradigm tailored to single-trajectory, nonstationary, non-ergodic regimes. It achieves high-fidelity model discovery on benchmark nonstationary linear and quadratic systems—including Black–Scholes dynamics—significantly improving estimation accuracy over existing approaches. This framework enables real-time, interpretable modeling of complex dynamical systems in domains such as finance and biophysics, where ergodic assumptions are fundamentally violated.

Estimating noise from single SDE trajectoryIdentifying SDE dynamics using sparse stochastic algorithmRecovering Brownian motion without ergodicity assumptions

This work addresses the computationally expensive inverse problem of parameter estimation for stochastic differential equations (SDEs) by proposing an efficient solution framework that, for the first time, integrates Wiener chaos expansion (WCE) with stochastic gradient descent (SGD). By projecting the stochastic solution onto a deterministic system of propagators via an orthogonal Hermite polynomial basis, the method constructs a regularized discrepancy functional amenable to SGD optimization. This transformation effectively converts the original stochastic inverse problem into a deterministic optimization task, substantially reducing computational complexity and data requirements. Numerical experiments on several nonlinear SDE models—including a biological individual growth model—demonstrate that the approach accurately and robustly recovers parameters even from sparse and noisy observational data, exhibiting strong scalability and practical promise.

Inverse ProblemNoisy ObservationsParameter Estimation

A theoretical-practical gap persists in score-based diffusion models. Method: We propose a unified, reproducible SDE-based modeling framework that systematically integrates score matching, SDE/ODE solvers, denoising score estimation, and consistency modeling; notably, we introduce reinforcement learning into diffusion sampling for inference-path optimization. Contributions: (1) We establish theoretical consistency between sampling and score estimation under the SDE formulation; (2) we provide concise proofs of key theorems alongside practical algorithmic implementation guidelines; (3) we release modular, open-source code enabling rapid validation and extension to novel architectures. This work bridges the efficiency of score matching with scalable, RL-enhanced inference, delivering a foundational toolkit that balances theoretical rigor and engineering practicality for the design, analysis, and application of diffusion models.

Discusses sampling and score matching in diffusion modelingExplains score-based diffusion models using stochastic differential equationsProvides technical introduction for practitioners designing new models

Towards Identifiability of Interventional Stochastic Differential Equations

May 21, 2025
AZ
Aaron Zweig
🏛️ Columbia University | New York Genome Center

This work addresses the structural identifiability of parameters in stochastic differential equation (SDE) models under multiple interventions—i.e., whether SDE parameters can be uniquely recovered from samples of post-intervention stationary distributions. Theoretically, we establish the first uniqueness guarantee for SDE parameter recovery under multi-intervention settings; for linear SDEs, we derive a tight lower bound on the minimum number of required interventions; for weak-noise nonlinear SDEs, we obtain an upper bound on identifiability. Methodologically, we propose a parametric framework featuring learnable activation functions, integrating intervention modeling, stationary distribution analysis, and weak-noise asymptotic theory. Experiments on synthetic data demonstrate that our approach accurately recovers ground-truth parameters, and the theory-guided learnable architecture significantly improves both estimation accuracy and robustness.

Identifiability of SDE models under interventionsNecessary interventions for linear and nonlinear SDEsProvable bounds for unique SDE parameter recovery

Existing continuous-time SDE approximations of SGD fail to accurately characterize its escape dynamics from stationary points—especially local minima—exhibiting significant deviation on quadratic objectives. To address this, we propose the Hessian-Aware Stochastic Differential Equation (HA-SDE), the first SDE framework that jointly incorporates local Hessian information into both drift and diffusion terms, enabling precise modeling of SGD’s local dynamics near stationary points. Theoretically, HA-SDE exactly reproduces the SGD iterate distribution in the quadratic case—achieving zero-order optimal approximation error—and yields highly accurate escape probabilities and trajectory predictions. Moreover, it substantially weakens dependence on higher-order smoothness constants of the objective. This work establishes a tighter, geometrically informed continuous-time benchmark for analyzing SGD’s generalization mechanisms and optimization dynamics.

Characterize local SGD dynamics near stationary points preciselyModel SGD escaping behaviors accurately with Hessian-aware SDEReduce approximation error dependence on objective smoothness

Latest Papers

What's happening recently
View more

This work establishes a unified theoretical framework for diffusion models from the perspective of differential equations. Starting from a conditional Gaussian forward process, it derives the corresponding forward stochastic differential equation (SDE) and ordinary differential equation (ODE), and constructs a dynamical system that transports the data distribution to a standard Gaussian prior via marginalization. The framework then introduces a reverse SDE and a probability flow ODE, both driven by the marginal score function, thereby unifying score matching and noise prediction objectives. It rigorously demonstrates the equivalence of DDPM and DDIM in their training objectives while clarifying their fundamental distinction in sampling mechanisms—DDPM corresponds to a discretized reverse SDE, whereas DDIM implements a reverse ODE. Furthermore, the framework seamlessly incorporates mainstream sampling techniques such as DPM-Solver and classifier guidance, providing a coherent and rigorous continuous-time foundation for diffusion models.

differential equationsdiffusion modelsODE

This study addresses the estimation of the time-homogeneous drift function in multivariate stochastic differential equations (SDEs) with known diffusion coefficients, based on high-frequency observations from multiple trajectories. To this end, the authors propose formulating drift estimation as a conditional denoising problem conditioned on historical observations and introduce a conditional diffusion model that dynamically generates new trajectories from which the drift estimator is extracted. This approach represents the first application of conditional denoising diffusion models to drift estimation in SDEs. It significantly outperforms classical methods in high-dimensional settings without relying on any specific neural network architecture, while achieving comparable performance to existing approaches in low-dimensional cases, thereby demonstrating both its effectiveness and scalability.

denoising diffusion modelsdrift estimationhigh-frequency observations

This study addresses the challenge of estimating diffusion parameters in stochastic differential equation (SDE) models when data and model are compatible only at specific scales. The authors propose an adaptive subsampling method based on the statistics of monotonic runs. By demonstrating that, for a broad class of additive-noise SDEs, the length of monotonic runs at infinitesimal scales approximately follows a geometric distribution with success probability 1/2, they establish a general criterion for selecting the subsampling rate without relying on multiscale diffusion asymptotics. The optimal sampling scale matching the SDE’s infinitesimal behavior is automatically determined solely from the statistical properties of monotonic increasing or decreasing segments in the observed time series. Validation on surrogate modeling of fiber lay-down trajectories in nonwoven fabric production demonstrates that the method yields highly accurate and model-consistent diffusion parameter estimates, proving effective in real-world industrial applications.

data-model compatibilitydiffusion parameter estimationstochastic differential equations

This work investigates the optimization and generalization dynamics of stochastic gradient descent (SGD) in high-dimensional diagonal linear networks. By constructing a stochastic differential equation (SDE) to approximate SGD trajectories and deriving deterministic partial differential equations that govern the evolution of key statistical quantities such as risk and curvature, the study explicitly decouples the drift and gradient noise components of SGD for the first time in a high-dimensional setting. Building on this decomposition, the authors establish a globally well-posed non-asymptotic theoretical framework that guarantees exponential convergence to zero risk with high probability under appropriate parametrization. The theoretical predictions are corroborated by numerical experiments, demonstrating excellent agreement between analysis and empirical observation.

diagonal linear networksgeneralizationhigh-dimensional regime

This work proposes a unified variational generative modeling framework based on stochastic differential equations (SDEs) to efficiently address complex data generation tasks, including images, videos, and biomolecular structures. By incorporating both ordinary and stochastic differential equations, the authors derive the evidence lower bound (ELBO) from a variational inference perspective, systematically demonstrating that diffusion models, score matching, and flow matching are distinct parameterizations within this general framework. Through theoretical analysis grounded in the Fokker–Planck equation and empirical validation via one-dimensional density modeling experiments, the study provides clear comparisons among different parameterization strategies, confirming the proposed framework’s theoretical coherence, expressive capacity, and practical efficacy.

diffusion modelsgenerative machine learningscore matching

Hot Scholars

TS

Taiji Suzuki

The University of Tokyo
StatisticsMachine learning
DZ

Difan Zou

The University of Hong Kong
Machine LearningDeep LearningOptimizationStochastic Algorithms
DW

Denny Wu

New York University
Machine LearningStatistics
JW

Jingfeng Wu

University of California, Berkeley
deep learning theorymachine learningoptimizationstatistical learning theory
SI

Shawn Im

University of Wisconsin-Madison