Sparse Weak-Form Discovery of Stochastic Generators

📅 2026-03-21
📈 Citations: 0
Influential: 0
📄 PDF

career value

202K/year
🤖 AI Summary
This work proposes a weak-form sparse regression framework based on spatial Gaussian test functions to address structural bias introduced by conventional time-domain test functions in the identification of stochastic differential equations (SDEs). By unifying the estimation of drift and diffusion terms into two sparse linear systems sharing a common design matrix, the method uniquely integrates the weak formulation of Weak SINDy with the objective of stochastic SINDy. The use of spatial Gaussian kernels ensures zero conditional mean under noise, thereby eliminating regression bias at its source. Combined with ℓ¹ regularization, grouped cross-validation, and a two-step bias correction scheme, the approach effectively handles state-dependent diffusion. Validated on Ornstein–Uhlenbeck, double-well Langevin, and multiplicative noise systems, it accurately recovers all active generators (coefficient errors < 4%), achieves total variation distances below 0.01 for stationary densities, and precisely reproduces true relaxation timescales in autocorrelation functions.

Technology Category

Application Category

📝 Abstract
We introduce a framework for the data-driven discovery of stochastic differential equations (SDEs) that unifies, for the first time, the weak-form integration-by-parts approach of Weak SINDy with the stochastic system identification goal of stochastic SINDy. The central novelty is the adoption of spatial Gaussian test functions $K_j(x)=\exp(-|x-x_j|^2/2h^2)$ in place of temporal test functions. Because the kernel weight $K_j(X_{t_n})$ is $\mathcal{F}_{t_n}$-measurable and the Brownian innovation $ξ_n$ is independent of $\mathcal{F}_{t_n}$, every noise term in the projected response has zero conditional mean given the current state -- a property that guarantees unbiasedness in expectation and prevents the structural regression bias that afflicts temporal test functions in the stochastic setting. This design choice converts the SDE identification problem into two sparse linear systems -- one for the drift $b(x)$ and one for the diffusion tensor $a(x)$ -- that share a single design matrix and are solved jointly via $\ell_1$-regularised regression with grouped cross-validation. A two-step bias-correction procedure handles state-dependent diffusion. Validated on the Ornstein--Uhlenbeck process, the double-well Langevin system, and a multiplicative diffusion process, the method recovers all active polynomial generators with coefficient errors below 4\%, stationary-density total-variation distances below 0.01, and autocorrelation functions that faithfully reproduce true relaxation timescales across all three benchmarks.
Problem

Research questions and friction points this paper is trying to address.

stochastic differential equations
system identification
weak-form
regression bias
diffusion processes
Innovation

Methods, ideas, or system contributions that make the work stand out.

weak-form SDE discovery
spatial Gaussian test functions
stochastic SINDy
unbiased system identification
sparse regression
🔎 Similar Papers