ergodic stochastic control

Designs and analyzes long-run (ergodic) optimal control policies for stochastic dynamical systems by formulating and solving ergodic Hamilton–Jacobi–Bellman (HJB) equations and associated linear Poisson (corrector) equations. Uses stochastic-process conditioning to characterize stationary laws and implements numerical solvers (e.g., finite-difference PDE methods) to compute stationary policies and long-run cost or value rates.

ergodicstochasticcontrol

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the computational challenges in non-Markovian stochastic optimal control, where the value function is governed by a stochastic Hamilton–Jacobi–Bellman (SHJB) equation whose measurability-induced randomness impedes tractable solution. Under the assumption of control-independent stochastic integral coefficients, the paper introduces the first policy iteration framework tailored to semilinear SHJB equations. The approach iteratively linearizes the original problem into a sequence of linear equations, which are then solved using tools from stochastic analysis and functional approximation. Theoretical analysis establishes that the resulting sequence of approximations converges monotonically in the mean-square sense and exhibits exponential convergence rates, thereby significantly enhancing both the feasibility and efficiency of computing value functions in non-Markovian settings.

computational challengesHamilton-Jacobi-Bellman equationnon-Markovian

This work addresses the numerical solution of high-dimensional Hamilton–Jacobi–Bellman (HJB) equations arising in stochastic optimal control. Methodologically, we propose a novel neural network-based actor–critic algorithm: the critic network is designed with an architecture that automatically satisfies boundary conditions, and a bias-gradient technique is introduced to reduce computational cost; the actor update minimizes the integral of the Hamiltonian, circumventing explicit differentiation. Theoretically, we prove that, in the infinite-width limit, the training dynamics converge to an infinite-dimensional ordinary differential equation whose fixed point exactly coincides with the true HJB solution. Experimentally, the method achieves high accuracy and robustness on problems up to 200 dimensions—including cases with non-convex Hamiltonians and linear-quadratic regulators—significantly extending the dimensionality frontier for deep learning–based HJB solvers.

Analyzing actor-critic methods for solving high-dimensional HJB equations.Demonstrating algorithm performance in stochastic control up to 200 dimensions.Ensuring boundary conditions and reducing computational costs in critic architecture.

A Temporal Difference Method for Stochastic Continuous Dynamics

May 21, 2025
HS
Haruki Settai
🏛️ University of Tokyo

Existing Hamilton–Jacobi–Bellman (HJB)-driven reinforcement learning methods require full knowledge of system dynamics, rendering them inapplicable to model-free stochastic continuous-time systems. This paper proposes the first fully model-free, HJB-guided temporal-difference (TD) framework that directly approximates the HJB partial differential equation on stochastic differential equation (SDE) systems without accessing drift/diffusion coefficients. Leveraging Itô’s lemma and stochastic calculus, our method establishes a model-free gradient update rule and provides rigorous convergence guarantees. It unifies stochastic optimal control with model-free RL, overcoming the long-standing requirement in HJB-RL for exact dynamical knowledge. Experiments demonstrate substantial improvements in policy performance and sample efficiency across multiple continuous-control benchmarks, with greater stability and computational efficiency compared to transition-kernel-based approaches.

Bridging stochastic optimal control and model-free RLModel-free RL for continuous dynamics without known coefficientsTemporal difference method targeting HJB equation

Neural optimal controller for stochastic systems via pathwise HJB operator

Feb 23, 2024
ZJ
Zhe Jiao
🏛️ Northwestern Polytechnical University | Massachusetts Institute of Technology

High-dimensional stochastic optimal control remains challenging due to the curse of dimensionality and limitations of conventional approaches relying on probabilistic representations of the Hamilton–Jacobi–Bellman (HJB) equation. Method: This paper proposes a physics-informed deep learning framework grounded in a pathwise HJB operator, unifying modeling and solution. It introduces the pathwise HJB operator as a novel physical constraint and designs two tailored numerical schemes—accommodating both explicit and implicit optimal control structures—integrated with PINNs, dynamic programming, SDE discretization, and HJB path-integral representations. Contribution/Results: A unified theoretical analysis quantifies truncation, approximation, and optimization errors. The framework significantly improves control accuracy and generalization across diverse high-dimensional tasks while ensuring interpretability and analytical tractability, establishing a new paradigm for real-time optimal control of complex stochastic systems.

Developing efficient deep learning control methodsSolving high-dimensional stochastic control problemsUsing physics-informed neural networks for HJB equations

This work addresses the challenge of reward fine-tuning in diffusion models and Boltzmann distribution sampling by formulating generative model optimization as a stochastic optimal control problem governed by stochastic differential equations. Leveraging the Stochastic Maximum Principle (SMP), the paper rigorously derives, for the first time, a general Hamiltonian adjoint matching objective applicable to settings where both drift and diffusion coefficients depend on the control, and establishes its intrinsic connection to the Hamilton–Jacobi–Bellman (HJB) equation. By integrating the adjoint system with a continuous-time successive approximation algorithm, the method recovers the lightweight adjoint loss when the diffusion coefficient is state-independent, validates the necessity of higher-order terms in state-dependent cases, and provides a tractable iterative scheme based on SMP that circumvents intractable martingale terms.

adjoint matchingHamiltonianstate-dependent diffusion

Latest Papers

What's happening recently
View more

Implicit Numerical Scheme for the Hamilton-Jacobi-Bellman Quasi-Variational Inequality in the Optimal Market-Making Problem with Alpha Signal

Dec 23, 2025
AM
Alexey Meteykin
🏛️ Lomonosov Moscow State University | Vega Institute Foundation

This paper addresses the optimal market-making problem for a limit-order book incorporating alpha signals, formulated as a Hamilton–Jacobi–Bellman quasi-variational inequality (HJBQVI) coupling stochastic and impulse controls. To overcome the time-step restrictions and poor stability of conventional explicit numerical schemes, we propose, for the first time, an unconditionally stable implicit time discretization combined with a policy iteration algorithm. We rigorously establish its convergence to the unique viscosity solution via monotonicity–stability–consistency analysis and the comparison principle. The resulting method significantly enhances the robustness and computational reliability of market-making strategies under high volatility, enabling dynamic spread management and inventory risk hedging. It provides a provably convergent, efficient, and numerically stable framework for solving coupled stochastic–impulse control problems in electronic market making.

Develops implicit scheme for HJBQVI in market-making with alpha signalsEnsures unconditional stability and convergence to viscosity solutionSolves combined stochastic and impulse control in limit order books

This work addresses the challenge of characterizing optimal firm behavior in nonlinear stochastic market environments, where traditional Hamilton–Jacobi–Bellman (HJB) methods are often intractable. The authors introduce a Euclidean path integral control framework that reformulates the equilibrium problem as a forward-looking Lagrangian stochastic control system. By leveraging Itô processes and integrating factors, the approach directly generates optimal strategies without explicitly constructing a value function. For the first time within this framework, a non-cooperative feedback Nash equilibrium is derived and contrasted with mean-field game solutions, revealing fundamental differences from the Pontryagin maximum principle. Combining the Feynman–Kac representation with mean-field approximations, the method yields computationally tractable equilibria for large-scale stochastic markets, with numerical examples demonstrating both its efficacy and its marked divergence from classical HJB solutions.

Hamilton-Jacobi-Bellmanmean-field gamesnon-cooperative Nash equilibrium

This work addresses the challenges of stability and error control in solving Hamilton-Jacobi-Bellman (HJB) equations for continuous-time reinforcement learning by proposing a mesh-free method that integrates physics-informed neural networks, finite differences interpreted via shift operators, stochastic continuous collocation, and greedy policy improvement. Leveraging a hybrid error analysis framework, the approach explicitly disentangles multiple error sources and quantifies gradient amplification factors, thereby effectively circumventing the implicit viscosity-related blow-up issues prevalent in conventional methods. Experimental results on a 64-dimensional linear-quadratic regulator (LQR) problem and several nonlinear control tasks demonstrate that the proposed method significantly outperforms state-of-the-art model-based and model-free reinforcement learning baselines in terms of residual control, policy alignment, and model error, achieving stable and efficient synthesis of optimal feedback controllers.

error analysisHamilton--Jacobi--Bellmanmodel-based reinforcement learning

This work addresses the challenges of high-dimensional stochastic optimal control over long planning horizons, where computational complexity typically grows linearly with time and performance degrades. Focusing on linearly solvable control problems with gradient-drift structure, the authors transform the Hamilton–Jacobi–Bellman equation into a linear partial differential equation governed by an operator ℒ, and establish the unitary equivalence between ℒ and a Schrödinger operator. Leveraging this insight, they derive—for the first time—an analytical solution for symmetric linear quadratic regulators (LQR) with arbitrary terminal costs. By integrating spectral methods, analytical solutions from quantum harmonic oscillators, and neural networks, they design a novel loss function to efficiently learn the eigenfunctions of ℒ. The resulting approach achieves an order-of-magnitude improvement in control accuracy over existing methods on multiple long-horizon benchmark tasks, while reducing both memory and computational complexity from 𝒪(Td) to 𝒪(d).

eigenfunction approximationHamilton-Jacobi-Bellman equationhigh-dimensional control

Hot Scholars

LB

Lijun Bo

Professor, School of Mathematics and Statistics, Xidian University
Stochastic Differential EquationsMathematical Finance
DF

Dena Firoozi

University of Toronto
Mathematical FinanceStochastic ControlMean-Field GamesEstimation
SC

Sylvain Calinon

Idiap Research Institute
robot manipulationlearning from demonstrationfrugal learningoptimal control
CY

Cheng Yu

PhD student, CSE, Ohio State University
audiovisualspeech enhancementspeech separationonline system
LS

Leandro Sánchez-Betancourt

Mathematical Institute, and Oxford-Man Institute, University of Oxford
mathematical financealgorithmic tradingoptimal executionmarket making