solve hjb equations

Formulate, analyze, and solve Hamilton–Jacobi–Bellman (HJB) partial differential equations to compute optimal value functions and derive feedback control laws for continuous-time deterministic and stochastic dynamic optimization problems; this includes developing existence/uniqueness/viscosity-solution theory and boundary conditions. Implement and evaluate analytical and numerical HJB methods (e.g., characteristics, finite-difference/element schemes, discretization and policy iteration) to produce implementable controllers, assess robustness and performance, and verify stability and optimality properties.

solvehjbequations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Ensemble based Closed-Loop Optimal Control using Physics-Informed Neural Networks

Oct 20, 2025
JB
Jostein Barry-Straume
🏛️ Virginia Tech

Solving the Hamilton–Jacobi–Bellman (HJB) equation for optimal control of nonlinear dynamical systems remains challenging due to its analytical intractability and high computational cost in numerical methods. To address this, we propose a multi-stage physics-informed neural network (PINN)-based ensemble learning framework that jointly learns the optimal cost function and the corresponding closed-loop control policy in an end-to-end manner. Departing from conventional stability-inducing regularization terms, our method enables both single-policy and ensemble-policy deployment, significantly enhancing robustness against state disturbances, measurement noise, and initial-condition sensitivity. By tightly integrating prior physical knowledge—encoded via the HJB equation—with data-driven learning, the approach achieves knowledge- and data-cooperative optimal control. Extensive experiments demonstrate high accuracy and strong generalization across infinite-horizon settings, diverse initial conditions, and perturbed dynamics.

Controlling nonlinear systems with noisy states and perturbationsDeveloping ensemble PINN framework without stabilizer termsSolving Hamilton-Jacobi-Bellman equation for optimal control

Neural optimal controller for stochastic systems via pathwise HJB operator

Feb 23, 2024
ZJ
Zhe Jiao
🏛️ Northwestern Polytechnical University | Massachusetts Institute of Technology

High-dimensional stochastic optimal control remains challenging due to the curse of dimensionality and limitations of conventional approaches relying on probabilistic representations of the Hamilton–Jacobi–Bellman (HJB) equation. Method: This paper proposes a physics-informed deep learning framework grounded in a pathwise HJB operator, unifying modeling and solution. It introduces the pathwise HJB operator as a novel physical constraint and designs two tailored numerical schemes—accommodating both explicit and implicit optimal control structures—integrated with PINNs, dynamic programming, SDE discretization, and HJB path-integral representations. Contribution/Results: A unified theoretical analysis quantifies truncation, approximation, and optimization errors. The framework significantly improves control accuracy and generalization across diverse high-dimensional tasks while ensuring interpretability and analytical tractability, establishing a new paradigm for real-time optimal control of complex stochastic systems.

Developing efficient deep learning control methodsSolving high-dimensional stochastic control problemsUsing physics-informed neural networks for HJB equations

This work addresses the numerical solution of high-dimensional Hamilton–Jacobi–Bellman (HJB) equations arising in stochastic optimal control. Methodologically, we propose a novel neural network-based actor–critic algorithm: the critic network is designed with an architecture that automatically satisfies boundary conditions, and a bias-gradient technique is introduced to reduce computational cost; the actor update minimizes the integral of the Hamiltonian, circumventing explicit differentiation. Theoretically, we prove that, in the infinite-width limit, the training dynamics converge to an infinite-dimensional ordinary differential equation whose fixed point exactly coincides with the true HJB solution. Experimentally, the method achieves high accuracy and robustness on problems up to 200 dimensions—including cases with non-convex Hamiltonians and linear-quadratic regulators—significantly extending the dimensionality frontier for deep learning–based HJB solvers.

Analyzing actor-critic methods for solving high-dimensional HJB equations.Demonstrating algorithm performance in stochastic control up to 200 dimensions.Ensuring boundary conditions and reducing computational costs in critic architecture.

Is Bellman Equation Enough for Learning Control?

Mar 04, 2025
HY
Haoxiang You
🏛️ Yale University | Microsoft Research

This paper reveals the non-uniqueness of solutions to the Bellman equation in continuous-state spaces—particularly, linear systems admit at least $inom{2n}{n}$ distinct solutions—causing value-based learning to converge to spurious “optimal” policies that satisfy the Bellman equation but yield unstable closed-loop dynamics. Method: We propose a positive-definite neural network architecture that structurally enforces value function positivity via parameter constraints, thereby embedding Bellman equation solving within a Lyapunov stability framework. Contribution/Results: We provide the first quantitative characterization of solution multiplicity and establish a stability-driven, structured value-function modeling paradigm. Theoretically, our design guarantees existence and convergence to the unique stable optimal solution. Empirically, it achieves 100% stable convergence on both linear and nonlinear systems, significantly enhancing closed-loop robustness. Crucially, it delivers verifiable convergence guarantees—unattainable under conventional unconstrained value-function approximators.

Proposed neural architecture ensures convergence to stable, optimal solutions.Uniqueness of Bellman equation solutions fails in continuous state spaces.Value-based methods may converge to unstable, non-optimal solutions.

This work addresses the computational challenges in non-Markovian stochastic optimal control, where the value function is governed by a stochastic Hamilton–Jacobi–Bellman (SHJB) equation whose measurability-induced randomness impedes tractable solution. Under the assumption of control-independent stochastic integral coefficients, the paper introduces the first policy iteration framework tailored to semilinear SHJB equations. The approach iteratively linearizes the original problem into a sequence of linear equations, which are then solved using tools from stochastic analysis and functional approximation. Theoretical analysis establishes that the resulting sequence of approximations converges monotonically in the mean-square sense and exhibits exponential convergence rates, thereby significantly enhancing both the feasibility and efficiency of computing value functions in non-Markovian settings.

computational challengesHamilton-Jacobi-Bellman equationnon-Markovian

Latest Papers

What's happening recently
View more

This work addresses the challenges of stability and error control in solving Hamilton-Jacobi-Bellman (HJB) equations for continuous-time reinforcement learning by proposing a mesh-free method that integrates physics-informed neural networks, finite differences interpreted via shift operators, stochastic continuous collocation, and greedy policy improvement. Leveraging a hybrid error analysis framework, the approach explicitly disentangles multiple error sources and quantifies gradient amplification factors, thereby effectively circumventing the implicit viscosity-related blow-up issues prevalent in conventional methods. Experimental results on a 64-dimensional linear-quadratic regulator (LQR) problem and several nonlinear control tasks demonstrate that the proposed method significantly outperforms state-of-the-art model-based and model-free reinforcement learning baselines in terms of residual control, policy alignment, and model error, achieving stable and efficient synthesis of optimal feedback controllers.

error analysisHamilton--Jacobi--Bellmanmodel-based reinforcement learning

Implicit Numerical Scheme for the Hamilton-Jacobi-Bellman Quasi-Variational Inequality in the Optimal Market-Making Problem with Alpha Signal

Dec 23, 2025
AM
Alexey Meteykin
🏛️ Lomonosov Moscow State University | Vega Institute Foundation

This paper addresses the optimal market-making problem for a limit-order book incorporating alpha signals, formulated as a Hamilton–Jacobi–Bellman quasi-variational inequality (HJBQVI) coupling stochastic and impulse controls. To overcome the time-step restrictions and poor stability of conventional explicit numerical schemes, we propose, for the first time, an unconditionally stable implicit time discretization combined with a policy iteration algorithm. We rigorously establish its convergence to the unique viscosity solution via monotonicity–stability–consistency analysis and the comparison principle. The resulting method significantly enhances the robustness and computational reliability of market-making strategies under high volatility, enabling dynamic spread management and inventory risk hedging. It provides a provably convergent, efficient, and numerically stable framework for solving coupled stochastic–impulse control problems in electronic market making.

Develops implicit scheme for HJBQVI in market-making with alpha signalsEnsures unconditional stability and convergence to viscosity solutionSolves combined stochastic and impulse control in limit order books

This work addresses the lack of verifiable error guarantees in existing physics-informed neural networks (PINNs) when solving Lyapunov and Hamilton-Jacobi-Bellman (HJB) equations—specifically, whether small PDE residuals imply small solution errors remains unclear. To bridge this gap, we develop the first theoretical framework that provides rigorous, verifiable error bounds for PINN approximations of these critical partial differential equations arising in nonlinear system analysis and control. Our approach converts residual bounds into relative error bounds and a posteriori estimates for the true solution, proves that one-sided residual bounds suffice to guarantee the PINN approximation itself constitutes a valid Lyapunov function, and delivers computable upper and lower bounds on the optimal value function along with quantified optimality gaps for feedback policies in HJB problems. Numerical experiments demonstrate the effectiveness and practical utility of the proposed methodology.

Error BoundsHamilton-Jacobi-Bellman EquationsLyapunov Equations

This work proposes an end-to-end framework for solving fully nonlinear second-order partial differential equations—such as infinite-dimensional Hamilton–Jacobi–Bellman (HJB) and Kolmogorov equations—defined on separable Hilbert spaces, along with their associated optimal control problems, without resorting to finite-dimensional projections. The core approach, termed the Hilbert–Galerkin Neural Operator (HGNO), directly minimizes the L² norm of the PDE residual in the infinite-dimensional setting by integrating a deep Hilbert–Galerkin method with a Hilbert-space Actor-Critic reinforcement learning algorithm. Theoretically, the study establishes the first universal approximation theorem applicable to infinite-dimensional PDEs involving first- and second-order Fréchet derivatives and unbounded operators. Numerical experiments demonstrate the method’s efficacy and novelty by successfully solving deterministic and stochastic optimal control problems linked to the heat and Burgers equations.

Hilbert spacesHJB equationsinfinite-dimensional PDEs

This study addresses the challenge of efficiently solving the Hamilton–Jacobi–Bellman equation for optimal control of high-dimensional nonlinear control-affine systems. The authors propose a novel approach that leverages the Pontryagin maximum principle to generate training data comprising the value function, its gradient, and Hessian. By integrating hyperbolic cross sparse polynomial expansions with weighted least squares regression, they construct a high-fidelity approximation model. A key innovation lies in explicitly incorporating Hessian information into the supervised learning framework, complemented by a partial Hessian strategy that balances computational efficiency and approximation accuracy. Experimental results demonstrate that, in high-dimensional settings, the proposed method reduces the required number of training samples by nearly an order of magnitude compared to approaches using only function values, while significantly improving both value function approximation accuracy and closed-loop control performance.

Hamilton-Jacobi-Bellman PDEshigh-dimensional problemsoptimal control

Hot Scholars

SH

Sylvia Herbert

Assistant Professor, University of California, San Diego
Safe ControlControl TheoryAutonomous SystemsRobotics
EB

Erhan Bayraktar

Professor, University of Michigan, Department of Mathematics
Mathematical FinanceStochastic optimal controlprobabilityinsurance mathematics
DL

Donghwan Lee

KAIST
Decision makingcontroland optimization
SB

Somil Bansal

Assistant Professor, Stanford University
RoboticsArtificial intelligenceDynamic systems and control