physics-informed residual critic

Designs and implements a critic (value estimator) that combines a physics-based parametric model with a learned residual term, and builds residual-dynamics or residual-value decompositions to represent and correct model mismatch. Focuses on semi-parametric residual learning and physics-aware critic architectures to analyze and reduce value bias from unmodeled dynamics and improve critic estimates in reinforcement-learning / control settings.

physics-informedresidualcritic

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges of deploying Actor-Critic algorithms in real-world control systems, where poor reliability and high sensitivity to hyperparameters often hinder practical application. Focusing on a real-world water treatment plant control task, the authors conduct over 33,000 large-scale ablation experiments to systematically evaluate how key algorithmic components—such as policy update schemes, action distribution representations, gradient estimation methods, and update frequencies—affect performance stability and hyperparameter robustness. Their empirical analysis reveals, for the first time, that commonly adopted default configurations (e.g., Gaussian action distributions with pathwise derivatives) exhibit low reliability, whereas bounded action distributions combined with adaptive update strategies substantially enhance robustness. The work identifies high-stability algorithmic configurations that significantly reduce performance variance under limited tuning budgets, offering actionable, component-level design guidelines for industrial deployment.

Actor-Criticalgorithm reliabilityhyperparameter sensitivity

This study addresses the bias-variance trade-off inherent in single-step Bellman residuals for off-policy evaluation in reinforcement learning by proposing a mixed Bellman residual framework. Methodologically, the approach achieves a natural trade-off through a convex combination of first- and second-order residuals. Leveraging minimax optimization and sample splitting techniques, it constructs a data-dependent adaptive kernel critic along future feature directions, thereby overcoming the limitations of fixed function classes to jointly optimize approximation error and importance sampling variance. Simulation experiments and evaluations on MetaWorld tasks demonstrate that this mixed residual strategy significantly enhances the accuracy of value estimation in complex scenarios.

Bellman ResidualsOffline Policy EvaluationReinforcement Learning

This work addresses the curse of dimensionality and poor cross-parameter generalization in approximating value functions for high-dimensional generalized/differential games with state constraints. We propose a Hybrid Neural Operator (HNO) that maps game parameters directly to the value-function space, integrating supervised data with physics-informed sampling from the full-space-time Hamilton–Jacobi–Isaacs (HJI) equation. To enhance robustness in safety-critical settings, HNO incorporates nonlinear dynamics embedding and a Lipschitz-aware training strategy. Compared to supervised neural operators (SNOs), HNO achieves superior safety performance under identical computational budgets in 9D and 13D nonlinear dynamical systems. Moreover, it enables real-time inference for human–machine and multi-agent interactions while maintaining convergence stability and constraint satisfaction.

Addressing convergence issues in value approximations with state constraintsEnabling generalizable value function approximation across parametric spacesOvercoming curse of dimensionality in differential games

This work addresses the instability commonly observed in online reinforcement learning for dynamic systems, which often stems from the opaque optimization dynamics of the critic network. To this end, the authors propose a critic-matching loss landscape visualization method that projects the critic’s parameter trajectory onto a low-dimensional linear subspace, enabling the construction of a three-dimensional loss surface and a two-dimensional optimization path. The study introduces, for the first time, the concept of a critic-matching loss landscape along with quantitative metrics, and integrates these with a normalized system performance index to enable joint qualitative and quantitative analysis of the training process. Experiments on inverted pendulum and spacecraft attitude control tasks demonstrate that the method effectively reveals distinct loss landscape characteristics associated with stable convergence versus unstable learning, offering a novel tool for understanding and diagnosing online reinforcement learning behavior.

actor-criticinterpretabilityloss landscape

This study addresses the limitations of model predictive control (MPC), which relies on analytical models and struggles with complex state dependencies, as well as the low data efficiency inherent in reinforcement learning. To overcome these challenges, this work proposes a hybrid learning-based MPC framework that integrates local model-based planning with learned components. Specifically, the method employs a residual dynamics network to compensate for model bias and introduces an action-value function to incorporate long-horizon structural information. Furthermore, a GPU-accelerated batch iLQR solver is designed to enable efficient parallel computation. Experimental results demonstrate that the proposed framework significantly improves closed-loop control performance and data efficiency in tasks involving modeling discrepancies, while preserving the structured advantages characteristic of model-based approaches.

Model Predictive ControlOptimal ControlReinforcement Learning

Latest Papers

What's happening recently
View more

This work addresses the challenges of real-time optimal control in high-dimensional dynamical systems, where low sample efficiency, the curse of dimensionality in exploration, and gradient instability hinder performance. The authors propose the PEARL framework, which uniquely integrates the adjoint method with a neural network-based reward function. By exploiting the differentiability of system dynamics, PEARL employs an actor-adjoint algorithm that combines automatic differentiation with adjoint sensitivity analysis to efficiently compute policy gradients over short horizons. This approach enables physics-informed policy learning, significantly enhancing sample efficiency and generalization while operating directly in high-dimensional state-action spaces. Evaluated on unsteady flow navigation tasks, PEARL outperforms existing reinforcement learning methods and scales to high-dimensional control problems without requiring dimensionality reduction or multi-agent architectures.

dynamical systemshigh-dimensional controloptimal control

Accurately predicting the dynamics of deformable objects is crucial yet highly challenging for robotic manipulation. This work proposes the Physics-Guided Residual Dynamics (PGRD) framework, which integrates an optimizable mass-spring physical simulator with a neural network that corrects residual errors in the velocity domain, enabling stable and high-fidelity dynamics modeling. To capture long-range temporal dependencies, the method incorporates a sliding-window Transformer and represents system states using 3D Gaussian Splatting. Evaluated on diverse real-world deformable objects, PGRD significantly outperforms both purely physics-based and purely data-driven baselines. The approach demonstrates successful applications in model predictive control, language-guided manipulation, and interactive video prediction conditioned on action sequences.

deformable object simulationdynamics predictionlearning-based simulation

This study addresses the coordination of locomotion and manipulation (loco-manipulation) behaviors in humanoid robots within the context of multi-objective reinforcement learning. Leveraging the Unitree G1 platform, the authors design a curriculum learning framework comprising 13 progressively complex tasks to systematically evaluate the performance of unified versus dual-critic architectures. Their findings reveal that the choice of critic architecture is a pivotal factor—more influential than reward function design—in determining multi-objective task performance, with the dual-critic configuration effectively mitigating policy degradation. Experimental results in NVIDIA Isaac Lab demonstrate that the dual-critic approach achieves a 3.5× faster goal-completion speed, doubles throughput, and attains a 65.2% effective success rate, substantially outperforming the unified critic counterpart.

critic architecturehumanoid locomotionloco-manipulation

This work proposes a relative value learning framework that shifts the focus from absolute value estimation—commonly used in traditional reinforcement learning—to directly modeling pairwise relative value differences, which are sufficient for policy optimization. The framework introduces an antisymmetric function to represent value differences between state pairs and defines a novel pairwise Bellman operator with a unique fixed point. Building upon this foundation, the authors derive n-step and λ-return objectives and develop an unbiased Relative Generalized Advantage Estimator (R-GAE) for policy gradient computation. When integrated into Proximal Policy Optimization (PPO), the approach achieves performance on par with standard PPO across 49 Atari games, demonstrating that relative value estimation can serve as an effective alternative to absolute value critics.

Critic EstimationPolicy GradientReinforcement Learning

This work addresses the challenge in model-based reinforcement learning where policies trained in simulation fail in the real world due to model prediction errors. To mitigate this, the authors propose shifting the model-learning objective from predictive accuracy to policy robustness by formulating a zero-sum minimax game between the dynamics model and an adversarial policy. Leveraging online learning theory guarantees, a critic-based simplified algorithm, and the Error-MDP duality, they design a provably convergent active data selection mechanism. Evaluated on continuous control tasks, the method reduces prediction errors in critical regions by 1.5–2.2×, enabling policies trained purely in simulation to achieve near-optimal performance when deployed in the real environment.

model-based reinforcement learningpolicy robustnesspredictive accuracy

Hot Scholars

YD

Yi Ding

University of Electronic Science and Technology of China
Deep LearningMedical Image segmentation
WX

Wenjun Xia

Rensselaer Polytechnic Institute
Medical Imaging
RD

Roberto Diversi

University of Bologna
System IdentificationFault Diagnosis and PrognosisStatistical Signal ProcessingAutomatic Control
SG

Somdatta Goswami

Assistant Professor, Civil and Systems Engineering, Johns Hopkins University
Deep LearningPhysics-informed MLComputational MechanicsFracture Mechanics