Score
Designs and implements reinforcement learning agents and training pipelines that incorporate physics-based models or numerical solvers in the loop, using solver outputs as rewards, critics, or constraints to enforce physical validity during learning. Builds methods to fine-tune generators or policies with solver feedback, integrate structure checkers or differentiable/non‑differentiable solvers into optimization, and resolve one‑to‑many inverse mappings in physics‑constrained design or control problems.
本文探讨了在机器人学习中嵌入物理先验知识的方法,以解决数据有限、复杂交互和可靠操作需求的问题,通过整合物理法则来提高学习算法的泛化能力、可解释性和样本效率。
This work addresses the limitations of traditional PDE solvers, which rely heavily on expert knowledge and laborious development, as well as existing large language model (LLM) approaches that focus primarily on reasoning optimization while lacking fine-grained feedback on scientific computation accuracy. The authors propose RLVP, a novel framework that introduces, for the first time, a physics-consistency-based continuous reward mechanism combined with a hard constraint on program executability to train LLMs via reinforcement learning for generating high-accuracy solver code. This approach overcomes the shortcomings of conventional binary verification in scientific computing, significantly outperforming both pretrained and supervised fine-tuning baselines across multiple PDE benchmarks. Notably, even smaller models trained with RLVP surpass state-of-the-art prompting strategies of larger models and demonstrate strong zero-shot transfer across PDE types and compositional generalization of numerical modules.
This work addresses the low sample efficiency and inconsistent actions often observed in reinforcement learning for robotic control, which stem from neglecting known physical dynamics. To this end, the authors propose PIPER, a novel framework that seamlessly integrates physical priors into policy learning by incorporating a differentiable Lagrangian dynamics residual—computed via a standard simulator—as a soft regularization term directly into the policy objective. Crucially, this approach requires no modifications to the underlying simulator or reinforcement learning algorithm. By softly enforcing analytical physical constraints during policy updates, PIPER achieves a tight coupling between physical consistency and learning, significantly improving sample efficiency, training stability, and control accuracy. Empirical results across multiple robotic tasks demonstrate that policies trained with PIPER exhibit superior physical plausibility and overall performance compared to baseline methods.
In reinforcement learning, weak programmability, limited expressivity, and low sample efficiency persist under non-Markovian reward structures. To address these challenges, we propose the physics-informed Reward Machine (pRM). pRM explicitly encodes domain-specific physical priors as symbolic logic constraints into the reward machine’s architecture, enabling decoupled modeling of known environmental priors and unknown dynamics. It supports counterfactual experience generation and differentiable reward shaping, substantially enhancing both the programmability and expressive capacity of reward specifications. The framework is unified for both discrete and continuous physical environments. Experiments across multiple control benchmarks demonstrate that pRM significantly reduces sample complexity, accelerates policy convergence, and improves modeling fidelity for temporally dependent and history-sensitive non-Markovian rewards.
This work addresses the limitations of large language model (LLM) agents in generating physics solvers for simulating dynamical systems by introducing the first benchmark tailored to this task. Drawing upon foundational literature in computer graphics, we construct an evaluation suite comprising 168 tasks that assess agents' comprehensive ability to synthesize executable solvers from code scaffolding, encompassing physical understanding, mathematical reasoning, and software engineering. Furthermore, we define a three-dimensional evaluation framework incorporating execution checks, visual fidelity, and physical plausibility. Experimental results demonstrate that state-of-the-art models achieve an overall score of only 48.7%, revealing that generating accurate and physically consistent solvers remains a formidable challenge.
This work proposes nonlinear fluid instability problems—such as droplet breakup, interfacial mixing, and rogue wave formation—as a novel testbed for reinforcement learning (RL), addressing the challenge of deploying agents in open-world settings where only partial observations and localized interventions are possible within high-dimensional, non-stationary fluid environments. By integrating open-source simulators like Dedalus and JAX-CFD, the study constructs a scalable RL environment featuring well-defined non-stationary dynamics and invariant structures, along with tailored interfaces for state representation, action space, and reward design. Experiments successfully train agents capable of navigating static fluid environments, thereby establishing, for the first time, a systematic framework bridging fluid dynamics and reinforcement learning, and laying the groundwork for future intelligent interaction in complex natural and industrial flows.
This work addresses the challenge of ensuring safety in industrial cyber-physical systems when applying deep reinforcement learning, where black-box exploration may inadvertently violate hardware constraints and conventional reward shaping struggles to balance safety with task performance. To overcome this, the authors propose a physics-informed safety mechanism that embeds a differentiable dynamics model into the loss function of a Proximal Policy Optimization (PPO) policy network. By performing short-horizon forward simulations to predict trajectories, the method imposes soft penalties—decoupled from the task-specific reward—on potential safety violations. This approach regularizes the policy online without requiring intricate reward engineering. Evaluated on a one-degree-of-freedom helicopter simulation platform, the method significantly reduces pitch angle constraint violations while maintaining excellent trajectory tracking performance.
This work addresses the scarcity of frontier tasks in reinforcement learning, where existing task distributions are prone to saturation and naively generated tasks are often overly simple, unsolvable, or ambiguously defined. The authors propose PROPEL, a framework that enables efficient training of a task generator on the learnable frontier without repeatedly invoking a solver. By freezing a reference model and training a lightweight activation probe on a one-time annotated corpus to predict task pass rates, PROPEL replaces costly solver rollouts. Coupled with reinforcement learning, this approach optimizes the generator to produce tasks that remain valid while significantly improving difficulty alignment and training efficiency. Experiments show that on code generation tasks, the proportion of frontier tasks for Qwen2.5-3B and 7B increases from 10.1%/5.3% to 20.0%/12.6%, respectively; on SWE tasks, Qwen3.5-27B achieves a rise from 9.8% to 19.6% in target pass-rate tasks on unseen repositories.
This study addresses the challenge that existing Mixed-Integer Linear Programming (MILP) instance generation methods struggle to simultaneously ensure feasibility and computational hardness while lacking explicit hardness metrics. To overcome these limitations, this work proposes a reinforcement learning framework driven by solver feedback. Specifically, it constructs reward signals based on branch-and-bound node counts and optimality gaps, employing the Group Relative Policy Optimization (GRPO) algorithm to fine-tune large language models such as Gemma and Qwen. Furthermore, an asymmetric seedless self-play mechanism is introduced to iteratively guide the model in generating feasible yet computationally challenging instances. Experimental results demonstrate that the proposed approach significantly increases SCIP search node counts and post-cut gaps, effectively expanding the difficulty spectrum of generated instances while improving their feasibility rates.
This work addresses the semantic gap between CAD and CAE in industrial design—specifically, how to effectively translate simulation feedback into geometric modifications that satisfy multiple constraints. To bridge this gap, the authors propose a tool-augmented reinforcement learning framework that enables large language models to orchestrate a closed-loop workflow encompassing CAD modeling, CAE simulation, result interpretation, and geometric refinement. A multi-constraint joint reward mechanism ensures solution feasibility and robustness of the toolchain, while a novel executable CAD-CAE dataset covering 25 categories of industrial components supports training in realistic scenarios. Experimental results demonstrate that the proposed approach significantly enhances the performance of small open-source large language models in constraint-driven design, outperforming both leading open-source and proprietary models in terms of feasibility, efficiency, and stability.
This work addresses the challenges of real-time optimal control in high-dimensional dynamical systems, where low sample efficiency, the curse of dimensionality in exploration, and gradient instability hinder performance. The authors propose the PEARL framework, which uniquely integrates the adjoint method with a neural network-based reward function. By exploiting the differentiability of system dynamics, PEARL employs an actor-adjoint algorithm that combines automatic differentiation with adjoint sensitivity analysis to efficiently compute policy gradients over short horizons. This approach enables physics-informed policy learning, significantly enhancing sample efficiency and generalization while operating directly in high-dimensional state-action spaces. Evaluated on unsteady flow navigation tasks, PEARL outperforms existing reinforcement learning methods and scales to high-dimensional control problems without requiring dimensionality reduction or multi-agent architectures.