Score
Designs and builds inverse-dynamics models that infer the actions or control inputs needed to reach predicted future states, and implements safe fallback policies that generate behaviorally consistent alternative actions when the primary controller is uncertain or unsafe. Also analyzes and implements mechanisms to detect when to adaptively switch to the fallback based on safety or uncertainty criteria.
Formalizing safety constraints in reinforcement learning remains challenging due to the difficulty of specifying explicit safety criteria or accurate system dynamics a priori. Method: This paper proposes a model-free safety framework grounded in reversibility—using state reversibility as an implicit, knowledge-free safety criterion. It integrates Model Predictive Path Integral (MPPI) control with real-time reversibility assessment during policy training, dynamically intercepting irreversible (i.e., potentially unsafe) actions via black-box environment queries. A decoupled safety evaluation architecture ensures orthogonality between safety enforcement and policy optimization. Contribution/Results: The approach achieves 100% interception of unsafe actions while matching the training efficiency and task performance of PPO baselines. It is the first work to introduce reversibility as a principled foundation for RL safety control, offering a theoretically interpretable, lightweight, and general-purpose safety paradigm for implicit safety constraints.
This work addresses the challenge of learning policies that simultaneously achieve high reward and satisfy safety constraints in Markov decision processes where the cost function is unknown and constraints are unobservable. To this end, the authors propose SafeQIL, an algorithm that, for the first time, integrates safety assessment into a Q-learning-based inverse constrained reinforcement learning framework. SafeQIL jointly models task rewards and safety constraints by introducing a “commitment” mechanism over state-action pairs and optimizes the model via maximum likelihood estimation using expert demonstrations. Evaluated across multiple benchmark tasks, SafeQIL significantly outperforms existing methods, achieving both enhanced policy performance and guaranteed safety, thereby effectively balancing conservative constraint adherence with exploratory high-reward behavior.
研究解决了未知动态下机器人运动规划与控制问题,通过预计算开环控制序列和量化转换风险的方法处理观测盲区。
This work addresses the challenge of enforcing hard affine state constraints in black-box hybrid dynamical systems subject to instantaneous state jumps and unknown nonlinear dynamics. To this end, the authors propose a novel reinforcement learning strategy that incorporates an affine repulsion mechanism near constraint boundaries and introduces a secondary repulsion region just before the system’s reset map. This approach ensures strict satisfaction of safety constraints in closed-loop operation without requiring an explicit system model. Notably, it establishes the first provably safe reinforcement learning framework for black-box hybrid systems, effectively mitigating constraint violations induced by state discontinuities. Empirical evaluations on benchmark tasks—including a constrained pendulum and a juggling paddle system—demonstrate that the proposed method consistently guarantees constraint satisfaction while learning superior policies compared to existing reward-shaping and control barrier function–based approaches.
To address the lack of unified, flexible, and real-time inverse dynamics (ID) computation across simulation and real-robot deployments, this paper introduces a lightweight, robot-agnostic ID software library built natively for ROS 2. Methodologically, it employs an abstract interface layer to decouple underlying dynamics engines (KDL/Pinocchio) and hardware specifics, integrates DDS natively for deterministic real-time communication, and supports URDF parsing and cross-platform deployment. Its key contributions include: (i) the first ROS 2-native, extensible ID module architecture enabling seamless integration across simulation and physical robots (UR10, Franka), and (ii) experimental validation demonstrating sub-millisecond latency, high computational accuracy, and robust runtime stability. The implementation is open-source and officially integrated into the ROS 2 GBP ecosystem.
This study investigates the causal relationship between information acquisition and commitment behavior in safe adaptive control: specifically, whether a controller can distinguish between system models requiring distinct policies through informative experiments within finite time under unified safety constraints. To address this, the work introduces the notion of “pre-commitment information,” quantifies observable information via Kullback–Leibler divergence, and integrates causal analysis, semidefinite programming, and linear Gaussian system theory to establish fundamental limits on information gathering under safety constraints. The main contributions include proving that in linear quadratic regulation, if the oracle gap is Ω(T), then every uniformly safe policy incurs linear regret; furthermore, the paper provides sufficient conditions for recoverability together with corresponding upper-bound certificates.
This study addresses the challenge of learning unknown constraints, which typically relies on known dynamics or entails high-risk online exploration. To this end, we propose the CF-KKT framework, which integrates differentiable dynamics models with locally optimal demonstrations to combine the data efficiency of CIOC with the flexibility of ICRL. By leveraging Karush–Kuhn–Tucker (KKT) condition optimization and generating synthetic infeasible data through counterfactual actions, the proposed method recovers unknown constraints without requiring hazardous online exploration. Experimental evaluations on high-dimensional robotic control tasks demonstrate that the CF-KKT framework significantly improves both safety and data efficiency compared to offline ICRL baselines.
This work addresses a critical limitation in existing safe reinforcement learning approaches, which often neglect the inherent control-affine structure of residual dynamics, leading to either overly conservative behavior or compromised safety guarantees. To overcome this, the authors propose a novel framework that integrates control-affine dynamical modeling with certifiably safe policy synthesis. Specifically, they explicitly encode the control-affine structure into data-driven modeling by representing system dynamics using control-affine random Fourier features (ARFF). Uncertainty in the learned model is rigorously quantified via adaptive conformal prediction, enabling the construction of control barrier functions that yield theoretically verifiable safety certificates. Evaluated on inverted pendulum and 3D quadrotor simulations, the method achieves significantly improved task efficiency while maintaining strict safety guarantees.
This study addresses the challenge of coordinating end-effector tracking with base motion in legged manipulation. To this end, it proposes a coupled framework integrating a response-shaping training strategy with closed-loop model predictive control (MPC). This approach pioneers the combination of response consistency and policy-aware MPC by unifying reinforcement learning, system identification, and response shaping techniques, effectively resolving command–response inconsistencies under dynamic loads. Simulation results demonstrate an approximate 28% reduction in position and orientation errors. Furthermore, the proposed method is successfully validated on a physical robot, exhibiting continuous coordinated manipulation capabilities between the mobile base and the robotic arm.
本文通过引入REVERSAL-BENCH,利用可调参数控制环境的可逆性及提供状态恢复验证机制,解决无外部重置强化学习中因不可逆事件导致的学习中断问题。