Score
Designs and implements differentiable modules that encode safety as occupancy or potential fields and use their gradients to guide, nudge, or project sampled trajectories or control signals onto collision-free regions. Builds projection layers, safety shields, and potential-based guidance functions that compute time-varying occupancy probabilities and apply differentiable repulsive guidance or loss terms during sampling or optimization.
Existing provably safe reinforcement learning (RL) methods for safety-critical autonomous robotics primarily target sample-based RL, leaving high-performance, sample-efficient analytic-gradient RL algorithms—such as Proximal Policy Optimization (PPO)—without training-time safety guarantees, thereby exacerbating the safety gap between simulation and reality. Method: This paper introduces the first provably safe training-time framework for analytic-gradient RL. It proposes a differentiable safety module integrating state-action space projection mapping, gradient redefinition, and differentiable physics simulation to enable end-to-end co-optimization of safety constraints and policy learning. Contribution/Results: Evaluated on canonical control benchmarks, the approach achieves zero safety violations throughout training while matching the performance of unconstrained baselines. It effectively bridges the longstanding trade-off between safety assurance and policy performance, enabling safer deployment of gradient-based RL in real-world robotic systems.
This work addresses the limitation of conventional multi-agent navigation methods, which typically assume a static environment and overlook the potential of environmental configuration to enhance safety and efficiency. The authors propose a differentiable co-optimization framework that jointly optimizes environment parameters and agent trajectories through a bilevel optimization formulation: the lower level employs an interior-point method to minimize trajectory costs for agents, while the upper level adjusts environmental parameters via gradient ascent to improve navigation safety. A novel safety metric grounded in measure theory is introduced, and end-to-end gradient propagation is enabled by leveraging KKT conditions and the implicit function theorem. Experimental results demonstrate that the proposed approach significantly enhances both safety and efficiency of multi-agent navigation in warehouse logistics and urban traffic scenarios.
Autonomous driving in complex, dynamic traffic environments often suffers from safety compromises due to the disconnect between prediction uncertainty and motion planning. This work proposes a sample-conditioned differentiable planning framework that, for the first time, directly integrates multimodal trajectories generated by a conditional diffusion model into the planning optimization loop. The approach explicitly controls tail risk through an empirical Conditional Value-at-Risk (CVaR) constraint and incorporates directed graphs for structured scene modeling. Evaluated on the Waymo Open Motion and Argoverse 2 datasets, the method significantly outperforms existing approaches, achieving state-of-the-art performance across key metrics including safety, efficiency, and passenger comfort.
Generating safe and dynamically feasible trajectories for complex robotic systems—particularly in non-convex environments—remains challenging due to the difficulty of jointly satisfying safety and motion-dynamics constraints. Method: This paper proposes a training-free diffusion-based planning framework. Its core innovation is the first integration of a safety shielding mechanism directly into the denoising process of the diffusion model—bypassing post-hoc correction—and thereby ensuring end-to-end trajectory generation that intrinsically satisfies both safety and dynamical feasibility. The approach unifies a model-based diffusion architecture, explicit kinematic and dynamic modeling, and real-time safety verification. Results: Evaluated on high-dimensional nonlinear systems such as tractor-trailer models, the method achieves state-of-the-art task success rates, substantially improves safety guarantees, and completes single-shot planning in under one second.
To address the challenge of safe motion planning for mobile robots in obstacle-dense environments under dynamic uncertainty, this paper proposes a two-layer cooperative framework. At the upper layer, a dynamics-coupled Guidance Vector Field (GVF) is constructed to generate curvature-constrained, safe, and dynamically feasible reference trajectories. At the lower layer, an online dynamic model is established by fusing a deep Koopman operator with a sparse Gaussian process for real-time uncertainty compensation, while a game-theoretic safety barrier function ensures theoretical safety guarantees. This work is the first to jointly integrate deep Koopman learning and sparse GP regression for online dynamic correction, and introduces the novel concept of dynamics-aware GVF design. Extensive simulations and real-world experiments on quadrotor and UGV platforms demonstrate significant improvements in obstacle avoidance success rate and trajectory smoothness, achieving a safety rate exceeding 99.2% and a planning frequency of up to 50 Hz.
This work addresses the challenge of ensuring safety in industrial cyber-physical systems when applying deep reinforcement learning, where black-box exploration may inadvertently violate hardware constraints and conventional reward shaping struggles to balance safety with task performance. To overcome this, the authors propose a physics-informed safety mechanism that embeds a differentiable dynamics model into the loss function of a Proximal Policy Optimization (PPO) policy network. By performing short-horizon forward simulations to predict trajectories, the method imposes soft penalties—decoupled from the task-specific reward—on potential safety violations. This approach regularizes the policy online without requiring intricate reward engineering. Evaluated on a one-degree-of-freedom helicopter simulation platform, the method significantly reduces pitch angle constraint violations while maintaining excellent trajectory tracking performance.
本文提出了一种基于信号时序逻辑的安全感知模型预测路径积分控制方法,通过将STL约束编码为控制屏障函数,并结合MPPI控制器,以提高机器人在复杂任务中的安全性和效率。
This work addresses the problem of safe motion planning for robots operating in complex, cluttered environments under stochastic disturbances with unknown distributions. The authors propose a sampling-based, provably safe planning algorithm that constructs Wasserstein ambiguity tubes from trajectory data to tightly envelop the evolution of state distributions with high confidence. Building upon these tubes, the method incrementally grows a planning tree that satisfies chance constraints. A key innovation lies in replacing a single high-dimensional ambiguity tube with multiple lower-dimensional ones, substantially reducing conservatism and improving scalability. Additionally, an efficient, probabilistically complete bandit-style validity checker is introduced. Experimental results demonstrate that the approach reliably generates feasible trajectories meeting stringent safety thresholds in highly cluttered scenarios, significantly outperforming state-of-the-art methods.
This work addresses the challenge of enabling autonomous robots to simultaneously satisfy complex temporal tasks and safety requirements in uncertain environments, where existing approaches are either non-differentiable or neglect uncertainty in belief space. The paper introduces a differentiable probabilistic Signal Temporal Logic (pdSTL) framework that unifies probabilistic semantics with differentiable robustness for the first time. It computes conservative satisfaction bounds via interval-valued probability semantics and unfolds STL operators through an LSTM-like recursive structure, achieving linear-time, end-to-end differentiable monitoring and optimization. This approach enables safe policy optimization over belief trajectories with formal probabilistic guarantees. Evaluations in simulated obstacle avoidance and lane-changing scenarios, as well as real-world experiments on a Crazyflie quadrotor, demonstrate that pdSTL significantly improves both safety margins and optimization efficiency compared to deterministic differentiable STL.
This work addresses the challenge of coordinating autonomous vehicles and mobile robots with heterogeneous dynamics in high-density, unsignalized urban intersections. To this end, the authors propose a Differentiable Model Predictive Safety (DMPS) framework that integrates the foresight of model predictive control into end-to-end reinforcement learning. DMPS jointly optimizes latent dynamical trajectory prediction and a differentiable safety critic, and—uniquely—enables gradient-based safety guidance directly in the action space. Experimental results demonstrate that DMPS reduces collision rates to below 5.6% in mixed-traffic simulations while simultaneously maintaining energy efficiency and throughput, thereby significantly enhancing real-time collision avoidance capabilities in heterogeneous multi-agent systems.