hierarchical safe exploration

Designs, builds, or analyzes hierarchical exploration and control systems that ensure safety by coordinating high-level planners and low-level policies to generate and follow intermediate safe subgoals, bias exploration toward safe regions, and limit unsafe long-horizon actions. Work includes defining safety constraints and risk-aware exploration strategies, constructing hierarchical policy architectures and subgoal mechanisms, and evaluating tradeoffs between exploration efficiency and safety.

hierarchicalsafeexploration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of safety constraint violations in long-horizon reinforcement learning tasks, which often arise from accumulated errors and limited exploration. To mitigate these issues, the paper proposes a novel safety-aware hierarchical reinforcement learning framework that integrates a learnable world model with a two-level policy architecture. The high-level policy generates safety-oriented subgoals, while the low-level policy leverages imagined rollouts within the learned predictive environment to evaluate and correct unsafe actions before execution, thereby enforcing safety at both levels. This approach is the first to incorporate imagination-based mechanisms into hierarchical reinforcement learning, effectively reducing error accumulation. Empirical results demonstrate that the method significantly improves constraint satisfaction rates and consistently adheres to predefined safety budgets in high-dimensional navigation and manipulation tasks, outperforming state-of-the-art safe reinforcement learning baselines.

hierarchical reinforcement learninglong-horizon tasksreinforcement learning

This work addresses the challenge of ensuring safe robotic exploration and interaction in unknown, stochastic environments, where existing safety-aware control methods often fail due to their reliance on known system dynamics. To overcome this limitation, we propose the Safe Stochastic Explorer framework, which introduces Gaussian processes into safe exploration for the first time. By online learning an unknown safety function and leveraging predictive uncertainty to guide informative data acquisition, our approach enables scalable, goal-directed exploration in continuous state spaces. The method provides probabilistic safety guarantees while effectively balancing exploration efficiency with safety constraints. Extensive simulations and real-world hardware experiments demonstrate that the proposed framework significantly enhances the autonomous safety capabilities of robots operating in complex, uncertain environments.

goal-driven navigationsafe explorationsafety-critical autonomy

A Formal gatekeeper Framework for Safe Dual Control with Active Exploration

Oct 07, 2025
KB
Kaleb Ben Naveed
🏛️ University of Michigan

Existing robust trajectory planning methods under model uncertainty are overly conservative, while active exploration approaches often lack formal safety guarantees or fail to rigorously quantify information gain. Method: We propose a formal dual-control framework that jointly integrates robust planning and safety-constrained active exploration: information-gathering actions are activated only when online verification confirms they explicitly reduce either task cost or parameter uncertainty—without compromising safety. Contribution/Results: Our key innovation extends the gatekeeper architecture by unifying formal safety verification, robust motion planning, and uncertainty-aware online dual control. Evaluated in quadrotor simulations, the method generates trajectories that are simultaneously safe, informative, and cost-efficient, achieving significantly faster convergence of parametric uncertainty compared to baseline approaches.

Ensuring safety while reducing uncertainty through verifiable improvementsIntegrating robust planning with formal safety guarantees for dual controlPlanning safe trajectories under model uncertainty with active exploration

Safe Guaranteed Exploration for Non-linear Systems

Feb 09, 2024
MP
Manish Prajapat
🏛️ ETH Zurich

This paper addresses safe exploration for nonlinear robotic systems under unknown constraints. Method: We propose the first optimal control framework that jointly ensures safety and exploration completeness, integrating model predictive control (MPC), Lipschitz continuity analysis, goal-directed exploration, and receding-horizon replanning to enable online safety verification and proactive constraint avoidance. Contribution/Results: Theoretically, we establish the first finite-time sample complexity bound for general nonlinear systems—guaranteeing satisfaction of safety constraints with arbitrarily high probability and achieving exploration completeness within a finite number of samples. Empirically, we validate the framework in an autonomous driving simulation environment, demonstrating its safety, efficiency, and realizability of theoretical guarantees.

Efficient algorithm for complex dynamics in unknown domainsGuaranteed finite-time exploration with high safety probabilitySafe exploration for non-linear systems with unknown constraints

Reinforcement learning (RL) for robot navigation faces a fundamental trade-off between safety—particularly strict zero-collision guarantees—and task efficiency. Method: This paper proposes a Dynamic Safety Shield (DSS), wherein an RL agent adaptively tunes parameters online to tightly integrate model predictive control (MPC), constrained optimization, and deep RL (PPO/SAC), ensuring hard zero-collision constraints while jointly optimizing exploration efficiency and long-horizon task success. Contribution/Results: DSS achieves the first tight coupling of robust safety control with the RL policy’s long-horizon predictive capability, breaking the conventional safety–performance Pareto frontier. In simulation, it significantly outperforms state-of-the-art methods; real-robot experiments validate its practical efficacy. Compared to classical safety-shield approaches, DSS completes more navigation tasks; relative to constrained RL baselines, it reduces collision frequency by a substantial margin—all while maintaining strict safety guarantees.

Balances exploration and safety in navigation tasksEnsures safe RL training with minimal collisionsImproves goals-to-collisions ratio in dynamic environments

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving both rigorous safety guarantees and efficient coordination in safety-critical multi-agent systems. The authors propose a hierarchical multi-agent reinforcement learning framework in which a low-level controller enforces hard safety constraints through constrained manifold control under mild assumptions, while a high-level policy learns to coordinate agents effectively. This approach represents the first integration of constrained manifold control with hierarchical reinforcement learning in a multi-agent setting, offering provable safety, stable training dynamics, and strong generalization across varying numbers of agents and obstacle configurations. Experimental results demonstrate that the system maintains nearly 100% safety compliance while achieving competitive task performance.

coordinationgeneralizationmulti-agent systems

This work addresses safe trajectory planning under model uncertainty by proposing a Dual-gatekeeper framework that, for the first time, jointly incorporates safety constraints and task-performance budgets within a dual-controller architecture. The approach guarantees formal safety while triggering active exploration only when such exploration can be verified to improve long-term performance. By synergistically combining robust planning with conditional exploration, the method balances immediate task execution with the reduction of uncertainty. Experimental evaluations in quadrotor and autonomous racing scenarios demonstrate that the proposed framework generates trajectories that are not only provably safe and highly efficient but also exhibit strong adaptability, significantly outperforming existing baselines.

active explorationbudget-constrained optimizationdual control

Existing hierarchical decision-making approaches often struggle to simultaneously satisfy constraints and maintain computational efficiency due to misalignment between low-level policies and high-level objectives. This work proposes a principled inverse optimization–based hierarchical framework that, for the first time, systematically constructs structured low-level optimization problems from expert demonstrations, thereby aligning high-level task abstractions with low-level decision-making. By integrating inverse optimization, hierarchical reinforcement learning, and optimal control, the method achieves both interpretability and computational efficiency. Empirical evaluations on resource allocation and obstacle avoidance tasks demonstrate that the approach significantly outperforms end-to-end reinforcement learning, learning-augmented optimal control, and existing hierarchical methods, achieving state-of-the-art performance in both decision quality and computational speed.

Hierarchical Decision MakingInverse OptimizationOptimal Control

Hot Scholars

ST

Sebastian Trimpe

Professor, RWTH Aachen University
ControlMachine LearningNetworked SystemsRobotics
AC

Andrea Carron

Senior Lecturer, ETH Zurich
Model Predictive ControlLearning-based ControlRoboticsMachine Learning
GS

Guanya Shi

Assistant Professor, CMU RI | Amazon Scholar, FAR (Frontier AI & Robotics)
RoboticsRobot LearningReinforcement LearningControl
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management