Score
Design, build, and analyze algorithms, policies, and mechanisms that govern how an agent or system explores unknown actions, states, or options while bounding or minimizing safety violations and risk. This includes devising exploration–exploitation trade-offs, creative or curiosity-driven exploration heuristics, safety constraints or filters, risk-sensitive objectives, and uncertainty estimation and monitoring methods to keep exploration within acceptable limits.
This paper addresses the insufficient modeling and enforcement of safety constraints in Safe Reinforcement Learning (SafeRL) and Constrained Markov Decision Processes (CMDPs), particularly in single- and multi-agent settings. To this end, it establishes— for the first time—the unified mathematical framework bridging SafeRL and Safe Multi-Agent RL (SafeMARL). Methodologically, it proposes a systematic approach integrating constrained optimization, policy-gradient-based safety guarantees, safe exploration strategies, multi-agent game-theoretic modeling, and rigorous CMDP-theoretic analysis. The work distills five key open problems in SafeRL/SafeMARL, three of which specifically address novel challenges arising in multi-agent safety-aware cooperation and competition. As a result, it delivers a comprehensive technical guide encompassing formal definitions, theoretical theorems, algorithmic designs, and emerging research directions—constituting the first authoritative reference that simultaneously ensures theoretical rigor and practical applicability for SafeRL and SafeMARL.
This work addresses the challenge of ensuring safe robotic exploration and interaction in unknown, stochastic environments, where existing safety-aware control methods often fail due to their reliance on known system dynamics. To overcome this limitation, we propose the Safe Stochastic Explorer framework, which introduces Gaussian processes into safe exploration for the first time. By online learning an unknown safety function and leveraging predictive uncertainty to guide informative data acquisition, our approach enables scalable, goal-directed exploration in continuous state spaces. The method provides probabilistic safety guarantees while effectively balancing exploration efficiency with safety constraints. Extensive simulations and real-world hardware experiments demonstrate that the proposed framework significantly enhances the autonomous safety capabilities of robots operating in complex, uncertain environments.
本文针对危险环境下的机器人探索问题,提出了一种基于Prelec概率加权的行为信息目标方法,以平衡信息获取与风险。
Achieving safe exploration with zero constraint violations in reinforcement learning remains highly challenging. This work proposes the Safe Equilibrium Exploration (SEE) framework, which formalizes the objective of safe exploration as a dynamic equilibrium between the feasible region and environmental model uncertainty, and efficiently attains this equilibrium through alternating optimization of the two components. By integrating graph-structured uncertainty modeling, dynamic expansion of the feasible region, and an alternating optimization strategy, SEE achieves zero constraint violations in classical control tasks and rapidly converges to the equilibrium within only a few iterations, substantially improving the efficiency of safe exploration.
This work addresses the challenge of safety constraint violations in long-horizon reinforcement learning tasks, which often arise from accumulated errors and limited exploration. To mitigate these issues, the paper proposes a novel safety-aware hierarchical reinforcement learning framework that integrates a learnable world model with a two-level policy architecture. The high-level policy generates safety-oriented subgoals, while the low-level policy leverages imagined rollouts within the learned predictive environment to evaluate and correct unsafe actions before execution, thereby enforcing safety at both levels. This approach is the first to incorporate imagination-based mechanisms into hierarchical reinforcement learning, effectively reducing error accumulation. Empirical results demonstrate that the method significantly improves constraint satisfaction rates and consistently adheres to predefined safety budgets in high-dimensional navigation and manipulation tasks, outperforming state-of-the-art safe reinforcement learning baselines.
Deep reinforcement learning (DRL) achieves strong performance in control tasks, yet its exploratory training process frequently violates safety constraints, hindering real-world deployment; meanwhile, conventional safety-critical control methods rely on precise system dynamics models, which are often unavailable in practical scenarios. Method: We propose the first state-level safety policy optimization framework that requires no prior knowledge of system dynamics and guarantees zero safety violations throughout training. Our approach integrates a black-box safety monitor, a safety-set-guided exploration mechanism, and an imagined cost constraint into a differentiable policy gradient objective. Results: Experiments on high-dimensional robotic control tasks demonstrate strict adherence to state-level safety constraints, significantly outperforming existing safe DRL baselines while preserving efficient policy learning and robust safety assurance.
Current language model agents lack systematic means to distinguish and quantify exploration versus exploitation errors when internal policies are inaccessible. This work proposes a policy-agnostic evaluation framework that leverages embodied AI principles to construct a controllable, partially observable 2D grid environment. By integrating programmable map generation with task-agnostic directed acyclic graphs (DAGs) of unknown tasks, the framework dynamically modulates the difficulty of exploration or exploitation and defines quantifiable metrics for both error types based solely on observable behavior. This approach enables, for the first time, a decoupled assessment of language models’ exploration–exploitation behaviors, revealing significant yet divergent failure modes across state-of-the-art models. Notably, reasoning-oriented models exhibit superior performance, and their capabilities can be effectively enhanced through lightweight reasoning guidance.
Efficient autonomous exploration in sparse-reward environments remains a fundamental challenge in reinforcement learning. This work proposes a novel paradigm that decouples exploration from policy optimization: during the exploration phase, it abandons conventional reinforcement learning and instead employs a “Go-With-The-Winner” tree search guided by epistemic uncertainty to actively expand state coverage; subsequently, it distills the collected exploration trajectories into a deployable policy via supervised inverse dynamics learning. The approach requires neither expert demonstrations nor domain-specific knowledge and operates directly from end-to-end pixel inputs. It substantially outperforms existing methods on challenging Atari benchmarks such as Montezuma’s Revenge, Pitfall!, and Venture, achieving an order-of-magnitude improvement in exploration efficiency. Notably, it is the first method to solve high-dimensional continuous-control sparse-reward tasks—including MuJoCo Adroit and AntMaze—directly from pixels.
This work addresses safe trajectory planning under model uncertainty by proposing a Dual-gatekeeper framework that, for the first time, jointly incorporates safety constraints and task-performance budgets within a dual-controller architecture. The approach guarantees formal safety while triggering active exploration only when such exploration can be verified to improve long-term performance. By synergistically combining robust planning with conditional exploration, the method balances immediate task execution with the reduction of uncertainty. Experimental evaluations in quadrotor and autonomous racing scenarios demonstrate that the proposed framework generates trajectories that are not only provably safe and highly efficient but also exhibit strong adaptability, significantly outperforming existing baselines.
This work addresses the tendency of large language model agents to prematurely rely on prior knowledge in unfamiliar environments, leading to inadequate exploration and task failure. The authors propose an Explore-then-Act paradigm that decouples exploration from execution: agents first systematically gather environmental information within a fixed interaction budget, then leverage the acquired embodied knowledge to accomplish tasks. The study formally characterizes an agent’s autonomous exploration capability, introduces a verifiable exploration coverage metric, and devises an alternating training strategy that interleaves exploration and task objectives. Furthermore, a dual-trajectory reinforcement learning framework with a verifiable reward mechanism is introduced to optimize behavior policies. This approach substantially enhances generalization in unseen environments, overcoming the limitations of conventional methods whose narrow behavioral repertoires constrain downstream task performance.
This study addresses how an agent should dynamically allocate limited effort between novel and established solution approaches to maximize the probability of success when problem difficulty is unknown. By integrating Bayesian learning, dynamic optimization, and mechanism design theory, the authors formulate a principal–agent model that captures both the exploration–exploitation trade-off and moral hazard. The analysis reveals that the optimal policy alternates between trying new and existing methods, and that learning effects lead to front-loaded incentive schemes. These findings offer novel theoretical foundations for designing innovation strategies and dynamic incentives in creative endeavors such as scientific research and product development.