Safe Explicable Policy Search

📅 2025-03-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In human-AI collaboration, AI agents must simultaneously satisfy user expectations (explicability) and strict safety constraints—a longstanding challenge. Method: This paper proposes the first unified framework that explicitly embeds safety into explainable policy learning, integrating Constrained Policy Optimization (CPO) with Explainable Policy Search (EPS). Within a single constrained optimization formulation, it jointly maximizes explicability, enforces safety boundaries, and bounds suboptimality—ensuring safety both during training and deployment. Results: Evaluated on Safety-Gym simulations and a real-world robotic platform, the learned policies retain over 90% of baseline task performance while achieving high explicability and zero safety violations. This significantly enhances the reliability and deployability of explainable AI in practical human-AI collaborative settings.

Technology Category

Humans and AI: Explainable AI (XAI) for Human UnderstandingPhilosophy and Ethics of AI: Safety, Robustness & TrustworthinessNatural Language Processing: Safety and Robustness

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsResponsible Web: Machine-in-the-loop, human agency and autonomySecurity and Privacy: Security and privacy of machine learning and AI applications
📝 Abstract
When users work with AI agents, they form conscious or subconscious expectations of them. Meeting user expectations is crucial for such agents to engage in successful interactions and teaming. However, users may form expectations of an agent that differ from the agent's planned behaviors. These differences lead to the consideration of two separate decision models in the planning process to generate explicable behaviors. However, little has been done to incorporate safety considerations, especially in a learning setting. We present Safe Explicable Policy Search (SEPS), which aims to provide a learning approach to explicable behavior generation while minimizing the safety risk, both during and after learning. We formulate SEPS as a constrained optimization problem where the agent aims to maximize an explicability score subject to constraints on safety and a suboptimality criterion based on the agent's model. SEPS innovatively combines the capabilities of Constrained Policy Optimization and Explicable Policy Search. We evaluate SEPS in safety-gym environments and with a physical robot experiment to show that it can learn explicable behaviors that adhere to the agent's safety requirements and are efficient. Results show that SEPS can generate safe and explicable behaviors while ensuring a desired level of performance w.r.t. the agent's objective, and has real-world relevance in human-AI teaming.
Problem

Research questions and friction points this paper is trying to address.

Generating explicable behaviors in AI agents
Incorporating safety considerations in learning settings
Ensuring safe and explicable human-AI teaming
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combines Constrained Policy Optimization with Explicable Policy Search
Formulates as constrained optimization for safety and explicability
Evaluates in safety-gym and physical robot experiments
💼 Related Jobs
No related jobs found.
Akkamahadevi Hanni
Akkamahadevi Hanni
Arizona State University
Explainable AISafety in AIHuman-aware decision-making
J
Jonathan Montano
School of Mathematical and Statistical Sciences, Arizona State University, Tempe, USA
Y
Yu Zhang
School of Computing and Augmented Intelligence, Arizona State University, Tempe, USA