A Survey on Explainable Deep Reinforcement Learning

📅 2025-02-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Deep reinforcement learning (DRL) excels in sequential decision-making but suffers from limited interpretability and trustworthiness in high-stakes applications due to its black-box nature. To address this, we present a systematic survey of eXplainable Reinforcement Learning (XRL) and propose a novel, unified four-level interpretability framework—spanning feature-, state-, dataset-, and model-level explanations. We further introduce a cross-granularity evaluation体系 integrating attention mechanisms, saliency mapping, counterfactual reasoning, surrogate modeling, attribution analysis, and Reinforcement Learning from Human Feedback (RLHF). Additionally, we establish a taxonomy and applicability boundary matrix for XRL methods. Empirical evaluations demonstrate that our framework significantly enhances policy generalization, adversarial robustness, safety guarantees, and alignment with human preferences—thereby establishing an explanation-driven paradigm for trustworthy AI deployment.

Technology Category

Humans and AI: Explainable AI (XAI) for Human UnderstandingMachine Learning: Transparent, Interpretable, Explainable MLPhilosophy and Ethics of AI: Accountability, Interpretability & Explainability

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Deep Reinforcement Learning (DRL) has achieved remarkable success in sequential decision-making tasks across diverse domains, yet its reliance on black-box neural architectures hinders interpretability, trust, and deployment in high-stakes applications. Explainable Deep Reinforcement Learning (XRL) addresses these challenges by enhancing transparency through feature-level, state-level, dataset-level, and model-level explanation techniques. This survey provides a comprehensive review of XRL methods, evaluates their qualitative and quantitative assessment frameworks, and explores their role in policy refinement, adversarial robustness, and security. Additionally, we examine the integration of reinforcement learning with Large Language Models (LLMs), particularly through Reinforcement Learning from Human Feedback (RLHF), which optimizes AI alignment with human preferences. We conclude by highlighting open research challenges and future directions to advance the development of interpretable, reliable, and accountable DRL systems.
Problem

Research questions and friction points this paper is trying to address.

Enhance transparency in Deep Reinforcement Learning
Improve interpretability and trust in DRL systems
Integrate RL with Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explainable Deep Reinforcement Learning
Integration with Large Language Models
Reinforcement Learning from Human Feedback
🔎 Similar Papers
No similar papers found.