🤖 AI Summary
Deep reinforcement learning (DRL) excels in sequential decision-making but suffers from limited interpretability and trustworthiness in high-stakes applications due to its black-box nature. To address this, we present a systematic survey of eXplainable Reinforcement Learning (XRL) and propose a novel, unified four-level interpretability framework—spanning feature-, state-, dataset-, and model-level explanations. We further introduce a cross-granularity evaluation体系 integrating attention mechanisms, saliency mapping, counterfactual reasoning, surrogate modeling, attribution analysis, and Reinforcement Learning from Human Feedback (RLHF). Additionally, we establish a taxonomy and applicability boundary matrix for XRL methods. Empirical evaluations demonstrate that our framework significantly enhances policy generalization, adversarial robustness, safety guarantees, and alignment with human preferences—thereby establishing an explanation-driven paradigm for trustworthy AI deployment.
📝 Abstract
Deep Reinforcement Learning (DRL) has achieved remarkable success in sequential decision-making tasks across diverse domains, yet its reliance on black-box neural architectures hinders interpretability, trust, and deployment in high-stakes applications. Explainable Deep Reinforcement Learning (XRL) addresses these challenges by enhancing transparency through feature-level, state-level, dataset-level, and model-level explanation techniques. This survey provides a comprehensive review of XRL methods, evaluates their qualitative and quantitative assessment frameworks, and explores their role in policy refinement, adversarial robustness, and security. Additionally, we examine the integration of reinforcement learning with Large Language Models (LLMs), particularly through Reinforcement Learning from Human Feedback (RLHF), which optimizes AI alignment with human preferences. We conclude by highlighting open research challenges and future directions to advance the development of interpretable, reliable, and accountable DRL systems.