Score
Designs, builds, and analyzes learned decision policies with an emphasis on human interpretability by producing sparse, explainable representations (e.g., lasso-based or other simple model forms) and extracting explicit rules or features that drive decisions. Develops methods to learn interpretable policies, explain why they outperform heuristics, identify decision triggers and patterns, and validate the robustness of those explanations across different configurations.
To address the opacity of deep neural network–based decision-making in reinforcement learning (RL) agents, this paper presents a systematic survey of explainable RL (XRL). We propose a dual-axis taxonomy—“what to explain” (target dimension: e.g., policy, value function, trajectory) and “how to explain” (method dimension: e.g., surrogate modeling, attention visualization, counterfactual generation)—to structurally categorize over 250 XRL studies. This framework unifies disparate classification logics in prior work and exposes critical gaps in adaptability to dynamic environments, human interpretability, and real-time explanation capability. Our analysis identifies three pressing research directions: (1) multi-granularity explanation integration, (2) human-in-the-loop evaluation mechanisms, and (3) lightweight, trustworthy explanation paradigms tailored to real-world RL applications—including robotic control and autonomous driving. The survey thus provides both theoretical foundations and practical guidelines for advancing XRL research and deployment.
Deep reinforcement learning (DRL) excels in sequential decision-making but suffers from limited interpretability and trustworthiness in high-stakes applications due to its black-box nature. To address this, we present a systematic survey of eXplainable Reinforcement Learning (XRL) and propose a novel, unified four-level interpretability framework—spanning feature-, state-, dataset-, and model-level explanations. We further introduce a cross-granularity evaluation体系 integrating attention mechanisms, saliency mapping, counterfactual reasoning, surrogate modeling, attribution analysis, and Reinforcement Learning from Human Feedback (RLHF). Additionally, we establish a taxonomy and applicability boundary matrix for XRL methods. Empirical evaluations demonstrate that our framework significantly enhances policy generalization, adversarial robustness, safety guarantees, and alignment with human preferences—thereby establishing an explanation-driven paradigm for trustworthy AI deployment.
This study investigates whether internal representations of a learned decision-making system encode interpretable concepts that explain its behavior. In the Game of Hidden Rules—a setting without explicit rule labels—the authors train a tokenized autoregressive Transformer agent and, for the first time, apply sparse autoencoders (SAEs) to extract structured features from its decision embeddings. The results demonstrate that individual SAE activation dimensions selectively correspond to fundamental concepts such as shapes and buckets, accounting for the vast majority of relevant decisions. Moreover, these representations reveal interpretable exploratory actions and feedback-driven strategies for switching between hidden rules, offering novel evidence for the interpretability of unsupervised rule-inference agents.
High-precision machine learning models (e.g., XGBoost, neural networks) suffer from poor interpretability in critical domains such as operations research. Method: This paper proposes a bottom-up “simple structure” identification framework that automatically partitions the data into subpopulations where feature interactions are significantly weakened, then fits lightweight interpretable models (e.g., shallow decision trees) within each subpopulation. Contribution/Results: We formally define “simple structure” for the first time and develop an end-to-end pipeline integrating data-partition ensembles, subgroup-adaptive modeling, and quantitative evaluation of decision-boundary interpretability. Experiments on synthetic data demonstrate that our approach matches XGBoost’s predictive accuracy while substantially enhancing both local interpretability and global transparency; moreover, its decision boundaries align more closely with domain intuition, overcoming the performance–interpretability trade-off that limits conventional globally interpretable models in settings with complex feature interactions.
In high-reliability domains (e.g., healthcare), quantifying the interpretability of reinforcement learning (RL) policies remains challenging due to the lack of objective evaluation criteria and heavy reliance on costly human assessments. Method: We propose the first fully automated, human-free interpretability evaluation paradigm, built upon a simulatability-based empirical framework that integrates program distillation, imitation learning, and symbolic program generation, complemented by computationally tractable interpretability metrics. Contributions/Results: (1) The first scalable, human-free quantitative assessment of RL policy interpretability; (2) Empirical evidence that interpretability and task performance are non-negatively correlated—and synergistically improved in certain settings; (3) Refutation of the existence of a universally optimal policy class across tasks; (4) Strong agreement between automated evaluations and user studies, with all evaluation protocols and baseline code publicly released.
This work addresses the fundamental trade-off in machine learning between high predictive performance and low interpretability inherent in “black-box” models (e.g., deep neural networks, ensemble methods). It rigorously distinguishes post-hoc explanation—applied after model training—from inherently interpretable modeling—designed for transparency from inception. To reconcile accuracy and interpretability, we propose a hybrid modeling paradigm centered on symbolic knowledge embedding, integrating differentiable symbolic modules, knowledge distillation, and symbolic reasoning into the model architecture itself. This enables joint optimization of fidelity and interpretability at the design stage. Extensive experiments across diverse domains demonstrate that our approach matches the predictive accuracy of state-of-the-art black-box models while generating human-understandable, logically grounded decision rules. As a result, it substantially enhances model trustworthiness and deployment viability in safety- and accountability-critical applications.
This paper addresses the challenge of simultaneously achieving scalability, local interpretability, and multi-attribute, multi-class fairness in rule-based classification models. To this end, we propose the first column generation–based rule learning framework. Methodologically, we introduce column generation—previously unexplored in rule learning—integrating a linear programming master problem, a decision-tree–inspired column generation heuristic, a surrogate pricing subproblem solver, and weighted rule optimization; we further formulate generalized fairness constraints supporting multiple sensitive attributes and multi-class outcomes. Our key contributions are: (1) enabling local interpretability via rule weights, and (2) unifying support for complex fairness constraints and scalable search over large rule spaces. Extensive experiments on benchmark datasets demonstrate that our approach achieves significant trade-off improvements among accuracy, interpretability, and fairness, substantially enhancing the practicality of rule models in real-world, large-scale applications.
This work addresses the limited interpretability of reinforcement learning policies, which undermines their trustworthiness, and the excessive complexity of decision rules produced by existing policy-to-tree conversion methods. The authors propose a structure- and usage-aware pruning framework that transforms trained policies into compact, auditable decision trees, significantly enhancing interpretability while preserving high task performance. The approach introduces policy re-execution evaluation and proxy metrics for interpretability, systematically uncovering viable pathways from complex policies to concise rule sets. Empirical validation on classic control and MuJoCo benchmarks demonstrates the method’s effectiveness: interpretability consistently improves throughout the pruning process with negligible loss in return.
This work addresses the fundamental challenge in reinforcement learning (RL) of simultaneously achieving high interpretability and strong policy performance. We propose an adaptive linear method grounded in spectral filtering, which theoretically analyzes regularization’s role in the bias–variance trade-off via spectral-domain analysis and designs a data-driven spectral filter to dynamically select optimal regularization strength. Built upon ridge regression, our approach generalizes to a unified spectral-domain linear framework for both policy evaluation and optimization, and integrates a quantifiable interpretability analysis module. Evaluated on synthetic benchmarks as well as real-world industrial environments—specifically Kwai and Taobao—the method maintains near-optimal policy performance while substantially improving decision transparency and generalization robustness. To our knowledge, this is the first RL approach that rigorously unifies high decision quality with formal interpretability guarantees under a theoretically sound framework.
Algorithmic decision-making often faces a trade-off between fairness and interpretability. Method: This paper proposes a synergistic optimization framework that integrates sensitive-attribute decorrelation preprocessing with interpretable policy trees. It introduces a novel feature-space inverse transformation mechanism that mitigates the influence of sensitive attributes on decisions while preserving original feature semantics—ensuring policy transparency—and enhances fairness and prediction stability through structural tree optimization. Contribution/Results: Evaluated on Swiss labor market policy allocation, the method significantly improves group-level fairness—e.g., statistical parity increases by 23%—while incurring only a marginal reduction in employment rate (<1.5%). These results demonstrate its effectiveness and practical viability in real-world policy deployment.
This paper addresses the problem of automatic identification and interpretability modeling of rule-based MLP neurons in OthelloGPT. We propose a neuron behavior decomposition method based on regression decision trees, modeling neuron activation as piecewise logical rules over board states—thereby mapping black-box activations to human-readable strategic patterns. Our approach innovatively integrates decision tree fitting, causal intervention, and ablation analysis to rigorously verify the causal role of individual neurons in predicting specific moves. Experiments show that 913 out of 2,048 neurons (44.6%) in layer 5 exhibit high-fidelity rule-based approximations (R² > 0.7). Targeted ablation of key neurons degrades corresponding move prediction accuracy by 5–10×, confirming their functional specificity and interpretability.
This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.