interpret learned policies

Designs, builds, and analyzes learned decision policies with an emphasis on human interpretability by producing sparse, explainable representations (e.g., lasso-based or other simple model forms) and extracting explicit rules or features that drive decisions. Develops methods to learn interpretable policies, explain why they outperform heuristics, identify decision triggers and patterns, and validate the robustness of those explanations across different configurations.

interpretlearnedpolicies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on Explainable Deep Reinforcement Learning

Feb 08, 2025
ZC
Zelei Cheng
🏛️ Northwestern University

Deep reinforcement learning (DRL) excels in sequential decision-making but suffers from limited interpretability and trustworthiness in high-stakes applications due to its black-box nature. To address this, we present a systematic survey of eXplainable Reinforcement Learning (XRL) and propose a novel, unified four-level interpretability framework—spanning feature-, state-, dataset-, and model-level explanations. We further introduce a cross-granularity evaluation体系 integrating attention mechanisms, saliency mapping, counterfactual reasoning, surrogate modeling, attribution analysis, and Reinforcement Learning from Human Feedback (RLHF). Additionally, we establish a taxonomy and applicability boundary matrix for XRL methods. Empirical evaluations demonstrate that our framework significantly enhances policy generalization, adversarial robustness, safety guarantees, and alignment with human preferences—thereby establishing an explanation-driven paradigm for trustworthy AI deployment.

Enhance transparency in Deep Reinforcement LearningImprove interpretability and trust in DRL systemsIntegrate RL with Large Language Models

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether internal representations of a learned decision-making system encode interpretable concepts that explain its behavior. In the Game of Hidden Rules—a setting without explicit rule labels—the authors train a tokenized autoregressive Transformer agent and, for the first time, apply sparse autoencoders (SAEs) to extract structured features from its decision embeddings. The results demonstrate that individual SAE activation dimensions selectively correspond to fundamental concepts such as shapes and buckets, accounting for the vast majority of relevant decisions. Moreover, these representations reveal interpretable exploratory actions and feedback-driven strategies for switching between hidden rules, offering novel evidence for the interpretability of unsupervised rule-inference agents.

concept discoverydecision-making agentsGame of Hidden Rules

Learning Ensembles of Interpretable Simple Structure

Feb 26, 2025
GA
Gaurav Arwade
🏛️ Iowa State University

High-precision machine learning models (e.g., XGBoost, neural networks) suffer from poor interpretability in critical domains such as operations research. Method: This paper proposes a bottom-up “simple structure” identification framework that automatically partitions the data into subpopulations where feature interactions are significantly weakened, then fits lightweight interpretable models (e.g., shallow decision trees) within each subpopulation. Contribution/Results: We formally define “simple structure” for the first time and develop an end-to-end pipeline integrating data-partition ensembles, subgroup-adaptive modeling, and quantitative evaluation of decision-boundary interpretability. Experiments on synthetic data demonstrate that our approach matches XGBoost’s predictive accuracy while substantially enhancing both local interpretability and global transparency; moreover, its decision boundaries align more closely with domain intuition, overcoming the performance–interpretability trade-off that limits conventional globally interpretable models in settings with complex feature interactions.

Enhances interpretability in complex machine learning models.Identifies simple structures within complex feature interactions.Improves decision support with transparent and accurate models.

Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs

Mar 11, 2025
HK
Hector Kohler
🏛️ Université de Lille | Inria | TU Darmstadt

In high-reliability domains (e.g., healthcare), quantifying the interpretability of reinforcement learning (RL) policies remains challenging due to the lack of objective evaluation criteria and heavy reliance on costly human assessments. Method: We propose the first fully automated, human-free interpretability evaluation paradigm, built upon a simulatability-based empirical framework that integrates program distillation, imitation learning, and symbolic program generation, complemented by computationally tractable interpretability metrics. Contributions/Results: (1) The first scalable, human-free quantitative assessment of RL policy interpretability; (2) Empirical evidence that interpretability and task performance are non-negatively correlated—and synergistically improved in certain settings; (3) Refutation of the existence of a universally optimal policy class across tasks; (4) Strong agreement between automated evaluations and user studies, with all evaluation protocols and baseline code publicly released.

Comparing interpretability and performance trade-offs across different policy classes.Developing a methodology to assess policy interpretability using simulatability proxies.Evaluating interpretability of reinforcement learning policies without human studies.

Investigating the Duality of Interpretability and Explainability in Machine Learning

Oct 28, 2024
MG
Moncef Garouani
🏛️ Université Toulouse Capitole | Université de Toulouse | Aix-Marseille University

This work addresses the fundamental trade-off in machine learning between high predictive performance and low interpretability inherent in “black-box” models (e.g., deep neural networks, ensemble methods). It rigorously distinguishes post-hoc explanation—applied after model training—from inherently interpretable modeling—designed for transparency from inception. To reconcile accuracy and interpretability, we propose a hybrid modeling paradigm centered on symbolic knowledge embedding, integrating differentiable symbolic modules, knowledge distillation, and symbolic reasoning into the model architecture itself. This enables joint optimization of fidelity and interpretability at the design stage. Extensive experiments across diverse domains demonstrate that our approach matches the predictive accuracy of state-of-the-art black-box models while generating human-understandable, logically grounded decision rules. As a result, it substantially enhances model trustworthiness and deployment viability in safety- and accountability-critical applications.

Addressing the need for transparent and trustworthy machine learning modelsClarifying the difference between explaining black box models and using inherently interpretable onesEvaluating hybrid methods combining symbolic knowledge with neural networks for interpretability

Rule Generation for Classification: Scalability, Interpretability, and Fairness

Apr 21, 2021
TE
Tabea E. Rober
🏛️ University of Amsterdam | Erasmus University Rotterdam

This paper addresses the challenge of simultaneously achieving scalability, local interpretability, and multi-attribute, multi-class fairness in rule-based classification models. To this end, we propose the first column generation–based rule learning framework. Methodologically, we introduce column generation—previously unexplored in rule learning—integrating a linear programming master problem, a decision-tree–inspired column generation heuristic, a surrogate pricing subproblem solver, and weighted rule optimization; we further formulate generalized fairness constraints supporting multiple sensitive attributes and multi-class outcomes. Our key contributions are: (1) enabling local interpretability via rule weights, and (2) unifying support for complex fairness constraints and scalable search over large rule spaces. Extensive experiments on benchmark datasets demonstrate that our approach achieves significant trade-off improvements among accuracy, interpretability, and fairness, substantially enhancing the practicality of rule models in real-world, large-scale applications.

Balancing accuracy with interpretability and fairnessEnsuring interpretability and fairness in rule learningScaling rule-based classification to large datasets

Latest Papers

What's happening recently
View more

This work addresses the limited interpretability of reinforcement learning policies, which undermines their trustworthiness, and the excessive complexity of decision rules produced by existing policy-to-tree conversion methods. The authors propose a structure- and usage-aware pruning framework that transforms trained policies into compact, auditable decision trees, significantly enhancing interpretability while preserving high task performance. The approach introduces policy re-execution evaluation and proxy metrics for interpretability, systematically uncovering viable pathways from complex policies to concise rule sets. Empirical validation on classic control and MuJoCo benchmarks demonstrates the method’s effectiveness: interpretability consistently improves throughout the pruning process with negligible loss in return.

decision-tree pruninginterpretable reinforcement learningpolicy transparency

Balancing Interpretability and Performance in Reinforcement Learning: An Adaptive Spectral Based Linear Approach

Oct 04, 2025
QY
Qianxin Yi
🏛️ Xi'an Jiaotong University | Hong Kong Baptist University

This work addresses the fundamental challenge in reinforcement learning (RL) of simultaneously achieving high interpretability and strong policy performance. We propose an adaptive linear method grounded in spectral filtering, which theoretically analyzes regularization’s role in the bias–variance trade-off via spectral-domain analysis and designs a data-driven spectral filter to dynamically select optimal regularization strength. Built upon ridge regression, our approach generalizes to a unified spectral-domain linear framework for both policy evaluation and optimization, and integrates a quantifiable interpretability analysis module. Evaluated on synthetic benchmarks as well as real-world industrial environments—specifically Kwai and Taobao—the method maintains near-optimal policy performance while substantially improving decision transparency and generalization robustness. To our knowledge, this is the first RL approach that rigorously unifies high decision quality with formal interpretability guarantees under a theoretically sound framework.

Achieving theoretical guarantees while enhancing practical decision qualityBalancing interpretability and performance in reinforcement learningDeveloping adaptive spectral linear method with regularization control

Fairness-Aware and Interpretable Policy Learning

Sep 15, 2025
NB
Nora Bearth
🏛️ University of St.Gallen | Örebro University | CEPR | CESIfo | IAB | IZA

Algorithmic decision-making often faces a trade-off between fairness and interpretability. Method: This paper proposes a synergistic optimization framework that integrates sensitive-attribute decorrelation preprocessing with interpretable policy trees. It introduces a novel feature-space inverse transformation mechanism that mitigates the influence of sensitive attributes on decisions while preserving original feature semantics—ensuring policy transparency—and enhances fairness and prediction stability through structural tree optimization. Contribution/Results: Evaluated on Swiss labor market policy allocation, the method significantly improves group-level fairness—e.g., statistical parity increases by 23%—while incurring only a marginal reduction in employment rate (<1.5%). These results demonstrate its effectiveness and practical viability in real-world policy deployment.

Enhancing policy allocations without sacrificing interpretabilityIntegrating fairness and interpretability into algorithmic decision makingRemoving dependencies between sensitive attributes and decision features

Automatically Finding Rule-Based Neurons in OthelloGPT

Oct 28, 2025
AS
Aditya Singh
🏛️ University of Chicago | Carnegie Mellon University | Independent

This paper addresses the problem of automatic identification and interpretability modeling of rule-based MLP neurons in OthelloGPT. We propose a neuron behavior decomposition method based on regression decision trees, modeling neuron activation as piecewise logical rules over board states—thereby mapping black-box activations to human-readable strategic patterns. Our approach innovatively integrates decision tree fitting, causal intervention, and ablation analysis to rigorously verify the causal role of individual neurons in predicting specific moves. Experiments show that 913 out of 2,048 neurons (44.6%) in layer 5 exhibit high-fidelity rule-based approximations (R² > 0.7). Targeted ablation of key neurons degrades corresponding move prediction accuracy by 5–10×, confirming their functional specificity and interpretability.

Automatically identifying rule-based neurons in OthelloGPT using decision treesConverting neuron activations into human-readable logical game rulesVerifying causal relevance of discovered patterns through targeted neuron ablation

This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.

causal machine learningcausal relationshipsdecision support

Hot Scholars

MK

Minwu Kim

New York University Abu Dhabi
Artificial IntelligenceMachine Learning
AR

Abbas Rahimi

Research Staff Member, IBM Research-Zurich
Machine ReasoningNeurosymbolic AIAI HardwareHW/SW Codesign
GD

Gregory Duthé

Postdoctoral Researcher, ETH Zürich
wind energygeometric deep learninggraph neural networksfluid mechanics
XZ

Xingchen Zou

The University of Hong Kong
LLM/VLM AgentAI4ScienceUrban Computing
TL

Tao Long

Columbia University
HCIhuman-AI interaction{crea.produc}tivity supports :)