🤖 AI Summary
This work addresses the limited interpretability of reinforcement learning policies, which undermines their trustworthiness, and the excessive complexity of decision rules produced by existing policy-to-tree conversion methods. The authors propose a structure- and usage-aware pruning framework that transforms trained policies into compact, auditable decision trees, significantly enhancing interpretability while preserving high task performance. The approach introduces policy re-execution evaluation and proxy metrics for interpretability, systematically uncovering viable pathways from complex policies to concise rule sets. Empirical validation on classic control and MuJoCo benchmarks demonstrates the method’s effectiveness: interpretability consistently improves throughout the pruning process with negligible loss in return.
📝 Abstract
Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. Converting a trained policy into explicit decision-tree rules improves transparency and the resulting artifacts often remain too complex for human understanding. We present a pruning process that simplifies such rule-based policies while preserving task performance and making edits to the policy auditable. The process defines a small set of structural and usage-aware operators and evaluates candidate edits by re-executing the policy to measure return and interpretability proxies. This exposes an transformation process from complex to compact policy structures. We investigate this approach on classic control and MuJoCo benchmarks, where pruning traces reveal consistent interpretability improvements while maintaining high performance.