Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning

📅 2025-11-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the poor interpretability of deep reinforcement learning (DRL) policies, this paper proposes a model-agnostic interpretability distillation framework. First, it adaptively partitions the state space using Voronoi diagrams; then, it distills locally linear policies on each partition. Unlike prior approaches, it imposes no structural assumptions, thus balancing interpretability and representational capacity while preserving policy transparency and achieving performance alignment—or even improvement. The key innovation lies in coupling geometric partitioning with knowledge distillation, endowing local linear models with both theoretical traceability and empirical effectiveness. Experiments on Gridworld and classical control benchmarks demonstrate that the distilled policies exhibit clear decision logic—e.g., piecewise-linear control laws—and achieve average performance gains of 1.2%–3.7% over the original DRL policies, significantly outperforming existing interpretable baselines.

Technology Category

Machine Learning: Reinforcement LearningComputer Vision: Interpretability, Explainability, and TransparencySearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Deep Reinforcement Learning is one of the state-of-the-art methods for producing near-optimal system controllers. However, deep RL algorithms train a deep neural network, that lacks transparency, which poses challenges when the controller has to meet regulations, or foster trust. To alleviate this, one could transfer the learned behaviour into a model that is human-readable by design using knowledge distilla- tion. Often this is done with a single model which mimics the original model on average but could struggle in more dynamic situations. A key challenge is that this simpler model should have the right balance be- tween flexibility and complexity or right balance between balance bias and accuracy. We propose a new model-agnostic method to divide the state space into regions where a simplified, human-understandable model can operate in. In this paper, we use Voronoi partitioning to find regions where linear models can achieve similar performance to the original con- troller. We evaluate our approach on a gridworld environment and a classic control task. We observe that our proposed distillation to locally- specialized linear models produces policies that are explainable and show that the distillation matches or even slightly outperforms the black-box policy they are distilled from.
Problem

Research questions and friction points this paper is trying to address.

Creating explainable RL policies from opaque deep neural networks
Balancing model simplicity and accuracy in policy distillation
Partitioning state space for specialized linear models using Voronoi
Innovation

Methods, ideas, or system contributions that make the work stand out.

Voronoi partitioning divides state space
Locally-specialized linear models replace black-box
Model-agnostic distillation enhances explainability
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Senne Deproost
Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium
D
Dennis Steckelmacher
Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium
A
Ann Nowé
Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium