Interpretability by Design for Efficient Multi-Objective Reinforcement Learning

📅 2025-06-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multi-objective reinforcement learning (MORL) suffers from non-unique mappings between policy parameter space and multi-objective performance space, poor interpretability, and low efficiency in Pareto frontier search. To address these challenges, this paper proposes an interpretable MORL framework based on local linear mapping—first embedding bidirectional parameter–performance interpretability into algorithm design. Specifically, it models the parameter-to-performance mapping locally as linear, enabling real-time semantic interpretation of policy objectives; supports zero-shot cross-domain policy transfer without retraining; and facilitates interpretable gradient-guided approximation of the Pareto frontier. Evaluated on multiple benchmark tasks, our method achieves significant improvements: +18.7% in Pareto frontier coverage and 2.3× acceleration in convergence speed, while outperforming state-of-the-art methods in both explanation quality and search efficiency.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Reinforcement LearningMultiagent Systems: Multiagent Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Multi-objective reinforcement learning (MORL) aims at optimising several, often conflicting goals in order to improve flexibility and reliability of RL in practical tasks. This can be achieved by finding diverse policies that are optimal for some objective preferences and non-dominated by optimal policies for other preferences so that they form a Pareto front in the multi-objective performance space. The relation between the multi-objective performance space and the parameter space that represents the policies is generally non-unique. Using a training scheme that is based on a locally linear map between the parameter space and the performance space, we show that an approximate Pareto front can provide an interpretation of the current parameter vectors in terms of the objectives which enables an effective search within contiguous solution domains. Experiments are conducted with and without retraining across different domains, and the comparison with previous methods demonstrates the efficiency of our approach.
Problem

Research questions and friction points this paper is trying to address.

Optimizing conflicting goals in multi-objective reinforcement learning
Finding diverse policies forming Pareto front in performance space
Enabling effective search using locally linear parameter-performance mapping
Innovation

Methods, ideas, or system contributions that make the work stand out.

Locally linear map between parameter and performance spaces
Approximate Pareto front for parameter interpretation
Effective search within contiguous solution domains
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qiyue Xia
School of Informatics, University of Edinburgh, UK
J
J. M. Herrmann
School of Informatics, University of Edinburgh, UK