MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the opacity of machine learning models and the difficulty domain experts face in utilizing explanation tools by proposing a reward-driven multi-agent framework. Within this architecture, a Proposer agent dynamically selects configuration tools to overcome the limitations of single-tool approaches, while an Actor agent performs end-to-end optimization via perturbation-based fidelity evaluation and modality-adaptive penalty mechanisms to generate natural language explanations. By integrating multi-agent systems, reinforcement learning, and large language models, the framework enables effective cross-modal processing. Experimental results demonstrate that the proposed method improves explanation fidelity over baselines by 28%, 21%, and 34% on tabular, textual, and visual tasks, respectively, significantly enhancing overall model interpretability.
📝 Abstract
Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools. To this end, we present MEA, a multi-agent framework that removes the explanation knowledge barrier entirely: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent is optimized end-to-end against faithfulness, transforming the outputs into natural language explanations grounded in model behavior across tabular, text, and vision modalities. Further, we introduce diverse question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. We find that frontier LLMs systematically produce unfaithful explanations. By optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with reward-driven optimization yielding faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. More broadly, our findings suggest that AI agents themselves can serve as a scalable, adaptable interface to ML explainability, opening a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field.
Problem

Research questions and friction points this paper is trying to address.

Model Explainability
Faithful Explanations
Multi-Agent System
Post-hoc Explanation
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent System
Faithful Explanations
Reward-Driven Optimization
Perturbation-based Metric
Cross-modal Explainability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuyang Cheng
University of Virginia
R
Raghav Kaushik Ravi
Vellore Institute of Technology
S
Srivarshinee Sridhar
Vellore Institute of Technology
S
Sriparna Saha
Indian Institute of Technology, Patna
A
Akash Ghosh
Indian Institute of Technology, Patna
Chirag Agarwal
Chirag Agarwal
Assistant Professor, UVA
XAITrustworthyMLArtificial Intelligence