🤖 AI Summary
This study addresses the opacity of machine learning models and the difficulty domain experts face in utilizing explanation tools by proposing a reward-driven multi-agent framework. Within this architecture, a Proposer agent dynamically selects configuration tools to overcome the limitations of single-tool approaches, while an Actor agent performs end-to-end optimization via perturbation-based fidelity evaluation and modality-adaptive penalty mechanisms to generate natural language explanations. By integrating multi-agent systems, reinforcement learning, and large language models, the framework enables effective cross-modal processing. Experimental results demonstrate that the proposed method improves explanation fidelity over baselines by 28%, 21%, and 34% on tabular, textual, and visual tasks, respectively, significantly enhancing overall model interpretability.
📝 Abstract
Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools. To this end, we present MEA, a multi-agent framework that removes the explanation knowledge barrier entirely: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent is optimized end-to-end against faithfulness, transforming the outputs into natural language explanations grounded in model behavior across tabular, text, and vision modalities. Further, we introduce diverse question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. We find that frontier LLMs systematically produce unfaithful explanations. By optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with reward-driven optimization yielding faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. More broadly, our findings suggest that AI agents themselves can serve as a scalable, adaptable interface to ML explainability, opening a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field.