🤖 AI Summary
This study addresses the high computational overhead and low inference efficiency in medical multimodal large language models arising from redundant communication topologies during multi-agent collaboration. To this end, we propose MedPrune, a framework that formulates the diagnostic process as a heterogeneous communication graph. Specifically, it introduces the first reinforcement learning-based node sparsification mechanism and designs an edge sparsification strategy that jointly optimizes task performance and topological complexity, thereby enabling dynamic, adaptive pruning of collaborative topologies. Experimental results demonstrate that MedPrune significantly outperforms existing baselines under both full-set and few-shot settings, substantially improving token efficiency while exhibiting strong adversarial robustness.
📝 Abstract
While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.