Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in multi-agent systems where communication gains, architectural advantages, and additional inference compute are inherently coupled, with final accuracy obscuring the underlying details of decision revision. To resolve this, we propose the ICR framework, which reconceptualizes communication evaluation as a selective "independent reasoning–communication revision" process. By employing text and latent-space auditing alongside message-free control experiments, our approach effectively disentangles the contribution of communication benefits from that of merely increasing inference steps. Our findings reveal a double-edged sword effect wherein richer messages amplify both beneficial and detrimental influences, and demonstrate that reception strategies significantly alter revision behaviors, challenging the prevailing view of channel quality as an intrinsic property. Ultimately, this work establishes a unified evaluation benchmark for multi-agent communication.
📝 Abstract
Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.
Problem

Research questions and friction points this paper is trying to address.

Multi-agent communication
LLM
Evaluation ambiguity
Final accuracy
Communication auditing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Communication
ICR Framework
Communication Auditing
Selective Revision
LLM