Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate

📅 2025-09-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates how model capability heterogeneity in multi-agent debate induces reasoning failures and accuracy degradation. We construct a controlled experimental framework to systematically vary the scale and capability of LLMs participating in debates, enabling fine-grained tracking of reasoning propagation and answer evolution. Our study reveals, for the first time, that even when strong models are present, heterogeneous capability configurations frequently cause consensus preference to override error-correction motivation—leading correct answers to drift toward incorrect ones. Moreover, persuasive yet fallacious reasoning is widely adopted in the absence of alignment incentives and anti-misinformation mechanisms. The core contribution is the identification of “consensus-driven degradation”—a novel failure mode—demonstrating that capability diversity alone does not ensure robustness. We further establish that improving multi-agent reasoning reliability requires co-designing incentive structures and anti-misinformation capabilities. (149 words)

Technology Category

Multiagent Systems: Agreement, Argumentation & NegotiationKnowledge Representation and Reasoning: ArgumentationCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. The prior work has exclusively focused on debates within homogeneous groups of agents, whereas we explore how diversity in model capabilities influences the dynamics and outcomes of multi-agent interactions. Through a series of experiments, we demonstrate that debate can lead to a decrease in accuracy over time -- even in settings where stronger (i.e., more capable) models outnumber their weaker counterparts. Our analysis reveals that models frequently shift from correct to incorrect answers in response to peer reasoning, favoring agreement over challenging flawed reasoning. These results highlight important failure modes in the exchange of reasons during multi-agent debate, suggesting that naive applications of debate may cause performance degradation when agents are neither incentivized nor adequately equipped to resist persuasive but incorrect reasoning.
Problem

Research questions and friction points this paper is trying to address.

Debate harms accuracy in multi-agent reasoning systems
Diverse model capabilities negatively impact debate outcomes
Agents favor agreement over correcting flawed reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explores diversity in model capabilities
Debate decreases accuracy over time
Models shift from correct to incorrect answers
A
Andrea Wynn
Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA
Harsh Satija
Harsh Satija
Vector Institute
Artificial IntelligenceMachine LearningReinforcement Learning
G
Gillian Hadfield
University of Toronto, Toronto, ON, Canada