Score
Design, build, and evaluate deliberative multi-agent systems in which multiple reasoning agents exchange proposals, critiques, and votes to produce a single solution or decision. This includes specifying interaction and communication protocols, aggregation and consensus mechanisms, role assignment and conflict‑resolution policies, and metrics for convergence and quality to ensure the multi‑agent deliberation outperforms or complements monolithic models and diversifies reasoning strategies.
This study addresses the challenges in multi-agent debate (MAD) research stemming from inconsistent terminology and the absence of a systematic design framework, which hinder cross-study comparisons and obscure implicit design choices. Through a systematic literature review of 141 MAD works, this paper proposes the first three-dimensional taxonomy encompassing participants, interaction mechanisms, and consensus protocols, and introduces formal notation to structurally describe MAD configurations. The resulting classification reveals prevalent yet empirically unvalidated design paradigms in the field, identifies critical limitations, and lays the groundwork for establishing controllable benchmarks and enabling automated tuning. Furthermore, it advocates for evolving the taxonomy into an executable, machine-readable specification framework to standardize and advance MAD research.
Multi-agent large language models (MA-LLMs) lack systematic performance benchmarking, hindering principled design and deployment. Method: We propose the first problem-solving–oriented analytical framework for MA-LLM influencing factors, integrating multi-agent systems theory, prompt engineering, distributed negotiation modeling, and empirical attribution analysis to systematically decouple three core variables: role design, communication topology, and decision-making mechanisms. Contribution/Results: We (i) uncover coupling patterns among scalability, communication structure, and decision paths; (ii) rigorously characterize the performance advantage boundary of MA-LLMs over single-agent LLMs; (iii) identify three pervasive bottlenecks—excessive communication overhead, role homogenization, and consensus drift; and (iv) derive actionable architectural optimizations that jointly enhance computational efficiency and collaborative robustness. This work establishes foundational insights for principled MA-LLM development and deployment.
This work addresses the limitations of existing large language model–based multi-agent systems, which lack structured deliberation mechanisms and struggle to produce accountable decisions while preserving dissent. The authors propose the Deliberative Collective Intelligence (DCI) framework—the first formal computational model of human-like deliberation—featuring four reasoning roles, 14 typified cognitive behaviors, a shared workspace, and the DCI-CF convergence-flow algorithm. This framework enables phased collaborative reasoning and generates structured decision packages comprising the selected option, residual disagreements, minority reports, and restart conditions. Experiments using Gemini 2.5 Flash demonstrate that DCI significantly outperforms unstructured debate by +0.95 on 40 unconventional tasks, achieves a hidden information integration score of 9.56, produces fully structured decision packages in 100% of cases (with minority reports in 98%), at a computational cost approximately 62 times that of a single agent.
This work addresses the challenge of multi-agent negotiation and collaboration in partially observable environments with asymmetric observations, where current large language models (LLMs) exhibit limited performance. The study formalizes negotiation-based collaboration as a joint decision-making problem, introduces a cross-domain and scalable evaluation benchmark, and systematically assesses LLMs’ capabilities in tasks requiring information exchange to achieve shared rewards. By designing negotiation protocols, integrating external tools, and explicitly modeling partial observability, the research reveals significant shortcomings of LLMs in negotiation alignment and complex reasoning. Nevertheless, it also uncovers that negotiation mechanisms possess reflective and error-correction potential, which, in certain scenarios, can enhance performance—even surpassing centralized baselines.
This study investigates the selection and evaluation of decision-making protocols in multi-agent debate. Using controlled experiments within a unified multi-agent debate framework, we systematically assess seven classical decision protocols—including majority voting and consensus—on knowledge-intensive tasks (MMLU, GPQA) and reasoning tasks (StrategyQA, MuSR). We propose two novel mechanisms: All-Agents Drafting (AAD), which enhances answer diversity, and Collective Improvement (CI), which mitigates limitations inherent to single-protocol reliance. Empirical results show that voting-based protocols improve performance by 13.2% on reasoning tasks, while consensus-based protocols yield a 2.8% gain on knowledge tasks. AAD and CI achieve up to 3.3% and 7.4% absolute accuracy improvements, respectively. Our findings establish decision mechanisms as a critical determinant of multi-agent collaborative efficacy, providing both empirical grounding and a new methodology for adaptive protocol selection.
This work addresses the limitations of single large language models in complex legal reasoning—specifically their narrow perspective and constrained deliberative capacity—by proposing a multi-agent collaborative deliberation framework inspired by courtroom procedures and legal argumentation theory. The framework introduces two novel legally grounded interaction heuristics that simulate adversarial debates among agents representing diverse viewpoints, thereby fostering multi-perspective critical reasoning. Experimental results demonstrate that, while overall performance remains comparable to baseline models, the proposed approach significantly outperforms existing methods on cases requiring nuanced analysis from multiple legal standpoints, particularly excelling in resolving complex legal problems where baseline approaches fail.
This study addresses the tendency of existing multi-agent policy simulations—based on homogeneous large language models—to generate artificial consensus that fails to capture genuine value disagreements. To overcome this limitation, the authors propose the AI Council, a three-stage deliberation framework that assigns heterogeneous 7–9B parameter models distinct value perspectives and incorporates a state-of-the-art model for consistency validation. Their approach demonstrates, for the first time, that architectural heterogeneity significantly reduces policy choice concentration (e.g., from 70.9% to 46.1% in child welfare and from 46.0% to 22.9% in housing policy, p<0.001). The work further identifies a pervasive trade-off between fidelity and diversity under consistency validation and introduces the “credible tension ratio” as a novel metric to evaluate deliberative capacity in smaller models.
This work addresses the tendency of multi-agent systems in value-laden tasks to overlook normative uncertainty embedded in disagreement by overemphasizing consensus. To bridge the gap between subsymbolic reasoning in large language models and symbolic knowledge, the authors propose a symbolic knowledge representation layer that abstracts agents’ reasoning trajectories and decisions into four distinct disagreement states grounded in consistency and conclusiveness. Building upon this framework, they introduce a defeasible policy routing mechanism that enables disagreement-aware agent scheduling in content moderation tasks. This approach significantly enhances the system’s strategic reasoning capabilities in value-sensitive scenarios by explicitly modeling and leveraging normative divergence among agents.
This study evaluates the reasoning capabilities of large language models (LLMs) in democratic deliberation scenarios that lack objective answers and require integration of diverse value systems, thereby exposing the limitations of current evaluation frameworks grounded in verifiable tasks and procedural metrics. For the first time, the empirically validated Deliberative Reasoning Index (DRI) from political science is adapted to assess multi-agent LLM deliberations. The authors conduct a systematic analysis of 1,980 dialogues across 12 civic issues involving 11 state-of-the-art models. Results reveal that while LLM groups achieve procedural discourse quality comparable to humans, they exhibit less than one-third the level of perspective diversity. Moreover, on ethically contentious topics, LLMs struggle to reach consensus, with deliberation often exacerbating disagreement. Role-based prompting fails to replicate human-like deliberative dynamics, challenging the assumption that LLMs can function as autonomous deliberative agents.
This work addresses a critical yet overlooked issue in multi-agent large language model (LLM) negotiations: the conflation of superficial consensus with genuine agreement, which often masks substantial loss of key facts and reduced stance diversity, thereby compromising reliability. To tackle this, the authors propose DelibTrace, a novel framework that decomposes negotiation topics into atomic issues, annotates essential facts, and tracks contextual evolution across dialogue turns. Leveraging this approach, they identify and formally characterize “deliberation hallucination”—a phenomenon wherein agents converge on misleading or incomplete agreements. The study introduces a new evaluation paradigm centered on factual retention rate and stance heterogeneity. Experiments reveal that up to 72% of critical facts can be lost during multi-round negotiations, consensus is heavily biased by base model priors, and even a single malicious agent can corrupt the shared context, leading to collectively degraded knowledge—“the more they negotiate, the less they know.”
This work investigates the design of efficient decision protocols in multi-agent large language model systems to enhance task performance while balancing training and inference costs. The authors propose MALLM, a unified framework that systematically compares three structured decision protocols—voting, consensus, and adjudication—across knowledge-intensive benchmarks (e.g., MMLU, GPQA) and complex reasoning tasks (e.g., StrategyQA, Math-lvl-5). Empirical results demonstrate that consensus protocols achieve superior performance on knowledge-dense tasks, whereas voting and adjudication mechanisms are better suited for intricate logical reasoning. Furthermore, increasing response diversity among agents significantly improves overall decision quality, highlighting the importance of heterogeneity in multi-agent collaborative reasoning.
This work addresses the limitation of existing multi-agent forecasting approaches, where information homogenization often induces herding behavior, thereby constraining belief updating and performance gains. To overcome this, the authors propose InfoDelphi, a novel framework that systematically introduces designed information asymmetry by partitioning evidence into shared and mutually exclusive private subsets, endowing each agent with unique knowledge. Effective calibration is achieved through correlation-aware evidence routing, iterative reasoning-based negotiation, and confidence-weighted aggregation. Theoretically, this mechanism reduces inter-agent error correlation, highlighting input diversity as essential for negotiation gains. On the PolyGym benchmark, InfoDelphi improves Brier scores by 12–18% over both the strongest single-agent and multi-agent baselines and increases accuracy by 4–8 percentage points. Ablation studies confirm that information asymmetry is critical to the negotiation efficacy.