Score
Design and implement multi-agent debate frameworks and protocols that generate, surface, and reconcile articulated assumptions and reasoning—this includes building agentic two-way or multi-round debate processes, bi-directional questioning schemes, and detectors for global vs. local disagreement signals to produce agreed articulation parameters. Develop paired affirmative/negated prompting methodologies and bias-aware evaluation procedures, including metrics and analysis pipelines that measure and reduce affirmation and other biases in binary and scalar response tests.
This study addresses the challenges in multi-agent debate (MAD) research stemming from inconsistent terminology and the absence of a systematic design framework, which hinder cross-study comparisons and obscure implicit design choices. Through a systematic literature review of 141 MAD works, this paper proposes the first three-dimensional taxonomy encompassing participants, interaction mechanisms, and consensus protocols, and introduces formal notation to structurally describe MAD configurations. The resulting classification reveals prevalent yet empirically unvalidated design paradigms in the field, identifies critical limitations, and lays the groundwork for establishing controllable benchmarks and enabling automated tuning. Furthermore, it advocates for evolving the taxonomy into an executable, machine-readable specification framework to standardize and advance MAD research.
This paper investigates how model capability heterogeneity in multi-agent debate induces reasoning failures and accuracy degradation. We construct a controlled experimental framework to systematically vary the scale and capability of LLMs participating in debates, enabling fine-grained tracking of reasoning propagation and answer evolution. Our study reveals, for the first time, that even when strong models are present, heterogeneous capability configurations frequently cause consensus preference to override error-correction motivation—leading correct answers to drift toward incorrect ones. Moreover, persuasive yet fallacious reasoning is widely adopted in the absence of alignment incentives and anti-misinformation mechanisms. The core contribution is the identification of “consensus-driven degradation”—a novel failure mode—demonstrating that capability diversity alone does not ensure robustness. We further establish that improving multi-agent reasoning reliability requires co-designing incentive structures and anti-misinformation capabilities. (149 words)
研究通过四种测量方法分析多代理LLM辩论中不同语气对答案质量和意见一致性的影响,发现虽然辩论改变了代理的表达,但对最终答案质量提升的证据较弱。
This study addresses the challenge of disentangling the effects of protocol design from intrinsic model capabilities in multi-agent debate settings. Through controlled experiments in a macroeconomic scenario, the authors compare four interaction protocols: Within-Round (WR), Cross-Round (CR), a novel Rank-Adaptive Cross-Round (RA-CR), and a non-interactive baseline. The RA-CR protocol innovatively incorporates an external evaluator to dynamically adjust speaking order and progressively mute the weakest-performing agent each round, thereby significantly accelerating consensus convergence. Results demonstrate that RA-CR achieves the highest consensus efficiency, WR exhibits the greatest peer citation rate, and the non-interactive baseline preserves the highest argument diversity, collectively revealing an inherent trade-off between interactivity and convergence in multi-agent deliberation.
本文研究辩论判断理论,通过分析和实验两种方法(LLM作为法官与计算论证的形式语义)解决辩论结果评判的可重复性、稳健性、基础性和可解释性问题。
To address the scalability bottleneck in multi-agent debate—specifically, the exponential growth in token consumption with increasing agent count and debate rounds—this paper proposes a *grouped multi-agent debate* architecture. Agents are partitioned into disjoint subgroups that conduct parallel internal debates; inter-group information exchange and a dynamic consensus mechanism then aggregate intermediate results efficiently. This approach breaks the traditional linear scaling constraint and represents the first systematic integration of grouping principles into multi-agent debate frameworks. Extensive experiments across multiple logical reasoning benchmarks demonstrate that our method reduces token consumption by up to 51.7% relative to baseline methods, while simultaneously improving accuracy by up to 25%. The architecture thus achieves a significant trade-off improvement between computational efficiency and reasoning performance.
论文提出Meta-Moderator框架,通过元认知过程动态调控多智能体辩论,解决现有方法中冗余讨论和证据聚合不可靠的问题。
本文提出MABPD,通过多智能体结构化辩论检测新闻文章中的媒体偏见,无需监督训练即可达到接近监督模型的性能。
This study addresses the significant first-speaker bias in sequential multi-agent debate, which causes stronger models to lose their reasoning advantage when speaking later. To tackle this issue, we first quantify the impact of such bias across varying strong-weak model configurations and construct a prompt intervention framework grounded in Big Five personality theory. Specifically, we propose low agreeableness as a behavioral modulation strategy to reshape agent influence dynamics. Experimental results demonstrate that low-agreeableness prompting effectively restores the discursive power of stronger models and improves final decision accuracy, whereas extraversion merely increases response redundancy without yielding systematic effects. This work offers a novel paradigm for optimizing the fairness and effectiveness of large language model debate mechanisms.
This work challenges the conventional view in multi-agent systems that treats disagreement among AI agents as mere noise to be eliminated, arguing instead that such divergence may reflect genuine value pluralism—particularly in culturally and subjectively charged tasks like hate speech moderation. The study proposes a novel, reasoning-structure-based taxonomy classifying AI disagreements into four types and implements it using five large language model (LLM) agents with diverse perspectives, generating reasoning traces on the Measuring Hate Speech dataset. By combining embedding and classification models, the framework identifies patterns such as “convergent disagreement.” Experimental results show that when agent conclusions align, human annotator disagreement drops significantly (d > 0.8), and the structure of AI disagreements strongly correlates with human conflicts. These findings demonstrate the approach’s efficacy in guiding human–AI collaborative decision-making and advocate for shifting multi-agent systems from consensus-seeking toward explicit uncertainty representation.
Existing multi-agent debate frameworks rely on static architectures, incurring high computational overhead and lacking flexibility in dynamically adjusting agent roles and coordination mechanisms. This work proposes MoD, a unified self-debate framework based on a Mixture-of-Experts (MoE) architecture that simulates diverse debating behaviors within a single model. MoD decouples role assignment from process control via a dual-routing mechanism and employs momentum-based switching for smooth expert selection. Debate personas are encapsulated as lightweight expert modules, eliminating inter-agent communication costs. Experimental results demonstrate that MoD significantly outperforms both single-model baselines and conventional multi-agent systems across multimodal benchmarks, achieving higher accuracy while reducing inference latency by 3.7× and token consumption by 87%.