When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the persistent bottleneck wherein multi-agent debate frequently fails to outperform majority voting. We reveal that proposal supply and verification-aware readout constitute the core mechanisms determining debate superiority. To this end, this work formalizes the notion of recoverable margin and introduces the Latent Verification Debate (LVD) model, leveraging neuralδΈ› agents to construct high-coverage societies for optimized reasoning. Experimental results demonstrate that societies selected based on coverage objectives significantly enhance complementary proposal supply and improve aggregation accuracy. The source code has been made publicly available.
πŸ“ Abstract
Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the first requirement through recoverable headroom, which measures cases where a correct proposal is available but the majority answer is wrong. For the second, we develop Latent Verification Debate (LVD), an accounting model in which candidate proposals receive answer-specific verification evidence before final generation. Controlled fixed-proposal interventions estimate this latent effect in equivalent peer-support units and show that correct evidence changes answer probabilities and generated decisions while proposal supply remains fixed. To improve proposal supply, we construct societies from neural-thicket agents using labeled and label-free coverage objectives. Across two backbones and matched-budget reasoning benchmarks, coverage-selected societies increase complementary proposal supply and improve aggregate accuracy in repeated stochastic evaluations. Round-level controls further show that interaction provides gains beyond applying the same finalizer directly to the initial proposals. These results identify proposal coverage and truth-sensitive evidence use as complementary conditions for debate to outperform voting. Code is available at https://github.com/Wang-ML-Lab/when-debate-helps.
Problem

Research questions and friction points this paper is trying to address.

multi-agent debate
proposal supply
majority voting
reasoning
readout mechanism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Debate
Latent Verification Debate
Recoverable Headroom
Proposal Coverage
Neural-Thicket Agents