🤖 AI Summary
This study addresses the challenge that multi-agent debate frameworks struggle to outperform strong single-agent and consistency baselines under equivalent cost constraints. To overcome this limitation, we propose a conditional progressive pruning framework that integrates a lightweight pruning strategy with a multi-round communication protocol. This approach optimizes the collaboration efficiency of large language model-driven agents, fully unleashing the potential of multi-turn interactions. Experimental results demonstrate that the proposed framework significantly surpasses existing methods across multiple mainstream benchmarks. Notably, it is the first to consistently outperform the consistency baseline under strict cost constraints, thereby establishing the performance advantages of multi-agent debate and validating its effectiveness and superiority.
📝 Abstract
Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. However, existing MAD frameworks fail to beat strong Single Agent and Consistency-based baselines under the same strict cost limit, which shakes the foundation of the MAD field. We propose Conditional Progressive Pruning (CPP), a lightweight pruning framework that fully exploits multi-round MAD. CPP outperforms all existing MAD frameworks on multiple dominated benchmarks. It is also the first to fully outperform consistency methods. Our code, detailed agent interaction records will be released soon.