🤖 AI Summary
Large language models (LLMs) suffer from limited single-agent reasoning capacity, while multi-agent debate (MAD) approaches incur excessive communication overhead. Method: This paper proposes MARS, a role-based multi-agent reasoning framework that emulates the academic paper review process—partitioning agents into Authors, Reviewers, and a Meta-Reviewer. MARS employs a role-separation architecture and asynchronous feedback integration to enable non-interactive collaboration, eliminating redundant inter-Reviewer communication. Contribution/Results: Its core innovation lies in adopting an academic-review-inspired paradigm that decouples generation from evaluation, thereby balancing reasoning quality and efficiency. Experiments across multiple benchmarks show that MARS achieves accuracy comparable to MAD while reducing token consumption and inference latency by approximately 50%.
📝 Abstract
Large language models (LLMs) have achieved impressive results in natural language understanding, yet their reasoning capabilities remain limited when operating as single agents. Multi-Agent Debate (MAD) has been proposed to address this limitation by enabling collaborative reasoning among multiple models in a round-table debate manner. While effective, MAD introduces substantial computational overhead due to the number of agents involved and the frequent communication required. In this paper, we propose MARS (Multi-Agent Review System), a role-based collaboration framework inspired by the review process. In MARS, an author agent generates an initial solution, reviewer agents provide decisions and comments independently, and a meta-reviewer integrates the feedback to make the final decision and guide further revision. This design enhances reasoning quality while avoiding costly reviewer-to-reviewer interactions, thereby controlling token consumption and inference time. We compared MARS with both MAD and other state-of-the-art reasoning strategies across multiple benchmarks. Extensive experiments with different LLMs show that MARS matches the accuracy of MAD while reducing both token usage and inference time by approximately 50%. Code is available at https://github.com/xwang97/MARS.