🤖 AI Summary
This work addresses the limitation of existing multi-agent collaboration benchmarks, which predominantly assume global observability and lack systematic evaluation under partial observability. We introduce a novel benchmark for multi-agent collaboration under partial observability constraints, encompassing reasoning, scheduling, and gaming tasks. Through a pioneering hierarchical evaluation framework, this benchmark systematically assesses the collaborative efficacy of protocol, memory, and routing mechanisms while providing deterministic quantitative metrics such as communication costs. Leveraging large language model (LLM)-driven multi-agent systems and environment simulations, cross-model experiments reveal critical performance-overhead trade-offs across diverse configurations. These findings offer empirical guidance for the design of efficient multi-agent systems operating in partially observable environments.
📝 Abstract
Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems depends not only on the agents themselves, but also on how collaboration mechanisms are designed and organized. Note that real-world collaboration is typically partially observable, where each agent can only access partial information about the environment due to physical or privacy-related constraints. However, many existing multi-agent benchmarks assume global observability, and leave limited support for systematically evaluating collaboration mechanisms. To bridge this gap, we introduce MASBench, a multi-agent collaboration benchmark designed under partially observable constraints. It is organized into three progressive task categories: Reasoning, Scheduling, and Game. Through this structure, we progressively evaluate three representative collaboration mechanisms: Protocol, Memory, and Routing. MASBench further provides deterministic evaluation metrics, including performance score, communication cost, and cost effectiveness, to characterize both collaboration outcomes and communication overhead. Experiments across diverse LLM backbones and mechanism configurations offer empirical guidance for effective MAS design. Code is available at: https://github.com/BUPT-GAMMA/MASBench