Score
Designs and applies quantitative and qualitative analyses to characterize how multiple autonomous agents collaborate, including defining and measuring interaction patterns, communication protocols, coordination strategies, and metrics such as task completion rates and robustness to noisy priors. Builds evaluations that compare collaboration modes and team sizes, analyze collaboration dynamics over time, and assess the performance and failure modes of collaborative multi-agent LLM systems.
Existing LLM-driven multi-agent collaboration research lacks a systematic, unified framework for modeling dynamic inter-agent relationships and structural principles underlying collective intelligence. Method: We propose the first comprehensive five-dimensional collaboration taxonomy—covering agents, interaction types, organizational structures, coordination strategies, and protocols—and explicitly model cooperative, competitive, and coetitive (cooperative-competitive) dynamics. Leveraging protocol-based collaboration modeling, structured framework design, and cross-domain empirical validation (5G/6G networks, Industry 5.0, question answering, and socio-cultural scenarios), we construct a knowledge graph spanning theory, architecture, and applications. Contribution/Results: Our work identifies decentralization and role-driven strategies as critical drivers of collective intelligence emergence; uncovers core challenges and frontiers in artificial collective intelligence; and establishes a reusable, principled methodology for collaborative AI systems—bridging theoretical rigor with practical deployability.
Existing research on LLM-driven multi-agent systems (LLM-MAS) lacks a systematic, communication-centric perspective on how natural language interaction enables collective intelligence and adaptive collaboration. Method: We propose the first “communication-centered” analytical framework that bridges system-level dimensions (architecture, paradigm) and mechanism-level components (goals, strategies, content generation), uncovering how coupling among communication elements shapes collaborative flexibility and emergent intelligence. Based on this, we establish the first taxonomy of natural language communication specifically for LLM-MAS. Contribution/Results: The taxonomy identifies three core challenges—scalability, security, and multimodal integration—and derives corresponding design principles and development roadmaps. Our work provides a theoretical foundation and practical guidance for building robust, interpretable, and cross-domain collaborative multi-agent systems.
In dynamic, partially observable environments, multi-agent systems must leverage effective communication to reduce uncertainty and enable collaboration. This work proposes a “Five Ws” analytical framework—addressing who communicates, when, what content is shared, and why—to systematically unify and analyze the evolution and design logic of communication mechanisms across three major paradigms: multi-agent reinforcement learning (MARL), emergent communication, and large language models (LLMs). By integrating insights across these paradigms, the study reveals fundamental trade-offs and shared challenges concerning interpretability, generalization, and scalability. It further distills practical communication design patterns and outlines a novel direction toward hybrid collaborative systems that synergistically integrate learning, language, and control.
Existing research on large language model–based multi-agent systems treats collaboration, fault attribution, and self-evolution in isolation, neglecting their intrinsic causal interdependencies and thereby hindering the realization of sustainable collective intelligence. This work proposes LIFE, a unified framework that structures system development into four phases: capability grounding, collaborative integration, fault attribution, and autonomous evolution. For the first time, it models the dependencies and constraints among these phases within a coherent causal architecture. Through a systematic literature review, formal modeling, and taxonomy construction, the study delineates key technical trajectories, establishes a conceptual roadmap and classification scheme spanning all four phases, and identifies critical challenges at phase boundaries. This provides a theoretical foundation for developing autonomous multi-agent systems capable of continuous diagnosis, reconfiguration, and optimization.
Current evaluations of multi-agent systems powered by large language models predominantly emphasize task outcomes or individual agent capabilities, often overlooking core collaborative competencies such as establishing consensus under constrained communication, maintaining mutual understanding, aligning individual and collective goals, and repairing misalignments. To address this gap, this work introduces CollabSim, a novel framework that integrates Computer-Supported Cooperative Work (CSCW) theory into the assessment of multi-agent collaboration for the first time. CollabSim provides a configurable simulation environment, defines collaboration dimensions grounded in CSCW principles, and employs action-level probes to analyze agents’ internal states. Experiments across four large language models demonstrate that CollabSim enables fine-grained, condition-controlled evaluation of collaborative abilities, effectively uncovering the impact of interaction conditions, inter-model differences, and the task-dependence of agent design on collaborative performance.
This work addresses the degradation of reasoning quality and unreliable verification in LLM-based multi-agent systems for scientific computing—specifically linear-elastic finite element analysis—caused by collaborative dynamics. We systematically identify three systemic failure modes: confirmation bias, premature consensus, and verification–validation decoupling, leading to undetected physics-inconsistent code. Building upon the AutoGen framework, we design a role-specialized tri-agent system (Coder/Executor/Critic) and evaluate it via controlled dialogue experiments under a dual-criteria assessment paradigm: physical consistency and executable correctness. Results show that functional complementarity outweighs team size; Critic involvement achieves 100% correctness in both physics and visualization; and confirmation bias is detected with 85–92% accuracy. Based on these findings, we propose three actionable design principles—role differentiation, multi-level verification, and anti-premature-convergence interaction—to establish a foundation for engineering-grade trustworthy multi-agent systems.
This study investigates how large language model (LLM)-driven multi-agent systems autonomously reach numerical consensus without predefined negotiation protocols, and applies this capability to zero-shot autonomous aggregation in multi-robot systems. Method: We propose an LLM-based multi-agent negotiation framework integrating numerical consensus modeling, network topology simulation, and real-world validation via ROS-integrated robotic platforms. Contribution/Results: We systematically discover— for the first time—that LLM agents inherently converge toward averaging-based consensus strategies without explicit instruction; that agent personality traits and communication topology critically modulate negotiation dynamics; and that this emergent mechanism generalizes directly to zero-shot collaborative planning. Experiments demonstrate high-robustness autonomous aggregation in both simulation and physical deployments, achieving a 92.7% convergence rate. These results validate the interpretability, generalizability, and practical deployability of LLM-mediated consensus behavior in embodied multi-agent coordination.
This work systematically identifies and addresses four open challenges in large language model (LLM)-driven multi-agent systems: inefficient dynamic task allocation, insufficient robustness in collaborative reasoning, difficulty in hierarchical context modeling, and weak long-range memory coordination. To tackle these, we propose a novel architecture integrating iterative debate mechanisms, hierarchical context encoding, memory-augmented retrieval, and blockchain-based smart contract integration. We establish the first comprehensive challenge taxonomy covering collaborative reasoning, dynamic context modeling, and cross-layer memory coordination—distilling six fundamental unsolved problems. Furthermore, we introduce the first verifiable, scalable, and interpretable LLM multi-agent paradigm tailored to real-world distributed environments (e.g., blockchain systems). Our framework unifies theoretical advancement and practical deployment, providing a principled roadmap for both research and engineering.
This study investigates the scaling behavior of homogeneous multi-agent systems built upon a single large language model as the number of agents increases. To this end, we propose the Sequential Iterative Multi-Agent System (SIMAS) framework, which enables systematic analysis of collaborative dynamics through sequential iterative communication, diverse task benchmarks, large language models of varying scales, and structured debate topologies. Our findings reveal that multi-agent performance exhibits diminishing returns with increasing agent count, governed by a trade-off between collaborative gains and coordination overhead. Collective intelligence is shown to depend critically on interaction design rather than sheer agent quantity. Moreover, effective collaboration requires a sufficiently capable base model, and the optimal number of agents varies significantly with task type. These results remain consistent across multiple interaction architectures.
This study investigates why large language models fail to cooperate even in zero-cost collaborative settings—where assisting others incurs no loss or gain and explicit cooperation is requested. By constructing a simplified multi-agent environment and integrating causal decomposition, communication interventions, and reasoning trace analysis, the work reveals for the first time that model capability exhibits no positive correlation with willingness to cooperate. The authors demonstrate that explicit coordination protocols and minimal shared incentives substantially enhance collaborative performance: under such protocols, a lower-capability model (o3-mini) achieves 50% of optimal collective performance, whereas a stronger model (o3) reaches only 17%. Moreover, even slight incentives effectively mitigate weak cooperative tendencies, underscoring the necessity of purpose-built mechanisms to foster reliable collaboration in artificial agents.
This study addresses the performance limitations of self-organized multi-agent teams powered by large language models (LLMs) when operating without predefined collaboration protocols. Such teams often underperform their best individual member by up to 37.6%, primarily due to ineffective utilization of expert knowledge. Integrating organizational psychology theory with multi-agent dialogue analysis, human-inspired experiments, and machine learning benchmarks, this work reveals that the key bottleneck stems not from failure to identify experts but from integrative compromises during negotiation. Furthermore, team performance exhibits a negative correlation with group size. While consensus-seeking behavior diminishes the utility of expert agents, it concurrently enhances system robustness against adversarial agents, highlighting a trade-off between expertise exploitation and collective resilience.
This study addresses the challenge non-technical researchers face in designing and analyzing complex experiments in multi-agent team dynamics. We propose VirTLab: an interactive 2D simulation platform powered by large language models (LLMs). Integrating team cognition theory with scalable agent modeling, VirTLab enables users—without programming expertise—to define environments, agent roles, tasks, and interaction rules, facilitating flexible simulation of coordination mechanisms, collective behavior, and emergent phenomena in human-AI collaboration. Its key contribution lies in balancing ecological validity and accessibility: spatialized agent behavior modeling, role-driven communication protocols, and real-time visualization empower both technical and non-technical researchers to conduct empirically grounded experiments. Evaluation demonstrates high fidelity between VirTLab’s simulated outputs and observed human team behavior, significantly lowering the barrier to entry for multi-agent experimentation.
This work addresses the challenge of error propagation in multi-agent collaboration caused by the absence of critical reasoning and verification information in early-stage communication, which significantly degrades system performance. To mitigate this issue, the authors propose Category-Aware Recovery Augmentation—a method that enhances communication quality within large language model–based multi-agent frameworks by identifying and explicitly embedding essential reasoning and validation content. Experimental results demonstrate that the approach successfully recovers 86.2% of previously failed cases across diverse tasks, underscoring the pivotal role of high-fidelity communication in collaborative efficacy. The study not only validates the importance of structured information exchange but also establishes a novel paradigm for designing robust communication protocols in multi-agent systems.