Score
Designs, implements, and evaluates multi-agent systems that coordinate specialized perception and reasoning agents (including LLM-based agents) into collaborative frameworks and hierarchical workflows, specifying agent interfaces, communication protocols, and orchestration policies. Builds fusion and aggregation methods to align heterogeneous agent outputs into unified, interpretable representations (e.g., evidence chains, spatially coherent alignments) and analyzes cooperative inference, robustness, and generalization across scenarios.
Existing LLM-driven multi-agent collaboration research lacks a systematic, unified framework for modeling dynamic inter-agent relationships and structural principles underlying collective intelligence. Method: We propose the first comprehensive five-dimensional collaboration taxonomy—covering agents, interaction types, organizational structures, coordination strategies, and protocols—and explicitly model cooperative, competitive, and coetitive (cooperative-competitive) dynamics. Leveraging protocol-based collaboration modeling, structured framework design, and cross-domain empirical validation (5G/6G networks, Industry 5.0, question answering, and socio-cultural scenarios), we construct a knowledge graph spanning theory, architecture, and applications. Contribution/Results: Our work identifies decentralization and role-driven strategies as critical drivers of collective intelligence emergence; uncovers core challenges and frontiers in artificial collective intelligence; and establishes a reusable, principled methodology for collaborative AI systems—bridging theoretical rigor with practical deployability.
Existing LLM-based agent reasoning frameworks lack systematic categorization and comparative analysis. Method: This paper presents the first comprehensive survey, proposing a unified taxonomy covering single-agent, tool-augmented, and multi-agent paradigms; formalizing reasoning structures, control flows, and interaction mechanisms via a rigorous descriptive language; and conducting a systematic literature review complemented by cross-domain comparative analysis (research, healthcare, software engineering) and empirical evaluation. Contribution/Results: The study identifies applicability boundaries, performance bottlenecks, and validation methodologies for each paradigm, establishing the first end-to-end mapping from methodological design to real-world application scenarios. It delivers a structured knowledge graph and authoritative reference for theoretical modeling, benchmark development, and engineering deployment of agent reasoning systems.
This work investigates how collaborative architecture design affects collective reasoning in multi-agent large language model (LLM) systems, with a focus on expertise allocation as a critical bottleneck. Methodologically, we conduct systematic ablation studies examining the interplay among three dimensions: domain-aligned expert specialization, collaboration paradigms (structured workflows vs. diversity-driven knowledge fusion), and system scale. Our results show that domain-dependent expert alignment substantially improves reasoning accuracy; diversity-aware knowledge integration outperforms rigid task decomposition; and communication overhead constitutes the primary scalability bottleneck. Based on these findings, we propose a configurable multi-agent design framework that quantifies the compute–performance trade-off under scale expansion and empirically validates the significant gains from expert alignment on context-intensive reasoning tasks.
Existing LLM-based multi-agent systems perform well on isolated tasks but lack key cognitive capabilities inherent in human teams—namely, knowledge sharing, recursive reasoning, structured critical evaluation, and theory of mind (ToM)-driven mental state inference—hindering high-order cognitive collaboration. This paper proposes a novel framework integrating adaptive ToM with systematic critical assessment: it dynamically models agents’ beliefs and intentions, supports multi-level recursive perspective-taking, and embeds formal logical flaw detection and bias mitigation mechanisms to enable deep inter-agent collaboration. Experiments demonstrate significant improvements in reasoning coherence, knowledge integration efficiency, and collective rigor across complex decision-making tasks, outperforming baseline models on multiple collaborative reasoning benchmarks. The core contribution is the first unified modeling of evolvable ToM and formal critical evaluation, establishing a new paradigm for building multi-agent systems capable of genuine cognitive synergy.
Current evaluations of multi-agent systems powered by large language models predominantly emphasize task outcomes or individual agent capabilities, often overlooking core collaborative competencies such as establishing consensus under constrained communication, maintaining mutual understanding, aligning individual and collective goals, and repairing misalignments. To address this gap, this work introduces CollabSim, a novel framework that integrates Computer-Supported Cooperative Work (CSCW) theory into the assessment of multi-agent collaboration for the first time. CollabSim provides a configurable simulation environment, defines collaboration dimensions grounded in CSCW principles, and employs action-level probes to analyze agents’ internal states. Experiments across four large language models demonstrate that CollabSim enables fine-grained, condition-controlled evaluation of collaborative abilities, effectively uncovering the impact of interaction conditions, inter-model differences, and the task-dependence of agent design on collaborative performance.
This work addresses the challenge in multi-agent systems where coarse-grained, opaque collaboration strategies hinder simultaneous optimization of performance and efficiency. We first systematically decompose collaboration into four fine-grained dimensions: agent governance, participation control, interaction dynamics, and dialogue history management. To jointly quantify accuracy and computational cost, we propose the Token-Accuracy Ratio (TAR) metric. Methodologically, we design a context-aware policy controller and a dynamic history compression–abstraction generation technique. Experiments on DEI and SES benchmarks show that the optimal policy combination improves accuracy by 19.3% while reducing token consumption by 32.7%. Our analysis reveals synergistic optimization principles: centralized governance, mentor-led participation, ordered interaction, and teacher-refined summarization. This work shifts the multi-agent design paradigm from architectural innovation toward principled, mechanism-level policy innovation.
Existing LLM-based agent tool integration suffers from architectural fragmentation, insufficient modularity, and terminological inconsistency. Method: This paper proposes the first unified modeling framework for LLM-based agents. Its core innovations include: (1) formally defining the “core agent” as a structured composition of five cohesive modules—planning, memory, persona, execution, and safety—and introducing an active/passive typology to enable multi-core coordination; (2) proposing a novel layered architecture with explicit safety modeling; and (3) establishing a composable, extensible hybrid agent design paradigm. Contribution/Results: Leveraging this framework, we systematically map and clarify the architectural essence of 13 state-of-the-art agents; validate five novel active/passive hybrid architectures; and identify critical cross-agent integration challenges alongside concrete optimization pathways. The framework advances principled, scalable, and secure agent design.
This work addresses the lack of a systematic understanding of the mechanisms that enhance collaborative reasoning performance in multi-agent systems and the difficulty in identifying the key factors that make them superior to single-agent approaches. The authors propose a unified theoretical framework that, for the first time, decomposes multi-agent reasoning gains into three orthogonal dimensions: exploration, information, and aggregation. Building upon this decomposition, they introduce the PRISM framework, which jointly optimizes these dimensions through role-based diversity generation, execution-feedback-driven cross-evaluation of evidence, and an iterative synthesis mechanism with closed-loop verification. Empirical evaluations demonstrate state-of-the-art performance across tasks including mathematical reasoning, code generation, and function calling, while significantly improving computational efficiency. This study thus provides both actionable theoretical insights and a practical paradigm for designing effective multi-agent systems.
Existing research on large language model–based multi-agent systems treats collaboration, fault attribution, and self-evolution in isolation, neglecting their intrinsic causal interdependencies and thereby hindering the realization of sustainable collective intelligence. This work proposes LIFE, a unified framework that structures system development into four phases: capability grounding, collaborative integration, fault attribution, and autonomous evolution. For the first time, it models the dependencies and constraints among these phases within a coherent causal architecture. Through a systematic literature review, formal modeling, and taxonomy construction, the study delineates key technical trajectories, establishes a conceptual roadmap and classification scheme spanning all four phases, and identifies critical challenges at phase boundaries. This provides a theoretical foundation for developing autonomous multi-agent systems capable of continuous diagnosis, reconfiguration, and optimization.
Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology, we synthesize a task-conditioned temporal workflow graph that jointly specifies agent connectivity and edge-level communication semantics. We introduce ReActNet, a training-free framework that compiles a query and a set of role-specialized agents into a sequence of directed communication graphs. Each graph snapshot corresponds to one reasoning stage, and each edge carries a natural-language instruction specifying the message that a source agent should provide to a target agent. The compiled temporal graph is then executed through structured message passing: agents update their reasoning states by integrating their previous states with messages from controller-assigned neighbors, and a final aggregator synthesizes the resulting states into the answer. This design separates graph compilation from graph execution, making multi-agent coordination explicit, inspectable, and task-conditioned without requiring reinforcement learning or gradient-based topology optimization. Across knowledge reasoning, mathematical problem solving, code generation, and GAIA-style assistant tasks, ReActNet consistently improves over fixed-topology and learned-topology baselines while maintaining competitive inference cost. These results suggest that effective multi-agent orchestration depends not only on which agents communicate, but also on engineering executable workflow graphs that encode when, why, and how information should flow during reasoning.
Existing large language model agent systems struggle to meet the demands of production environments—such as simplicity, controllability, and predictable inference costs—due to their high complexity, unbounded reasoning expenses, and unpredictable behavior. To address these limitations, this work proposes a practical, utility-driven agent design framework that employs “pseudo-tools” to enforce modularity, replaces dynamic planning with fixed workflows, and integrates a dedicated learning algorithm to jointly optimize component performance. The approach innovatively applies multi-objective optimization to balance inference cost and response quality, while supporting result fusion across multiple systems. Experimental results demonstrate that the proposed method significantly reduces inference costs and improves accuracy across diverse tasks, outperforming handcrafted dynamic planning baselines.
This work addresses the tight coupling among organizational structure, coordination mechanisms, and collaboration algorithms in multi-agent large language model systems, which hinders independent configuration and evaluation. To resolve this, the paper introduces IMACS, a novel framework that formalizes classical organizational theories—specifically Belbin team roles, Mintzberg’s coordination mechanisms, and the RACI responsibility model—into configurable components, thereby achieving orthogonal decoupling across three layers: team composition, agent alignment, and collaboration algorithms. The framework unifies six collaboration protocols under a common interface and incorporates a context-aware multi-armed bandit-based adaptive routing mechanism for dynamic, task-driven protocol selection. Experimental results demonstrate that adaptive routing significantly outperforms fixed protocols, with optimal configurations varying across model families, thus validating both the necessity of decoupled design and the efficacy of online learning.