Score
Designs and implements systems in which large language models are used to generate, arbitrate, and adapt coordination plans among multiple agents or system components; this includes interpreting semantic intents and constraints, negotiating resource allocations, producing multi‑objective control plans, and updating decisions from telemetry feedback.
This work addresses the fragmentation and lack of systematic frameworks in large language model (LLM)-based agent research. We propose the first methodology-centered, three-dimensional taxonomy—spanning *architecture design*, *collaboration mechanisms*, and *evolutionary pathways*—to unify key technical areas including multi-agent systems, prompt engineering, tool utilization, reflection mechanisms, reinforcement learning, and evaluation frameworks. By elucidating the intrinsic links between agent design principles and emergent behaviors in complex environments, we construct a structured knowledge graph and an open-source literature repository (Awesome-Agent-Papers). Our contributions are: (1) the first methodology-centered taxonomy for LLM agents; (2) a comprehensive, stack-wide survey—from monolithic agents to collective coordination; and (3) a reproducible evaluation benchmark alongside a forward-looking research roadmap. This work establishes both theoretical foundations and practical guidance for the LLM agent community.
This paper investigates the potential and challenges of integrating large language models (LLMs) across the full lifecycle of agent-based modeling (ABM)—from problem formulation and model design to implementation, analysis, interpretation, and dissemination. Adopting a process-oriented approach, it systematically maps LLM capabilities—including text generation, information extraction, logical reasoning, and natural language interaction—to each ABM stage, offering the first comprehensive mapping of such synergies. The study proposes a critical integration framework that delineates current practical boundaries, technical limitations (e.g., interpretability, reliability, domain adaptability), and key risks (e.g., hallucination, bias, lack of formal validation). By synthesizing empirical insights and conceptual analysis, the work establishes foundational principles and methodological guidelines for the safe, transparent, and trustworthy incorporation of LLMs in computational social science and complex systems modeling. (149 words)
Solving control engineering problems traditionally requires deep domain expertise, posing a significant barrier to accessibility and automation. Method: This paper proposes LLM-Agent-Controller—the first general-purpose, multi-agent large language model system tailored for control theory. It employs a domain-specific multi-agent architecture integrating retrieval-augmented generation (RAG), chain-of-thought reasoning, self-critique refinement, efficient memory management, and role-based collaboration. Users submit queries in natural language; the system autonomously executes end-to-end tasks including dynamic modeling, controller synthesis, stability analysis, time-domain response evaluation, and simulation. Contribution/Results: Innovations include supervised workflow orchestration and zero-barrier interaction, enabling real-time problem solving without prior control theory knowledge. Evaluated on five canonical control problem classes, the system achieves an overall success rate of 83% and an average per-agent success rate of 87%. Performance scales significantly with base model capability, demonstrating robust generalization and practical applicability.
This study investigates whether large language models (LLMs) can replace hand-coded logic to model adaptive behavior and self-organized emergence in multi-agent systems. We propose LLM-NetLogo, a synergistic framework integrating structured prompting with knowledge-driven dual-mode prompting to enable real-time perception–decision–response capabilities for agents in dynamic environments. Implemented via GPT-4o and the NetLogo Python extension, the framework achieves high-fidelity replication and behavioral extension of two canonical collective phenomena: ant foraging and bird flocking. Experiments demonstrate that LLMs effectively generate emergent collective behaviors beyond predefined rule constraints. To foster reproducibility and extensibility, we open-source all code, prompt templates, and simulation datasets. This work establishes the first empirically grounded, reproducible, and scalable paradigm for leveraging LLMs in complex systems modeling.
Large language models (LLMs) exhibit limited autonomous capability in solving open-ended problems, primarily due to overreliance on explicit algorithms and static knowledge. Method: We propose a novel end-to-end paradigm—spanning problem framing, solution exploration, implementation generation, and strategy assessment—that integrates prompt engineering, retrieval-augmented generation (RAG), and reinforcement learning from human feedback (RLHF). This synergy enhances LLMs’ proficiency in feature composition, dynamic anomaly response, and high-level strategy evaluation. Contribution/Results: We present the first systematic taxonomy of paradigm evolution for LLM-based implementation generation, identify critical technical bottlenecks, and establish a theoretical framework and technology roadmap for autonomous problem solving. Our work advances the development of LLM-driven general-purpose agents by enabling more robust, adaptive, and self-assessing reasoning capabilities.
Large language models (LLMs) exhibit limited planning capabilities in multi-step reasoning and goal-directed tasks. To address this, we propose the Modular Agent Planning (MAP) architecture—a cognitively inspired, reinforcement learning–informed framework that decomposes planning into specialized LLM modules: conflict monitoring, state prediction, and task decomposition. These modules operate in a recurrent, collaborative loop to enable dynamic, adaptive planning. MAP supports lightweight deployment and cross-task generalization without fine-tuning, seamlessly adapting to diverse LLM scales (e.g., Llama3-70B). Empirical evaluation on graph traversal, Tower of Hanoi, PlanBench, and StrategyQA demonstrates substantial improvements over zero-shot prompting, chain-of-thought, and tree-of-thought baselines—yielding higher planning accuracy and robustness. Our core contribution is the first systematic integration of modular, division-of-labor mechanisms into LLM-based planning, establishing a novel paradigm for structured, interpretable, and scalable agent reasoning.
This work addresses the challenge of coordination in large language model (LLM)-based multi-agent systems, where nondeterministic LLM behavior can lead to subtle, hard-to-detect errors such as deadlocks or message mismatches. The paper introduces ZipperGen, a novel framework that formally incorporates Message Sequence Charts (MSCs) into LLM-driven multi-agent coordination for the first time. It employs a domain-specific language to decouple communication structure from LLM behavior and uses syntax-guided projection to derive local agent programs from a global specification, guaranteeing deadlock freedom by construction. This approach enables a verifiable coordination mechanism that is disentangled from LLM nondeterminism and supports runtime generation of structurally sound workflows. The framework’s ability to independently verify coordination properties is demonstrated through its application to consensus protocol diagnostics.
This work addresses the absence of a systematic framework for evaluating the effectiveness, structural design, and team size selection of large language model (LLM) ensembles. It pioneers the integration of distributed systems theory into LLM team research by conceptualizing such teams as distributed intelligent systems. Combining multi-agent coordination mechanisms with system architecture analysis, the study establishes a cross-disciplinary analytical framework. This framework uncovers fundamental parallels between LLM teams and traditional distributed systems in terms of performance, scalability, and fault tolerance, thereby offering both theoretical foundations and practical guidance for the design, evaluation, and optimization of LLM teams.
This work addresses the imbalance between structure and flexibility in existing multi-agent collaboration frameworks powered by large language models, which often leads to inefficiency and resource waste. To overcome this limitation, the authors propose LATTE—a novel framework inspired by distributed systems—that introduces a dynamically evolving, shared task graph to explicitly encode subtask dependencies, agent assignments, and progress states. By integrating LLM-based multi-agent systems with distributed coordination protocols, LATTE enables adaptive collaboration structures, dynamic task allocation, and emergent task discovery. Empirical results demonstrate that LATTE significantly reduces token consumption, execution time, and communication overhead across diverse collaborative tasks, while minimizing file conflicts and redundant outputs. Moreover, it achieves accuracy on par with or superior to strong baselines such as MetaGPT.
This study presents the first systematic evaluation of the scientific validity of large language models (LLMs) in automatically generating reproducible and verifiable agent-based model code from ODD protocol specifications. Using the PPHPC predator–prey model as a benchmark, the authors assess Python implementations produced by 17 LLMs across four dimensions: executability, behavioral consistency, computational efficiency, and code maintainability. Results indicate that GPT-4.1 consistently generates statistically valid and efficient code, with Claude 3.7 Sonnet performing comparably but exhibiting lower stability. Crucially, the study demonstrates that executability does not guarantee behavioral fidelity, underscoring the essential role of formal verification in scientific modeling. While these findings reveal the potential of LLMs in specification-driven modeling, they also affirm that current models cannot yet replace human expertise in rigorous scientific model development.
This work addresses the lack of systematic architectural approaches for enterprise-scale multi-agent collaborative systems, particularly in complex scenarios integrating human and AI agents. The authors propose a three-layer design pattern—comprising LLM agents, autonomous agents, and agent communities—that integrates principles from distributed coordination, formal modeling, and a role-protocol-governance structure. For the first time, this framework introduces formal collaboration protocols and role-based governance mechanisms into agent communities, enabling executable specification and verification of organizational, legal, and ethical rules. Validated through a clinical trial matching case study, the architecture demonstrates governable and verifiable human-AI collaboration, offering both formal verification capabilities and actionable design guidance for enterprise deployment.