Score
Designs and implements simulation platforms that embed or couple large language models as agents, controllers, or oracle components, including hybrid LLM–simulator architectures and coupling layers that mediate inputs, observations, and outputs. Builds agent-facing RESTful APIs, observation pipelines, role-based information access and hidden validation logic, optional 2D/3D visualizers, and support for multi-agent scenarios and real-time simulator–LLM interaction.
Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.
本文提出一个多代理框架,使大型语言模型能够通过科学模拟模型进行受控实验,以优化制药过程设计,提高输出的具体性和实用性。
This survey systematically examines the current state and potential of large language model (LLM)-driven agents in software engineering (SE). Addressing the lack of a unified analytical framework and unclear human-agent collaboration mechanisms in prior work, we analyze 124 studies to propose the first two-dimensional taxonomy—“SE Tasks × Agent Capabilities”—spanning the full SE lifecycle: requirements analysis, coding, testing, and maintenance. Methodologically, we synthesize key techniques including tool use, multi-agent systems, reflection mechanisms, and external knowledge retrieval, and release an open-source literature repository, Agent4SE-Paper-List. Our contributions include: (1) clarifying evolutionary trajectories and bottlenecks in multi-agent coordination and human-agent interaction; (2) identifying six open challenges; and (3) proposing three future research directions—scalable evaluation, domain alignment, and trustworthy collaboration.
The integration of large language model (LLM)-driven multi-agent systems (LMAS) into the full software engineering (SE) lifecycle remains underexplored, with no comprehensive mapping of LMAS applications across SE phases or consensus on research priorities. Method: We conduct a systematic literature review and empirical case studies using mainstream LMAS frameworks (e.g., AutoGen, CrewAI) to analyze LMAS deployment across SE stages—requirements, design, development, testing, and operations. Contribution/Results: We present the first phase-wise LMAS application taxonomy for SE, identifying critical research gaps. We propose the “SE 2.0” vision and a dual-track research agenda emphasizing *individual agent capability enhancement* and *cross-agent collaboration optimization*. Experimental evaluation demonstrates that state-of-the-art LMAS frameworks significantly improve autonomy, robustness, and scalability in real-world SE tasks—yet expose limitations in consistency, traceability, and domain-specific reasoning. This work establishes a foundational theoretical framework and actionable implementation guidelines for LLM-augmented SE.
Traditional industrial automation systems suffer from high operational complexity and require labor-intensive reprogramming for process changes. Method: This paper proposes the first end-to-end large language model (LLM)-driven industrial control system. We introduce an industrial task agent framework integrating structured prompt engineering, multi-semantical-level event-driven modeling, and real-time data integration. Additionally, we propose a reusable, task-specific dataset generation methodology to support LLM domain adaptation and evaluation. Contribution/Results: We establish the first structured prompt paradigm for industrial automation; enable natural-language instruction parsing, dynamic production planning, and closed-loop device control; and support responsive adaptation to unforeseen operational conditions—substantially lowering the human operator barrier. A formal system design and proof-of-concept implementation have been completed, with source code and demonstration videos publicly released.
Evaluating LLM-based agents is challenging due to their dynamic, probabilistic, and continuously evolving nature—traditional predefined benchmarks fail to capture open-ended behaviors, emergent outcomes, and lifecycle adaptation. Method: We propose an evaluation-driven agent development paradigm, integrating online runtime evaluation with offline reconstruction evaluation. Our hybrid framework enables real-time feedback injection, human-AI collaborative closed-loop refinement, and iterative optimization across the full stack (pipeline, architecture, and LLM), incorporating both human and AI evaluators. Contribution/Results: We introduce the first evaluation-centric process model and reference architecture for LLM agent development. It uniquely supports open-behavior capture, emergent-result governance, and dynamic alignment. Experiments demonstrate that the framework effectively enables safe, controllable, and continuous agent iteration under objective drift, requirement changes, and regulatory evolution—achieving robust adaptability without compromising reliability or compliance.
This work systematically designs and evaluates large language model (LLM)-based multi-agent systems to enhance development efficiency, reliability, and scalability in real-world applications. By formalizing multi-agent design patterns, the study proposes a modular architecture, standardized communication protocols, and a controllable orchestration mechanism. Empirical evaluations are conducted across three practical domains: telecommunications security, cultural heritage management, and utility customer service automation. The project establishes the first architectural paradigm for LLM-driven multi-agent systems, enabling prototype delivery within two weeks and pilot deployment within one month—significantly reducing development costs and improving user accessibility. However, the study also identifies the inherent behavioral instability of LLMs as a critical barrier to robust production deployment.
This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.
This study addresses the proliferation of labels in current LLM applications, where brand marketing concepts frequently obscure genuine architectural distinctions. Through a comprehensive literature review and multidimensional feature analysis, this work establishes a unified terminology system, systematically delineates seven LLM integration architectures, and constructs an empirical corpus encompassing 22 systems. The findings reveal fundamental differences between Copilots and Agents regarding user intervention and execution planning across dimensions such as control flow. By clarifying the boundaries and limitations of each architecture, this research provides a rigorous, structured analytical framework for the classification of LLM applications.
Large language models (LLMs) suffer from hallucination, unreliability, and uncontrolled behavior, hindering their trustworthy deployment in safety-critical workflows; existing reliability-enhancement tools are fragmented and lack a systematic framework. This paper introduces LSL (LLM Scripting Language), a domain-specific scripting language that embeds formal specifications, verifiable constraints, and explainability mechanisms directly into the LLM interaction process—enabling structured output constraints, programmable behavioral control, and decoupled execution governance. LSL unifies domain-specific language (DSL) design, formal verification, and runtime checking, significantly improving output reliability, consistency, and traceability. Experiments demonstrate that LSL effectively mitigates hallucination across diverse tasks, supports safe and controllable LLM integration, and establishes a novel interaction paradigm for trustworthy AI systems.
This work proposes AgentForge, a lightweight and open-source Python framework designed to overcome the limitations of existing large language model (LLM) agent frameworks—namely architectural rigidity, vendor lock-in, and high complexity—that hinder rapid development. AgentForge employs a modular design to enable flexible construction of LLM-driven autonomous agents, featuring composable skill abstractions, a unified LLM backend interface, and declarative YAML-based configuration that expresses arbitrary sequential and parallel task flows as directed acyclic graphs (DAGs). The framework supports both cloud APIs and local inference engines, offering six built-in skills alongside an extensible mechanism for custom implementations. Experimental results demonstrate competitive task completion rates across four benchmark scenarios, with development time reduced by 62% compared to LangChain and by 78% versus direct API integration, while maintaining orchestration overhead below 100 ms—making it suitable for real-time applications.