Score
Designs, builds, and evaluates LLM-driven agent systems that augment language models with external tools and predictors by implementing controllers (MCP/agent) that invoke generation, prediction, and mutation tools, consult live predictor outputs between calls, record per-step decision traces, and orchestrate tool sequencing and error handling. This work includes engineering low-latency, multi-turn interaction flows, integrating tool APIs, and analyzing agent behavior for correctness, latency, and traceability.
Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.
This study addresses the question of whether performance bottlenecks of large language model (LLM) agents in long-horizon tasks stem from the base model, the execution framework, or their coupling. To this end, the authors propose a “model–framework” co-analysis perspective, decomposing the execution framework into six runtime responsibilities: observation, context, control, action, state, and verification. They further integrate this decomposition with the four paradigms of agent engineering—prompt engineering, workflow design, context engineering, framework engineering, and co-training—to systematically categorize existing approaches. The resulting analytical framework elucidates how runtime design choices influence task success rate, efficiency, and reliability, while highlighting critical challenges including value-aware evaluation, safety, framework generalization, and co-evolution of models and frameworks.
This work addresses the challenge of coordinating multiple tools with large language models to accomplish long-horizon tasks in complex, dynamic environments—a setting where prior research has largely been confined to single-tool or one-step interactions. We formulate multi-tool coordination as dynamic orchestration over extended trajectories and introduce a unified task definition alongside a structured analytical framework. Through six core dimensions—planning and execution, training methodologies, safety, efficiency, capability completeness, and evaluation benchmarks—we systematically review key techniques, including reasoning-time planning, trajectory-based training, safety-aware control, resource-constrained optimization, and modeling in open-ended environments. By synthesizing insights from software engineering, enterprise workflows, GUI automation, and mobile systems, we delineate the fundamental challenges and future directions for building reliable, scalable, and verifiable multi-tool agents.
This work exposes a critical vulnerability in large language models (LLMs): their tool invocation behavior is highly sensitive to textual tool descriptions—minor, semantically preserving edits can induce drastic shifts in model preference, with invocation rates increasing over 10× in some cases. We systematically identify and name this phenomenon “tool preference manipulation.” To rigorously characterize it, we design a controllable evaluation framework grounded in the Model Context Protocol (MCP) and conduct cross-model experiments across ten state-of-the-art models—including GPT-4.1 and Qwen2.5-7B—quantifying the impact of diverse description-editing strategies. Our results reveal both the prevalence and model-specific generalization patterns of this phenomenon, demonstrating that susceptibility varies significantly across architectures and scales. This provides the first empirical foundation for diagnosing and improving tool selection mechanisms, advancing the development of more robust, reliable, and trustworthy agent-based tool invocation systems.
To address the limited generalization capability of LLM-based agents in specialized domains such as life sciences and medicine—stemming from their reliance on manually pre-written tools—this paper introduces ToolMaker: the first end-to-end framework for automatically constructing LLM-callable tools from scientific code repositories. Its core innovation is a closed-loop, self-correcting paradigm for fully automated tool generation, integrating multi-step reasoning–driven agent orchestration, automatic dependency installation, iterative code generation and debugging, and unit-test–driven robustness validation. Evaluated on a benchmark of 15 complex tasks spanning medical and non-medical domains, ToolMaker achieves an 80% accuracy rate, substantially outperforming existing software-engineering–oriented LLM agents. It represents the first system capable of autonomously transforming research code into production-ready, LLM-executable tools.
The integration of large language model (LLM)-driven multi-agent systems (LMAS) into the full software engineering (SE) lifecycle remains underexplored, with no comprehensive mapping of LMAS applications across SE phases or consensus on research priorities. Method: We conduct a systematic literature review and empirical case studies using mainstream LMAS frameworks (e.g., AutoGen, CrewAI) to analyze LMAS deployment across SE stages—requirements, design, development, testing, and operations. Contribution/Results: We present the first phase-wise LMAS application taxonomy for SE, identifying critical research gaps. We propose the “SE 2.0” vision and a dual-track research agenda emphasizing *individual agent capability enhancement* and *cross-agent collaboration optimization*. Experimental evaluation demonstrates that state-of-the-art LMAS frameworks significantly improve autonomy, robustness, and scalability in real-world SE tasks—yet expose limitations in consistency, traceability, and domain-specific reasoning. This work establishes a foundational theoretical framework and actionable implementation guidelines for LLM-augmented SE.
Evaluating LLM-based agents is challenging due to their dynamic, probabilistic, and continuously evolving nature—traditional predefined benchmarks fail to capture open-ended behaviors, emergent outcomes, and lifecycle adaptation. Method: We propose an evaluation-driven agent development paradigm, integrating online runtime evaluation with offline reconstruction evaluation. Our hybrid framework enables real-time feedback injection, human-AI collaborative closed-loop refinement, and iterative optimization across the full stack (pipeline, architecture, and LLM), incorporating both human and AI evaluators. Contribution/Results: We introduce the first evaluation-centric process model and reference architecture for LLM agent development. It uniquely supports open-behavior capture, emergent-result governance, and dynamic alignment. Experiments demonstrate that the framework effectively enables safe, controllable, and continuous agent iteration under objective drift, requirement changes, and regulatory evolution—achieving robust adaptability without compromising reliability or compliance.
This work addresses the limitation of existing large language model (LLM) agents in multi-step tasks, which often employ short-sighted greedy strategies for tool selection and lack global planning capabilities that account for tool dependencies. To overcome this, the authors propose ToolTree, a Monte Carlo tree search-inspired framework for tool planning. ToolTree incorporates a two-stage LLM evaluation process and a bidirectional pruning mechanism applied before and after execution, enabling efficient yet forward-looking exploration of tool usage trajectories. The method explicitly models inter-tool dependencies to support multi-step reasoning and planning. Experimental results demonstrate that ToolTree achieves state-of-the-art performance on both open-set and closed-set tool planning tasks, with an average improvement of approximately 10% across four benchmark datasets.
This work addresses the limitations of large language model (LLM) agents in tool invocation, which are highly dependent on the quality of human-written tool interface descriptions and suffer significant performance degradation in cold-start scenarios with numerous candidate tools or absent execution traces. To overcome this, the authors propose Trace-Free+, a framework that leverages curriculum learning to transfer supervised knowledge from trace-rich environments to trace-free deployment settings. This approach guides the model to learn reusable tool-use patterns and automatically refine tool descriptions without relying on execution trajectories. Trace-Free+ supports cross-tool generalization and scales effectively to tool sets comprising hundreds of functions. Experiments on StableToolBench and RestBench demonstrate that Trace-Free+ substantially improves invocation accuracy on unseen tools, exhibiting strong cross-domain generalization and robustness at scale.
Existing large language model agent systems struggle to meet the demands of production environments—such as simplicity, controllability, and predictable inference costs—due to their high complexity, unbounded reasoning expenses, and unpredictable behavior. To address these limitations, this work proposes a practical, utility-driven agent design framework that employs “pseudo-tools” to enforce modularity, replaces dynamic planning with fixed workflows, and integrates a dedicated learning algorithm to jointly optimize component performance. The approach innovatively applies multi-objective optimization to balance inference cost and response quality, while supporting result fusion across multiple systems. Experimental results demonstrate that the proposed method significantly reduces inference costs and improves accuracy across diverse tasks, outperforming handcrafted dynamic planning baselines.
This work addresses the challenge that large language models struggle to efficiently plan tool usage in complex tasks due to implicit reasoning and dynamic environmental changes. It introduces a novel approach that models tool relationships at the schema level by constructing a tool–schema hypergraph, where each tool is represented as a hyperedge connecting input and output schema nodes. The method further incorporates a task-relevant context graph, a schema-aware task-directed acyclic graph (DAG), and a gap-driven expansion mechanism conditioned on system state to enable precise dynamic planning. Evaluated on the AppWorld benchmark, this framework significantly improves task completion rates while simultaneously reducing redundant API calls, LLM interactions, and token consumption.
Current research on large language model (LLM) agents lacks a unified formal framework, resulting in conceptual and methodological ambiguity that hinders implementation-agnostic analysis and comparison. To address this gap, this work proposes the Structural Context Model—a novel formalism that integrates a declarative implementation framework with a semantic dynamic analysis workflow. This approach yields, for the first time, an analyzable, self-consistent, and implementation-independent formal model of LLM agents, enabling systematic design and iteration across their entire lifecycle. Empirical validation on the highly challenging dynamic Monkey-and-Bananas problem demonstrates a 32-percentage-point improvement in task success rate for agents developed within this framework, underscoring both its theoretical rigor and practical engineering utility.