Score
Design and implement integrations that call, orchestrate, and monitor LLM APIs and connected tools, including building connectors for multimodal inputs, tool invocation pipelines, and reliable API orchestration logic. Create and iterate prompt strategies and tuning workflows—authoring prompt templates, prompting techniques, and LLM-in-the-loop control for simulation orchestration—and test integrations for correctness, robustness, and expected behavior under varied inputs.
In the era of large language models, agent workflows face critical challenges in scalability, controllability, and security. To address these, this paper presents a systematic literature review and proposes, for the first time, a dual-dimensional taxonomy—spanning functional capabilities (task planning, multi-agent collaboration, tool integration) and architectural characteristics (role definition, orchestration process, specification languages). Through comparative analysis of over twenty representative academic and industrial systems, we identify recurring design patterns and persistent technical bottlenecks. We further introduce security-enhanced orchestration optimization strategies and pinpoint core gaps, including the lack of standardization and insufficient multimodal integration. This work establishes a foundational theoretical framework and practical guidelines for the design, evaluation, and evolution of agent workflows, advancing the field toward structured, trustworthy, and multimodal-cooperative paradigms.
Current tool-augmented large language model (LLM) ecosystems suffer from fragmentation—characterized by coexisting heterogeneous protocols (e.g., OpenAI Function Calling, Toolformer), manual schema definition, and complex execution orchestration—leading to low development efficiency and high integration overhead. To address this, we propose a protocol-agnostic unified tool integration framework. Our approach introduces an abstract protocol layer for cross-standard compatibility, an automated schema inference mechanism to eliminate manual specification, and a dual-mode concurrent scheduler enabling seamless synchronous and asynchronous tool execution. Experimental evaluation demonstrates that, compared to baseline approaches, our framework reduces implementation code volume by 60–80%, achieves up to 3.1× improvement in end-to-end execution latency, and maintains full backward compatibility with mainstream LLM tool-calling ecosystems.
This work addresses the trade-off between response quality and execution cost in large language model (LLM) agents when using external tools, a challenge often exacerbated by existing approaches that are either overly rigid or prone to redundant calls and high latency. The authors propose a utility-guided explicit orchestration framework that models agent behavior as a multi-action decision problem, dynamically selecting among responding, retrieving, invoking tools, verifying outcomes, and halting execution. This framework explicitly balances utility, cost, uncertainty, and redundancy, eschewing implicit control via prompt engineering in favor of a lightweight redundancy suppression mechanism and a utility-driven dynamic scheduling algorithm. Experimental results demonstrate that the proposed approach effectively governs the cost–performance trade-off, significantly improving agent behavior across diverse baselines and strategy variants, thereby validating the efficacy of utility-aware design in controllable agent orchestration.
This work addresses the challenge of coordinating multiple tools with large language models to accomplish long-horizon tasks in complex, dynamic environments—a setting where prior research has largely been confined to single-tool or one-step interactions. We formulate multi-tool coordination as dynamic orchestration over extended trajectories and introduce a unified task definition alongside a structured analytical framework. Through six core dimensions—planning and execution, training methodologies, safety, efficiency, capability completeness, and evaluation benchmarks—we systematically review key techniques, including reasoning-time planning, trajectory-based training, safety-aware control, resource-constrained optimization, and modeling in open-ended environments. By synthesizing insights from software engineering, enterprise workflows, GUI automation, and mobile systems, we delineate the fundamental challenges and future directions for building reliable, scalable, and verifiable multi-tool agents.
This work addresses the fragmentation and incompatibility across multiple LLM providers in tool-calling interfaces, message formats, and streaming behaviors, which hinder the portability and reproducibility of agent systems. To resolve this, the authors propose Orchestral, a lightweight Python framework that enables cross-LLM agent development through a unified, type-safe interface abstraction. Its core innovations include automatic tool schema generation driven by Python type hints, a synchronous streaming execution model that balances determinism with interactivity, and a modular, decoupled architecture. Orchestral supports standardized message and tool representations, context compression, sandboxed workspaces, MCP integration, and sub-agent mechanisms, substantially reducing engineering complexity while enhancing system portability, maintainability, and functional completeness.
This work addresses the limitations of existing agent workflows, which predominantly rely on abstract structures from large language models and lack genuine tool integration, resulting in poor usability and stability. To overcome this, we propose FlowScout, a novel framework that explicitly models real-world tool invocations as nodes in a directed graph. FlowScout integrates tool coordination skeleton mining with a Monte Carlo Tree Search mechanism guided by execution feedback to automatically optimize workflow topology. Experimental results across four task domains demonstrate that FlowScout significantly outperforms baseline methods—including PM4Py, ReAct, and AFlow—with at least a 92.69% improvement in tool invocation accuracy, a minimum 17.66% gain in execution quality, and enhanced runtime stability.
This study systematically investigates the capability boundaries of large language models (LLMs) in security tool orchestration, with a focus on the relative impact of model choice, client implementation, toolset composition, and reasoning mechanisms on system performance. Leveraging the open-source orchestration framework HexStrike-AI, the authors conduct multi-configuration comparative experiments across 86 picoCTF challenges, complemented by failure diagnosis and targeted refinements—including tool corrections, behavioral adjustments, and capability extensions—to quantitatively demonstrate, for the first time, the critical role of the client component in determining the performance of a fixed LLM. Results indicate that performance bottlenecks primarily stem from reasoning or environmental constraints rather than missing tools, enabling an increase in overall solve rate from 55.4% to 72.0% (p < 0.001) with high reproducibility (17 out of 20 trials consistent). The work introduces a reproducible evaluate-and-improve feedback loop, establishing a new paradigm for intelligent security agent systems.
This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.