Score
Designs and implements frameworks, APIs, and runtime orchestration that enable an application or agent to discover, invoke, coordinate, and manage external tools and functions; this includes specifying invocation interfaces, invocation patterns (synchronous, asynchronous, parallel), workflows, retries, batching, and integration points. Builds the tooling and middleware that route inputs/outputs, enforce interface contracts, handle errors and monitoring, and support tool-calling techniques and orchestration across components.
In the era of large language models, agent workflows face critical challenges in scalability, controllability, and security. To address these, this paper presents a systematic literature review and proposes, for the first time, a dual-dimensional taxonomy—spanning functional capabilities (task planning, multi-agent collaboration, tool integration) and architectural characteristics (role definition, orchestration process, specification languages). Through comparative analysis of over twenty representative academic and industrial systems, we identify recurring design patterns and persistent technical bottlenecks. We further introduce security-enhanced orchestration optimization strategies and pinpoint core gaps, including the lack of standardization and insufficient multimodal integration. This work establishes a foundational theoretical framework and practical guidelines for the design, evaluation, and evolution of agent workflows, advancing the field toward structured, trustworthy, and multimodal-cooperative paradigms.
This work addresses the lack of a unified framework in current LLM agent workflows, which hinders method comparison and reproducibility. To resolve this, we propose the Agent Computation Graph (ACG) framework, which models workflows as computation graphs and adopts “structure determines timing” as a core principle. The framework explicitly distinguishes between reusable templates, runtime instance graphs, and execution traces, enabling a systematic categorization of static and dynamic optimization approaches. Through a comprehensive literature review and conceptual modeling, we develop a multidimensional evaluation framework that integrates structural properties, establishes precise terminology, and defines standardized evaluation criteria. This foundation supports a reproducible and highly comparable research paradigm for optimizing LLM agent workflows.
Existing intelligent agents often suffer from high system fragility and substantial execution overhead in multi-tool coordination due to inadequate scheduling mechanisms. This work proposes a hierarchical orchestration paradigm that obviates the need for fine-grained dependency graphs by providing coarse-grained global guidance, coupled with context-constrained intra-layer execution. A pattern-aware local reflection mechanism is introduced to enable runtime error detection and repair without triggering costly global replanning. The approach significantly enhances the robustness of tool invocation while reducing execution complexity and resource consumption, yielding a lightweight and reusable tool orchestration component.
Existing LLM agent systems suffer from tight coupling between logical workflows and underlying programming languages or deployment environments, resulting in high development/deployment complexity and poor maintainability. This paper proposes a declarative domain-specific language (DSL) tailored for LLM agent workflows—marking the first effort to universally abstract and unify common patterns such as RAG, API orchestration, and filtering. The DSL fully decouples workflow specification from execution semantics, enabling cross-language (Java/Python/Go) and cross-environment (cloud-native/on-premises) deployment. The system integrates multi-backend adapters, a lightweight workflow engine, and an automated metrics collection framework, natively supporting multi-strategy A/B testing and performance benchmarking. Evaluated in PayPal’s e-commerce setting, it reduces development time by 60% and accelerates deployment threefold; complex workflows shrink from >500 to <50 lines of code, achieve orchestration latency under 100 ms, and allow safe, non-engineer configuration.
Existing multi-agent collaborative systems are hindered by static workflows, sequential scheduling, and heterogeneous interfaces, leading to high complexity and poor scalability. This work proposes Agent-as-Tool, a unified paradigm that abstracts both agents and tools into a standardized, learnable action space, and introduces ParaManager—a lightweight coordinator enabling state-aware parallel subtask decomposition, delegation, and asynchronous execution. By unifying communication protocols and incorporating explicit state feedback, the framework facilitates efficient multi-agent collaboration. A two-stage training strategy—combining supervised fine-tuning with a recovery mechanism and reinforcement learning—optimizes task success rate, protocol compliance, response diversity, and reasoning efficiency. Experiments demonstrate that ParaManager achieves strong performance across multiple benchmarks and exhibits robust generalization to unseen agent pools.
Existing tool interfaces based on static endpoints struggle to express long-running workflows involving complex control flows such as loops, conditional branches, and retries. This work proposes replacing static endpoints with executable tool programs, enabling explicit effect typing and sophisticated workflow control through constraint-guided program construction, effect-aware exactly-once replay mechanisms, and configuration-driven execution policies. Implemented atop MCP-style services and a WebAssembly sandbox, the system demonstrates significant performance improvements in real-world scenarios, reducing end-to-end latency by up to 53.4% and client-side traffic by as much as 96.1%, with particularly pronounced gains under high network latency or increased workflow complexity.
This work addresses the limitations of traditional workflow platforms, which rely on static, pre-defined processes and struggle to accommodate the dynamic data integration demands of distributed systems. To overcome this, the authors propose a configuration-driven runtime orchestration framework that dynamically constructs execution graphs at request time through dependency-aware scheduling and parallel task execution, thereby circumventing the constraints of fixed workflows. This approach enables rapid adaptation to evolving integration scenarios without requiring system redeployment, significantly reducing latency. Empirical evaluation in a real-world Customer 360 enterprise use case demonstrates that the framework offers substantial advantages in flexibility, scalability, and efficient data aggregation compared to conventional solutions.
This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.
This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.
This work addresses the limitations of existing agent workflows, which predominantly rely on abstract structures from large language models and lack genuine tool integration, resulting in poor usability and stability. To overcome this, we propose FlowScout, a novel framework that explicitly models real-world tool invocations as nodes in a directed graph. FlowScout integrates tool coordination skeleton mining with a Monte Carlo Tree Search mechanism guided by execution feedback to automatically optimize workflow topology. Experimental results across four task domains demonstrate that FlowScout significantly outperforms baseline methods—including PM4Py, ReAct, and AFlow—with at least a 92.69% improvement in tool invocation accuracy, a minimum 17.66% gain in execution quality, and enhanced runtime stability.