Score
Designs, implements, and evaluates executable sequences of tasks and their orchestration (workflows), including task dependencies, data and control flows, triggers, scheduling, error handling, and monitoring. Builds workflow specifications, engines, integrations, and automation pipelines to coordinate services, human steps, and data transformations.
This work addresses the challenge that large language models (LLMs) struggle to reliably translate free-form reasoning into structured workflows when handling complex tasks. To this end, we propose the Execute-Summarize framework, which decouples task execution from workflow generation for the first time: the LLM first executes the task and records its execution trace, and a separate module then reconstructs a structured workflow solely from this trace. This approach significantly enhances both the accuracy and robustness of the resulting workflows. We also introduce FlowBench, a new benchmark designed to systematically evaluate workflow generation capabilities. Experimental results demonstrate that our framework substantially outperforms existing methods on FlowBench, offering a reliable paradigm for converting LLM-based reasoning into structured, executable processes.
This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.
研究通过分析GitHub Agentic Workflows的结构和维护方式,探讨了开发者如何定义和维护由AI代理执行的工作流程,并建议增加防御措施。
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
为解决自然语言工作流执行不可靠的问题,提出Artic编译器,将自然语言工作流转换为基于工件的工作流,提高任务解决率和一致性。
This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.
This work addresses the limitations of existing agent workflows, which predominantly rely on abstract structures from large language models and lack genuine tool integration, resulting in poor usability and stability. To overcome this, we propose FlowScout, a novel framework that explicitly models real-world tool invocations as nodes in a directed graph. FlowScout integrates tool coordination skeleton mining with a Monte Carlo Tree Search mechanism guided by execution feedback to automatically optimize workflow topology. Experimental results across four task domains demonstrate that FlowScout significantly outperforms baseline methods—including PM4Py, ReAct, and AFlow—with at least a 92.69% improvement in tool invocation accuracy, a minimum 17.66% gain in execution quality, and enhanced runtime stability.
This work addresses the limitations of general-purpose large language models in business process automation, where inconsistent functionality, frequent tool-calling errors, and unstable code quality hinder industrial-grade reliability and maintainability. Focusing on the task of translating BPMN diagrams into executable agent workflows, we propose the first specialized agent system designed for structured code generation. By integrating BPMN control-flow semantics, a deterministic path execution mechanism, and a lightweight code generation strategy, our approach achieves high-precision, low-latency, and zero-repair automation. Experimental results demonstrate that, compared to general-purpose models, our method improves tool-calling accuracy by 9–20 percentage points, reduces latency by 2–4×, decreases calling errors by a factor of three, lowers token-generation costs by over 95%, and entirely eliminates the need for repair iterations.
This study addresses the absence of a general principled framework for schema engineering workflows and the difficulty existing systems face in composably representing schema operations. To overcome these limitations, this work proposes a typed logic-based workflow framework. By defining typed representations of fixed components, the approach translates schema operations into composable query chains, thereby decoupling workflow definitions from execution constraints. Furthermore, it designs an LLM agent-driven execution pipeline to enable automated analysis. Experiments conducted on the WIPO dataset successfully execute three categories of query tasks, including exemplar mining. The results validate performance variations across different LLM configurations alongside their error localization capabilities. Ultimately, this research establishes a modular and orchestrable paradigm for schema engineering.
研究通过开发Trace2Flow将AI执行过程转化为可编辑的图形表示,以提高用户对AI执行过程的理解和复用能力。