Score
Designs, builds, and analyzes end-to-end operational workflows that incorporate AI components to augment or automate human tasks; this includes specifying how tasks are split between humans and models, the data inputs/outputs and transformation steps, model inference and orchestration, interface and handoff points, error handling and fallback behavior, and monitoring and evaluation metrics for the combined human–AI process.
In the era of large language models, agent workflows face critical challenges in scalability, controllability, and security. To address these, this paper presents a systematic literature review and proposes, for the first time, a dual-dimensional taxonomy—spanning functional capabilities (task planning, multi-agent collaboration, tool integration) and architectural characteristics (role definition, orchestration process, specification languages). Through comparative analysis of over twenty representative academic and industrial systems, we identify recurring design patterns and persistent technical bottlenecks. We further introduce security-enhanced orchestration optimization strategies and pinpoint core gaps, including the lack of standardization and insufficient multimodal integration. This work establishes a foundational theoretical framework and practical guidelines for the design, evaluation, and evolution of agent workflows, advancing the field toward structured, trustworthy, and multimodal-cooperative paradigms.
Large language models (LLMs) exhibit hallucination and operational fragility in technical service delivery, undermining reliability and safety. Method: This paper proposes a human–AI collaborative authoring framework tailored for technical services. Drawing inspiration from autonomous driving’s levels of autonomy, we introduce a taxonomy of six interaction paradigms—HOOTL (Human-Over-Open-Task Loop), HIC (Human-in-Control), HITP (Human-in-Task-Planning), HITL (Human-in-the-Loop), HOTL (Human-on-the-Loop), and HAM (Human-as-Monitor)—and develop a contingency-based mode selection framework grounded in task complexity, operational risk, and system reliability. Contribution/Results: Evaluated via LLM-agent experiments and multi-case studies, the framework provides technical service platforms with an actionable responsibility-allocation tool that enhances automation efficiency, safety, and contextual adaptability while preserving human oversight and control.
AI agents lack fundamental understanding of human work practices, limiting their effective integration into collaborative workflows. Method: This paper introduces the first cross-domain (data analysis, engineering, computing, writing, design) systematic framework for comparing human and AI workflows, featuring a novel screen-operation-trajectory-based workflow reconstruction toolkit that enables structured, interpretable behavioral modeling. Contribution/Results: Empirical analysis reveals that AI agents predominantly follow rigid, programmatic execution paths—diverging markedly from human interface interaction patterns, even in vision-intensive tasks. Although AI achieves 88.3% faster task completion and reduces costs by 90.4%–96.2%, it exhibits lower output quality, including data fabrication and misuse of advanced tools. Its strengths are confined to programmable subtasks; thus, AI is best positioned as a high-efficiency collaborator within human-led workflows—not as an autonomous replacement.
To address the lack of formal modeling support for human–AI collaborative workflows, this paper proposes a BPMN extension tailored for hybrid collaboration between humans and large language model (LLM)-based agents. We introduce the first BPMN metamodel and corresponding graphical notation explicitly supporting human–agent coordination, formally capturing task responsibility assignment, decision logic ownership, and execution strategy binding—thereby bridging a critical expressiveness gap in existing business process modeling languages for mixed-agent scenarios. Furthermore, we design a domain-specific language (DSL) and implement an open-source web-based modeling tool using TypeScript and React, enabling visual specification and formal verification of collaborative processes. Empirical evaluation through case studies demonstrates that the approach achieves strong expressive power, operational feasibility, and practical applicability in real-world human–LLM workflow engineering.
This work addresses the limitations of existing human-in-the-loop (HITL) mechanisms in intelligent agent workflows, which are often tightly coupled with application logic, resulting in poor reusability, weak consistency, and limited scalability. To overcome these challenges, the paper proposes a decoupled HITL system architecture that abstracts human oversight into an independent component. By introducing explicit interfaces and a structured execution model, the approach cleanly separates human–machine interaction from business logic. Furthermore, it introduces a novel four-dimensional framework—comprising intervention conditions, role resolution, interaction semantics, and communication channels—to enable context-aware, controllable human intervention. This design achieves, for the first time, protocol-level reusability of HITL mechanisms, supporting consistent and scalable autonomy governance in multi-agent environments and laying a foundational infrastructure for system-level human–agent collaboration.
Enterprise operational workflows are notoriously difficult to automate end-to-end due to their heavy reliance on human intervention and limited adaptability to change. This work proposes the first action-centric workflow graph framework, which achieves automated construction, execution, and evolution through a three-stage pipeline: structured workflow graphs are extracted from human operation traces, executed via multi-agent online traversal, and continuously optimized in a closed loop using an Adaptive Traversal Reinforcement (ATR) mechanism. Integrating large-scale offline graph construction, graph-guided retrieval, and large language model reasoning, the approach was deployed across four cloud database services. It substantially outperforms the Trace-RAG baseline in coverage breadth, factual accuracy, and diagnostic throughput, achieving an expert blind-review score of 4.95 out of 5.
This study addresses the ambiguity in responsibility and agency between AI coding agents and human developers during pull request (PR) lifecycles, where proactive AI actions intersect with human-led merge governance. The authors propose an “Initiator × Approver” taxonomy and construct a collaboration–assistance spectrum alongside state-machine models of various tools. Through systematic log analysis of 29,585 PRs, they disentangle the distinct roles of AI and humans in the PR workflow. Their findings reveal that over 96% of PRs in collaborative tools are initiated by AI, yet merge decisions remain almost exclusively under human control. While automated merges record execution behavior, they do not engage with core governance functions. This work thus provides the first clear delineation of operational boundaries and governance demarcation for AI coding agents.
This work addresses the challenges of semantic drift and poor auditability in multi-step AI agent automation, which often arise from error propagation and sensitivity to prompt variations. The authors propose a framework that treats a visual workflow graph as the single source of truth, leveraging formal semantic definitions at compile time to specify data scopes, execution logic, and monitoring policies, while employing contract-based verification—via preconditions and postconditions—to ensure component correctness. At runtime, the system supports replay, retry, and auditing through persistent event logging and explicitly isolates risk using lane-based trust boundaries. Evaluated across three clinical institutions over 8,728 executions, the prototype achieved a 97.08% completion rate, with most failures attributable to external system integration issues.