Score
Designs, builds, and evaluates software, scripts, and toolchains that automate repetitive tasks, workflows, or system operations, including task orchestration, scheduling, error handling, retries, and monitoring. Implements integrations, plugins, and user interfaces that connect systems and measure automation effectiveness and reliability.
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
This study addresses the significant burden developers face in authoring and maintaining GitHub Actions workflows, stemming from a lack of systematic understanding of real-world automation and reuse practices. Through a mixed-methods approach combining a survey of 419 practitioners with qualitative and quantitative analysis, this work presents the first developer-centric characterization of common automation tasks, patterns of reuse mechanism adoption, and maintenance pain points in workflow development. The findings reveal that while developers heavily rely on reusable Actions, they seldom adopt reusable workflows; version management challenges lead to rampant copy-pasting; and critical aspects such as security and performance monitoring remain under-automated. These insights provide empirical foundations for improving CI/CD toolchains and reuse mechanisms.
Early-stage software development—spanning requirements elicitation, testing, and deployment—is hindered by ill-defined tasks and dense manual intervention points, impeding automation. Traditional CI/CD pipelines address only post-coding phases, leaving semantic gaps between underspecified stages unbridged. Method: We propose “workflow-as-software,” a novel paradigm that models end-to-end development as programmable workflows. Leveraging large language models (LLMs) as universal semantic adapters, our approach automatically reconciles heterogeneous task semantics. It integrates domain-specific workflow orchestration, a lightweight domain-specific language (DSL), and semantic translation interfaces. Contribution/Results: Evaluated in production at Volvo, the method reduced test automation effort by 2–3 full-time engineers and compressed the end-to-end development-to-deployment cycle to two months. It marks the first demonstration of LLM-driven, fully automated software delivery across the entire lifecycle—from requirements to deployment—thereby extending automation beyond conventional CI/CD boundaries.
This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.
This work addresses the limitations of existing agent workflows, which predominantly rely on abstract structures from large language models and lack genuine tool integration, resulting in poor usability and stability. To overcome this, we propose FlowScout, a novel framework that explicitly models real-world tool invocations as nodes in a directed graph. FlowScout integrates tool coordination skeleton mining with a Monte Carlo Tree Search mechanism guided by execution feedback to automatically optimize workflow topology. Experimental results across four task domains demonstrate that FlowScout significantly outperforms baseline methods—including PM4Py, ReAct, and AFlow—with at least a 92.69% improvement in tool invocation accuracy, a minimum 17.66% gain in execution quality, and enhanced runtime stability.
Current evaluations of AI agents lack comprehensive assessment of cross-application coordination, autonomous API discovery, and policy adherence, failing to reflect the demands of real-world business automation. This work proposes the first unified benchmark that integrates cross-application workflow orchestration, autonomous API exploration, and regulatory compliance. Built upon the Zapier platform, the benchmark encompasses multi-scenario tasks spanning CRM, email, calendar, and other systems, requiring agents to autonomously discover APIs, follow multi-layered business rules, and execute cross-system data writes in realistic sales and marketing workflows. Evaluation relies solely on end-state correctness through an automated scoring mechanism. Experimental results reveal that even state-of-the-art models achieve success rates below 10%, highlighting a significant capability gap in deploying current AI systems for practical business automation.
This study addresses the ambiguity in responsibility and agency between AI coding agents and human developers during pull request (PR) lifecycles, where proactive AI actions intersect with human-led merge governance. The authors propose an “Initiator × Approver” taxonomy and construct a collaboration–assistance spectrum alongside state-machine models of various tools. Through systematic log analysis of 29,585 PRs, they disentangle the distinct roles of AI and humans in the PR workflow. Their findings reveal that over 96% of PRs in collaborative tools are initiated by AI, yet merge decisions remain almost exclusively under human control. While automated merges record execution behavior, they do not engage with core governance functions. This work thus provides the first clear delineation of operational boundaries and governance demarcation for AI coding agents.
This study addresses the critical issue of frequent failures in GitHub Actions workflows, which severely undermine automation reliability and maintainability. For the first time, it systematically maps 197 language constructs to 14 workflow capability features through a large-scale quantitative analysis of over 260,000 workflows across 49,000 repositories. By integrating language construct categorization with metadata mining, the work uncovers prevalent usage patterns, evolutionary trends, and their impact on workflow reliability. The findings reveal that only a small subset of constructs is heavily used, and that specific capability features are significantly associated with elevated failure rates and maintenance costs. These empirical insights provide actionable guidance for optimizing workflow design and improving robustness in continuous integration and delivery pipelines.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).