Score
Designs, builds, and analyzes version‑controlled operational workflows and practices that use Git as the single source of truth to drive automated deployment, configuration, and reconciliation of systems; defines repository layouts, CI/CD pipelines, branch and PR policies, automation tooling or controllers, and procedures for managing declarative infrastructure and application delivery.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
This paper addresses the conceptual ambiguity, ill-defined boundaries, and lack of implementation standards between Infrastructure-as-Code (IaC) and Pipeline-as-Code in DevOps practice. To resolve these issues, we systematically delineate their respective roles and synergistic mechanisms within the DevOps ecosystem and propose a reusable, standardized IaC-driven CI/CD implementation framework. Our approach integrates Terraform for infrastructure provisioning, Ansible for configuration management, GitLab CI for pipeline orchestration, and Docker/Kubernetes for containerized deployment—enabling an end-to-end automated delivery pipeline. Empirical evaluation demonstrates 99.8% configuration change accuracy, reduces environment provisioning time from hours to minutes, and significantly improves deployment consistency and delivery efficiency.
This study addresses the lack of systematic understanding regarding the evolution of GitHub Actions workflows. Through a mixed-methods approach, we conduct the first large-scale empirical analysis of over 3.4 million workflow file versions from more than 49,000 repositories spanning November 2019 to August 2025. We identify seven categories of conceptual changes and find that repositories typically contain a median of three workflow files, with 7.3% of workflows modified weekly—approximately 75% of which involve only a single change, predominantly in task configuration and specification. Our findings further indicate that current large language model (LLM) tools have not yet significantly influenced workflow maintenance frequency, offering empirical grounding for the design of fine-grained automated maintenance tools.
This study addresses the significant burden developers face in authoring and maintaining GitHub Actions workflows, stemming from a lack of systematic understanding of real-world automation and reuse practices. Through a mixed-methods approach combining a survey of 419 practitioners with qualitative and quantitative analysis, this work presents the first developer-centric characterization of common automation tasks, patterns of reuse mechanism adoption, and maintenance pain points in workflow development. The findings reveal that while developers heavily rely on reusable Actions, they seldom adopt reusable workflows; version management challenges lead to rampant copy-pasting; and critical aspects such as security and performance monitoring remain under-automated. These insights provide empirical foundations for improving CI/CD toolchains and reuse mechanisms.
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
This study addresses the critical issue of frequent failures in GitHub Actions workflows, which severely undermine automation reliability and maintainability. For the first time, it systematically maps 197 language constructs to 14 workflow capability features through a large-scale quantitative analysis of over 260,000 workflows across 49,000 repositories. By integrating language construct categorization with metadata mining, the work uncovers prevalent usage patterns, evolutionary trends, and their impact on workflow reliability. The findings reveal that only a small subset of constructs is heavily used, and that specific capability features are significantly associated with elevated failure rates and maintenance costs. These empirical insights provide actionable guidance for optimizing workflow design and improving robustness in continuous integration and delivery pipelines.
This work addresses the limitations of existing CI/CD workflow analyses, which often focus narrowly on stage identification and struggle to assess reliability, maintainability, and optimization priorities. To overcome this, we propose a large language model–based CI/CD analysis pipeline that integrates repository context enhancement, anti-pattern detection, stage mining, and actionable recommendation generation. Our approach uniquely combines diagnostic reasoning, context awareness, and human-in-the-loop review to deliver observability tailored to cybersecurity engineering. Leveraging few-shot prompting, YAML parsing, and statistical tests (chi-square and Cramér’s V), the method identifies 434,769 anti-patterns across 75,201 workflows and generates an average of 8.25 syntactically valid optimization suggestions per repository, achieving a 96.1% compliance rate with YAML syntax standards.
This study addresses the governance failures arising from the disconnect between enterprise architecture models and engineering workflows by proposing a Git-native, Architecture-as-Code (EA-as-Code) continuous governance framework. The framework models architectural facts in YAML, enforces type checking and rule-based governance according to the ArchiMate 3.2 standard, performs change impact analysis via graph traversal algorithms, and establishes an automated pipeline loop using GitHub Actions. Experimental results demonstrate that the system achieves an F1 score of 1.0 for fault detection and completes validation in approximately 52 seconds at a scale of 50,000 objects, confirming its effectiveness, practicality, and scalability in large-scale scenarios.
This work addresses the challenge that developers often make errors when performing complex Git operations, and existing large language models (LLMs) struggle to ensure correctness and safety due to their limited formal reasoning capabilities. To overcome this limitation, the paper proposes a novel approach that integrates automated planning with LLMs, enabling the system to interpret natural language instructions and formally model the state of a Git repository to generate safe, verifiable command sequences. By combining symbolic reasoning with language understanding, the method significantly improves the success rate of Git operations compared to pure LLM-based solutions. Experimental results demonstrate consistent performance gains across multiple evaluation metrics, thereby enhancing both the reliability and interpretability of AI-powered developer assistance tools.
研究通过分析GitHub Agentic Workflows的结构和维护方式,探讨了开发者如何定义和维护由AI代理执行的工作流程,并建议增加防御措施。