Score
Designs, documents, and implements version-control workflows and branching strategies to manage collaboration, merges, pull requests, and branch protection across single and multiple repositories. Builds and configures Git/GitHub workflow automation and integrations — including GitHub workflow configs and GitOps-style pipelines — to orchestrate CI/CD triggers, cross-repository coordination, and repo-level workflow management.
This study addresses the lack of systematic understanding regarding the evolution of GitHub Actions workflows. Through a mixed-methods approach, we conduct the first large-scale empirical analysis of over 3.4 million workflow file versions from more than 49,000 repositories spanning November 2019 to August 2025. We identify seven categories of conceptual changes and find that repositories typically contain a median of three workflow files, with 7.3% of workflows modified weekly—approximately 75% of which involve only a single change, predominantly in task configuration and specification. Our findings further indicate that current large language model (LLM) tools have not yet significantly influenced workflow maintenance frequency, offering empirical grounding for the design of fine-grained automated maintenance tools.
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
This study investigates the contextual applicability of Trunk-Based Development (TBD) versus Branch-Based Development (BBD) workflows to optimize developer productivity and software quality across diverse team settings. Method: We conducted a mixed-methods empirical study with 127 practitioners in Brazil, comprising semi-structured interviews (n=15) and an online survey (n=112), analyzed via qualitative coding and quantitative statistical methods. Contribution/Results: To our knowledge, this is the first systematic, industry-based comparison of TBD and BBD practices. We identify team size and developer experience as critical moderating factors: TBD significantly improves delivery efficiency in small, experienced teams, whereas BBD better accommodates large-scale or junior-heavy teams—at the cost of higher process management overhead. The findings yield a data-driven decision framework to guide workflow selection based on organizational context.
This study addresses the significant burden developers face in authoring and maintaining GitHub Actions workflows, stemming from a lack of systematic understanding of real-world automation and reuse practices. Through a mixed-methods approach combining a survey of 419 practitioners with qualitative and quantitative analysis, this work presents the first developer-centric characterization of common automation tasks, patterns of reuse mechanism adoption, and maintenance pain points in workflow development. The findings reveal that while developers heavily rely on reusable Actions, they seldom adopt reusable workflows; version management challenges lead to rampant copy-pasting; and critical aspects such as security and performance monitoring remain under-automated. These insights provide empirical foundations for improving CI/CD toolchains and reuse mechanisms.
This study presents the first large-scale empirical analysis of GitHub Actions (GHA) workflows across multilingual open-source projects (Java, Python, C++), addressing three core challenges: workflow structural complexity, cross-language heterogeneity, and deviations from official CI best practices. Leveraging static analysis, pattern mining, and compliance checking, we construct and analyze a dataset comprising over 10,000 real-world GHA workflows. Results reveal pervasive structural issues—including redundant steps, unjustified parallelization, and fragmented environment configurations—and significant inter-language disparities in workflow design patterns and adherence to guidelines. Quantitatively, Python projects exhibit the highest compliance rate, while C++ projects show the lowest. Based on these findings, we propose language-specific, lightweight refactoring guidelines and automated compliance detection strategies. This work establishes an empirical foundation and actionable pathways for improving CI maintainability, standardization, and cross-language interoperability in modern software development.
This study addresses the critical issue of frequent failures in GitHub Actions workflows, which severely undermine automation reliability and maintainability. For the first time, it systematically maps 197 language constructs to 14 workflow capability features through a large-scale quantitative analysis of over 260,000 workflows across 49,000 repositories. By integrating language construct categorization with metadata mining, the work uncovers prevalent usage patterns, evolutionary trends, and their impact on workflow reliability. The findings reveal that only a small subset of constructs is heavily used, and that specific capability features are significantly associated with elevated failure rates and maintenance costs. These empirical insights provide actionable guidance for optimizing workflow design and improving robustness in continuous integration and delivery pipelines.
研究通过分析GitHub Agentic Workflows的结构和维护方式,探讨了开发者如何定义和维护由AI代理执行的工作流程,并建议增加防御措施。
This study addresses widespread compliance issues in GitHub Actions workflows, such as excessive permissions and weak secret management. It proposes the first documentation-driven compliance checking framework, which derives a 30-item checklist from official documentation and implements a hybrid auditing pipeline combining large language models (LLMs) with expert oversight. The authors automatically evaluate 95 real-world Java workflows using four open-source LLMs, employ GPT-5 as a conflict arbitrator, and integrate manual review into a multi-tiered validation system. Experimental results reveal an overall compliance rate of only 28%, with permission control as low as 4%. The proposed approach reduces manual verification effort by 81% while achieving 87% agreement with expert judgments, significantly enhancing audit efficiency and reproducibility.
This study addresses the ambiguity in responsibility and agency between AI coding agents and human developers during pull request (PR) lifecycles, where proactive AI actions intersect with human-led merge governance. The authors propose an “Initiator × Approver” taxonomy and construct a collaboration–assistance spectrum alongside state-machine models of various tools. Through systematic log analysis of 29,585 PRs, they disentangle the distinct roles of AI and humans in the PR workflow. Their findings reveal that over 96% of PRs in collaborative tools are initiated by AI, yet merge decisions remain almost exclusively under human control. While automated merges record execution behavior, they do not engage with core governance functions. This work thus provides the first clear delineation of operational boundaries and governance demarcation for AI coding agents.
This study presents the first systematic evaluation of reliability differences among multiple AI agents—Claude, Devin, Cursor, Copilot, and Codex—in GitHub Actions CI/CD workflows. Leveraging the AIDev dataset, the authors collected 61,837 workflow runs via the GitHub Actions API and integrated CI logs, pull request metadata, and commit data to construct a taxonomy of 13 distinct failure causes. Their analysis reveals that Copilot and Codex achieve the highest success rates (93%–94%), while the frequency of AI contributions exhibits a significant negative correlation with workflow success. Moreover, high-frequency AI involvement is associated with an increased likelihood of specific failure types. These findings provide empirical evidence and a practical framework for integrating AI-generated code into CI/CD pipelines, particularly in high-stakes development scenarios.