Understanding and Detecting Flaky Builds in GitHub Actions

πŸ“… 2026-02-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the prevalence and impact of non-deterministic (i.e., β€œflaky”) builds in GitHub Actions, which significantly undermine the reliability of continuous integration and waste computational resources. Leveraging re-execution data from 1,960 open-source Java projects, the work presents the first systematic characterization of flaky builds, revealing that 67.73% of re-run builds are affected, spanning 51.28% of the projects, and identifying 15 distinct root causes. To mitigate this issue, the authors propose a job-level machine learning approach for detecting flaky failures. Evaluated against state-of-the-art baselines, the method achieves up to a 20.3% improvement in F1-score, substantially enhancing the accuracy of flaky build identification.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Other Foundations of Reasoning under UncertaintyKnowledge Representation and Reasoning: Action, Change, and Causality

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAIResponsible Web: Machine-in-the-loop, human agency and autonomy
πŸ“ Abstract
Continuous Integration (CI) is widely used to provide rapid feedback on code changes; however, CI build outcomes are not always reliable. Builds may fail intermittently due to non-deterministic factors, leading to flaky builds that undermine developers'trust in CI, waste computational resources, and threaten the validity of CI-related empirical studies. In this paper, we present a large-scale empirical study of flaky builds in GitHub Actions based on rerun data from 1,960 open-source Java projects. Our results show that 3.2% of builds are rerun, and 67.73% of these rerun builds exhibit flaky behavior, affecting 1,055 (51.28%) of the projects. Through an in-depth failure analysis, we identify 15 distinct categories of flaky failures, among which flaky tests, network issues, and dependency resolution issues are the most prevalent. Building on these findings, we propose a machine learning-based approach for detecting flaky failures at the job level. Compared with a state-of-the-art baseline, our approach improves the F1-score by up to 20.3%.
Problem

Research questions and friction points this paper is trying to address.

flaky builds
continuous integration
GitHub Actions
build reliability
intermittent failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

flaky builds
GitHub Actions
machine learning
continuous integration
empirical study
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
W
Wenhao Ge
School of Computer Science and Technology, Soochow University, China
C
Chen Zhang
School of Computer Science and Technology, Soochow University, China