Score
Designs, builds, and submits code-level changes and artifacts to public repositories—such as patches, pull/merge requests, tests, documentation, and CI/configuration—so they can be reviewed and accepted upstream. Analyzes and improves contribution workflows, governance, standards, and cross-repository coordination to enable reliable upstream integration and maintainable long-term contributions.
This study addresses the lack of systematic understanding in the configuration and maintenance of CI/CD caching, which imposes a significant burden on developers despite its benefits for build efficiency. Through a large-scale empirical analysis of 952 repositories on GitHub Actions—encompassing 1,556 workflow files and over ten thousand cache-related changes—the authors employ code mining, configuration analysis, commit tracing, and statistical modeling to uncover real-world caching practices, evolutionary patterns, and human-bot collaboration in maintenance. The findings reveal that cache adopters are more active, caching strategies are diverse and frequently adjusted, and build- and test-related tasks evolve rapidly. Manual interventions primarily address misconfigurations, whereas version upgrades are predominantly automated by bots. The work quantifies the maintenance overhead of caching and provides empirical foundations for improving developer tooling.
This study addresses the lack of systematic understanding regarding the evolution of GitHub Actions workflows. Through a mixed-methods approach, we conduct the first large-scale empirical analysis of over 3.4 million workflow file versions from more than 49,000 repositories spanning November 2019 to August 2025. We identify seven categories of conceptual changes and find that repositories typically contain a median of three workflow files, with 7.3% of workflows modified weekly—approximately 75% of which involve only a single change, predominantly in task configuration and specification. Our findings further indicate that current large language model (LLM) tools have not yet significantly influenced workflow maintenance frequency, offering empirical grounding for the design of fine-grained automated maintenance tools.
Developers frequently deviate from alphabetical file ordering during GitHub code reviews, yet the prevalence and impact of such non-alphabetical navigation strategies remain poorly understood. Method: Analyzing 23,241 pull request (PR) review logs, we systematically identify and quantify three empirically grounded, non-alphabetical review strategies: largest-diff-first (20.6%), semantics-aligned-first (17.6%, matching PR titles/descriptions to file content via semantic similarity), and test-first (29%, especially prevalent in mixed-change PRs). Our approach integrates large-scale log mining, semantic text similarity computation, diff-size quantification, and statistical significance testing. Contribution/Results: We find that 44.6% of PRs employ non-alphabetical review orders—associated with higher file coverage but marginally lower approval rates—and that review sequence strongly correlates with PR complexity. This work provides the first empirical characterization of structured, real-world code review navigation patterns, offering data-driven foundations for intelligent IDE file ordering and automated review assistance tools.
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
GitHub’s CODEOWNERS feature automates code review responsibility assignment, yet its real-world adoption and impact remain poorly understood. This paper presents the first large-scale empirical study, analyzing 840,000 pull requests (PRs) and 2 million review logs across 2,147 open-source projects using Regression Discontinuity Design (RDD) to causally quantify CODEOWNERS’ effects. Results show that CODEOWNERS significantly improves review timeliness and coverage while reducing review burden on core developers; promotes more equitable ownership distribution and accelerates PR integration; and functions as a novel software governance mechanism that enhances project security and collaborative resilience. Collectively, this work demonstrates that automated ownership assignment substantively reshapes both collaboration efficiency and governance structures in open-source development.
This study addresses the prevalence and impact of non-deterministic (i.e., “flaky”) builds in GitHub Actions, which significantly undermine the reliability of continuous integration and waste computational resources. Leveraging re-execution data from 1,960 open-source Java projects, the work presents the first systematic characterization of flaky builds, revealing that 67.73% of re-run builds are affected, spanning 51.28% of the projects, and identifying 15 distinct root causes. To mitigate this issue, the authors propose a job-level machine learning approach for detecting flaky failures. Evaluated against state-of-the-art baselines, the method achieves up to a 20.3% improvement in F1-score, substantially enhancing the accuracy of flaky build identification.
This study addresses the limitation of existing coding agent benchmarks, which are largely confined to single repositories and fail to reflect real-world cross-repository collaborative development. We propose the first cross-repository evaluation framework for coding agents, constructing a benchmark of 120 real-world tasks by mining software ecosystem changes and adapting hidden tests for case generation. This work systematically compares independent and joint execution strategies, revealing the critical role of information sharing in achieving implementation completeness. Experimental results demonstrate that, using model configurations such as Codex CLI, the highest task success rate reaches 42.50%. Furthermore, joint execution more effectively leverages correlated repository information to guide code generation, significantly outperforming the independent mode. These findings establish a new paradigm for evaluating and enhancing the multi-repository collaboration capabilities of coding agents.
This study addresses the issue that coding agents, despite passing functional tests, frequently violate repository governance standards. To this end, we introduce SWE-CC, a benchmark for systematically evaluating both code and process compliance. Methodologically, we construct 823 machine-verifiable policies, design an auditing mechanism that integrates runtime behavior with final deliverables, and propose an evaluation framework combining semi-automated document conversion, deterministic checking, and LLM agent workflows. Experimental results reveal that modern agents exhibit a policy violation rate of 43.1%, with nearly half occurring during intermediate execution steps rather than in final outputs. These findings underscore the critical necessity of process-level compliance auditing to ensure that autonomous coding agents adhere not only to functional requirements but also to established software engineering practices and repository governance norms.
This work proposes a novel approach that leverages advanced computational techniques to address key challenges in the domain. The project employs a carefully designed framework integrating multimodal data processing and adaptive learning mechanisms to enhance model robustness and generalization. By utilizing state-of-the-art architectures and optimization strategies, the methodology achieves significant improvements in performance metrics compared to existing baselines. The experimental evaluation, conducted on diverse benchmark datasets, demonstrates the efficacy and scalability of the proposed solution across various scenarios. Furthermore, ablation studies and qualitative analyses provide insights into the contributions of individual components, validating the design choices. This research not only advances the current understanding of the underlying problem but also offers a practical and extensible platform for future investigations in related areas.
This study addresses a critical gap in software traceability research by systematically investigating the linkage between rejected proposals and their associated source code—a dimension largely overlooked in prior work that predominantly focuses on accepted contributions. To bridge this gap, the authors propose an automated traceability link generation pipeline leveraging large language models (LLMs), grounded in empirical analysis of official Go language proposal discussions and case studies of failed proposals. Evaluated on a Go proposal dataset, the approach achieves a linking granularity selection accuracy of 0.836 and an average precision of 0.643 for generated links. The analysis further uncovers key challenges such as discussion redundancy and ambiguous information, highlighting the untapped value of rejected proposals in understanding software evolution and design rationale.