Score
Designs and implements analyses and tools that extract and interpret commit metadata, diffs, and push-event records from Git repositories to reconstruct commit provenance and lineage, classify change types and patterns, and quantify or visualize maintenance and contribution trends. Builds algorithms and reconciliation procedures to walk and align push-event chains, detect and characterize fast‑forward chain breaks and force‑push orphaned commits, and reconcile discrepancies between advertised and observed repository graphs.
This study presents the first large-scale empirical analysis of Git history rewriting and its threats to software supply chain integrity and reproducibility. Addressing risks—including push conflicts, broken provenance, and malicious code injection—arising from history-altering operations (e.g., rebase, filter-branch) in public repositories, the authors analyze 111 million open-source projects archived by Software Heritage. Leveraging static analysis and two in-depth case studies, they propose the first evidence-driven taxonomy of Git history rewriting and develop GitHistorian, an automated detection tool. Applied at scale, the methodology identifies 1.22 million projects exhibiting history rewriting (8.7 million operations total), revealing prevalent legitimate use cases such as license updates and sensitive information removal. The work establishes a novel, scalable methodology for supply chain security assessment and delivers an open, extensible infrastructure for detecting and characterizing historical tampering.
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
Software source code often harbours"hotspots": small portions of the code that change far more often than the rest of the project and thus concentrate maintenance activity. We mine the complete version histories of 91 evolving, actively developed GitHub repositories and identify 15 recurring line-level hotspot patterns that explain why these hotspots emerge. The three most prevalent patterns are Pinned Version Bump (26%), revealing brittle release practices; Long Line Change (17%), signalling deficient layout; and Formatting Ping-Pong (9%), indicating missing or inconsistent style automation. Surprisingly, automated accounts generate 74% of all hotspot edits, suggesting that bot activity is a dominant but largely avoidable source of noise in change histories. By mapping each pattern to concrete refactoring guidelines and continuous integration checks, our taxonomy equips practitioners with actionable steps to curb hotspots and systematically improve software quality in terms of configurability, stability, and changeability.
Existing tools struggle to support fine-grained analysis of software code evolution effectively. To address this limitation, this work proposes GitEvo—a multilingual, extensible analysis framework that uniquely integrates Git version metadata with syntactic code structures, such as abstract syntax trees (ASTs), enabling deep co-modeling of version history and code structure for the first time. GitEvo facilitates cross-language tracking of code evolution, computation of evolutionary metrics, and interactive visualization. Its effectiveness has been validated on real-world repositories, demonstrating its utility both as a foundation for empirical software engineering research and as an educational platform for understanding the patterns and principles of software evolution.
This study addresses the frequent misattribution of missing commits in open-source code mirrors, which are often ambiguously ascribed to either upstream force-push deletions or incomplete mirror ingestion. For the first time at full scale, this work disentangles these two root causes by integrating GHArchive event streams with the World of Code database, leveraging reference-chain backtracking, cross-source commit graph comparison, and a structured force-push detection algorithm. The resulting provenance dataset spans 1.1 billion commits, each labeled as “ingested,” “rewritten,” or “unobserved.” The analysis reveals that 53.35% of commits are ingested, 6.47% are rewritten, and 40.18% remain unobserved. The study releases a force-push dataset comprising 167 million events and statistics from 78 million repositories, correcting an underestimation of developer productivity by 10.82% and substantially improving the accuracy of mirror completeness assessment and contribution accounting.
Although Git tags are commonly regarded as immutable references, they can in fact be altered or deleted via force pushes, thereby jeopardizing build reproducibility and software supply chain security. This study presents the first large-scale empirical analysis of tag mutability across more than 300 million public repositories, leveraging the Software Heritage dataset to identify 10.2 million tag modification events spanning 189,000 repositories. Through cross-validation with the Nixpkgs package management system, the research confirms that seven packages experienced build failures directly attributable to tag changes. The work systematically exposes the non-immutability of Git tags, quantifies their real-world impact, and offers actionable recommendations for improving software integrity and secure development practices.
This study addresses the challenge that existing code hosting platforms fail to identify cross-platform forking relationships, leading to severe inflation of project popularity metrics due to duplicate counting. To resolve this, the authors propose a de-forking method grounded in globally shared commit relationships, employing star encoding, parallel Louvain clustering, and a cluster-size truncation strategy to systematically reconstruct cross-platform fork families. Their approach enables, for the first time, the identification of root projects outside GitHub, revealing that 5.41% of multi-platform fork families and 1.51% of forks originate from non-GitHub roots. The released de-forking mapping, p2PFull, covers World of Code V2604, achieves 99.01% edge consistency with GitHub’s native fork graph, and includes an exclusion list of 134 million child repositories alongside 455,000 hard-split records.
This study addresses the critical issue of frequent failures in GitHub Actions workflows, which severely undermine automation reliability and maintainability. For the first time, it systematically maps 197 language constructs to 14 workflow capability features through a large-scale quantitative analysis of over 260,000 workflows across 49,000 repositories. By integrating language construct categorization with metadata mining, the work uncovers prevalent usage patterns, evolutionary trends, and their impact on workflow reliability. The findings reveal that only a small subset of constructs is heavily used, and that specific capability features are significantly associated with elevated failure rates and maintenance costs. These empirical insights provide actionable guidance for optimizing workflow design and improving robustness in continuous integration and delivery pipelines.
This study addresses a critical gap in existing research, which has predominantly assessed commit signing adoption from a repository-centric perspective, thereby failing to capture developers’ actual signing behaviors and their real-world implications for software supply chain security. For the first time, we employ a platform-scale, developer-centered analytical framework, tracking 71,694 active GitHub users and their 16 million commits. By integrating large-scale behavioral logs, metadata analysis, audits of PGP/Git verification mechanisms, and key lifecycle assessments, we uncover systemic deficiencies in current signing practices: after excluding platform-generated signatures, fewer than 6% of developers have ever manually signed a commit; most manual signatures originate from the web interface; approximately 12.5% of locally signed commits are unverifiable due to missing public key uploads; and over 25% of users retain expired yet unrevoked keys.