commit history analysis

Designs and implements analyses and tools that extract and interpret commit metadata, diffs, and push-event records from Git repositories to reconstruct commit provenance and lineage, classify change types and patterns, and quantify or visualize maintenance and contribution trends. Builds algorithms and reconciliation procedures to walk and align push-event chains, detect and characterize fast‑forward chain breaks and force‑push orphaned commits, and reconcile discrepancies between advertised and observed repository graphs.

commithistoryanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Altered Histories in Version Control System Repositories: Evidence from the Trenches

Sep 11, 2025
SR
Solal Rapaport
🏛️ LTCI | Telecom Paris | Institut Polytechnique de Paris

This study presents the first large-scale empirical analysis of Git history rewriting and its threats to software supply chain integrity and reproducibility. Addressing risks—including push conflicts, broken provenance, and malicious code injection—arising from history-altering operations (e.g., rebase, filter-branch) in public repositories, the authors analyze 111 million open-source projects archived by Software Heritage. Leveraging static analysis and two in-depth case studies, they propose the first evidence-driven taxonomy of Git history rewriting and develop GitHistorian, an automated detection tool. Applied at scale, the methodology identifies 1.22 million projects exhibiting history rewriting (8.7 million operations total), revealing prevalent legitimate use cases such as license updates and sensitive information removal. The work establishes a novel, scalable methodology for supply chain security assessment and delivers an open, extensible infrastructure for detecting and characterizing historical tampering.

Analyzing impacts on repository integrity and reproducibilityIdentifying security risks from rewritten commit historiesInvestigating Git history alterations in public repositories

This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.

CI/CDfailure responseGitHub Actions

Software source code often harbours"hotspots": small portions of the code that change far more often than the rest of the project and thus concentrate maintenance activity. We mine the complete version histories of 91 evolving, actively developed GitHub repositories and identify 15 recurring line-level hotspot patterns that explain why these hotspots emerge. The three most prevalent patterns are Pinned Version Bump (26%), revealing brittle release practices; Long Line Change (17%), signalling deficient layout; and Formatting Ping-Pong (9%), indicating missing or inconsistent style automation. Surprisingly, automated accounts generate 74% of all hotspot edits, suggesting that bot activity is a dominant but largely avoidable source of noise in change histories. By mapping each pattern to concrete refactoring guidelines and continuous integration checks, our taxonomy equips practitioners with actionable steps to curb hotspots and systematically improve software quality in terms of configurability, stability, and changeability.

change historycode churnmaintenance activity

Existing tools struggle to support fine-grained analysis of software code evolution effectively. To address this limitation, this work proposes GitEvo—a multilingual, extensible analysis framework that uniquely integrates Git version metadata with syntactic code structures, such as abstract syntax trees (ASTs), enabling deep co-modeling of version history and code structure for the first time. GitEvo facilitates cross-language tracking of code evolution, computation of evolutionary metrics, and interactive visualization. Its effectiveness has been validated on real-world repositories, demonstrating its utility both as a foundation for empirical software engineering research and as an educational platform for understanding the patterns and principles of software evolution.

code evolutiondevelopment toolsempirical analysis

This study addresses the frequent misattribution of missing commits in open-source code mirrors, which are often ambiguously ascribed to either upstream force-push deletions or incomplete mirror ingestion. For the first time at full scale, this work disentangles these two root causes by integrating GHArchive event streams with the World of Code database, leveraging reference-chain backtracking, cross-source commit graph comparison, and a structured force-push detection algorithm. The resulting provenance dataset spans 1.1 billion commits, each labeled as “ingested,” “rewritten,” or “unobserved.” The analysis reveals that 53.35% of commits are ingested, 6.47% are rewritten, and 40.18% remain unobserved. The study releases a force-push dataset comprising 167 million events and statistics from 78 million repositories, correcting an underestimation of developer productivity by 10.82% and substantially improving the accuracy of mirror completeness assessment and contribution accounting.

collection gapcommit provenanceforce-push

Latest Papers

What's happening recently
View more

Although Git tags are commonly regarded as immutable references, they can in fact be altered or deleted via force pushes, thereby jeopardizing build reproducibility and software supply chain security. This study presents the first large-scale empirical analysis of tag mutability across more than 300 million public repositories, leveraging the Software Heritage dataset to identify 10.2 million tag modification events spanning 189,000 repositories. Through cross-validation with the Nixpkgs package management system, the research confirms that seven packages experienced build failures directly attributable to tag changes. The work systematically exposes the non-immutability of Git tags, quantifies their real-world impact, and offers actionable recommendations for improving software integrity and secure development practices.

Git tagsimmutabilityreproducible builds

This study addresses the challenge that existing code hosting platforms fail to identify cross-platform forking relationships, leading to severe inflation of project popularity metrics due to duplicate counting. To resolve this, the authors propose a de-forking method grounded in globally shared commit relationships, employing star encoding, parallel Louvain clustering, and a cluster-size truncation strategy to systematically reconstruct cross-platform fork families. Their approach enables, for the first time, the identification of root projects outside GitHub, revealing that 5.41% of multi-platform fork families and 1.51% of forks originate from non-GitHub roots. The released de-forking mapping, p2PFull, covers World of Code V2604, achieves 99.01% edge consistency with GitHub’s native fork graph, and includes an exclusion list of 134 million child repositories alongside 455,000 hard-split records.

code repositoriescross-forgefork detection

This study addresses the critical issue of frequent failures in GitHub Actions workflows, which severely undermine automation reliability and maintainability. For the first time, it systematically maps 197 language constructs to 14 workflow capability features through a large-scale quantitative analysis of over 260,000 workflows across 49,000 repositories. By integrating language construct categorization with metadata mining, the work uncovers prevalent usage patterns, evolutionary trends, and their impact on workflow reliability. The findings reveal that only a small subset of constructs is heavily used, and that specific capability features are significantly associated with elevated failure rates and maintenance costs. These empirical insights provide actionable guidance for optimizing workflow design and improving robustness in continuous integration and delivery pipelines.

execution failuresGitHub Actionslanguage constructs

This study addresses a critical gap in existing research, which has predominantly assessed commit signing adoption from a repository-centric perspective, thereby failing to capture developers’ actual signing behaviors and their real-world implications for software supply chain security. For the first time, we employ a platform-scale, developer-centered analytical framework, tracking 71,694 active GitHub users and their 16 million commits. By integrating large-scale behavioral logs, metadata analysis, audits of PGP/Git verification mechanisms, and key lifecycle assessments, we uncover systemic deficiencies in current signing practices: after excluding platform-generated signatures, fewer than 6% of developers have ever manually signed a commit; most manual signatures originate from the web interface; approximately 12.5% of locally signed commits are unverifiable due to missing public key uploads; and over 25% of users retain expired yet unrevoked keys.

commit signingdeveloper behaviorkey management

Hot Scholars

SZ

Stefano Zacchiroli

LTCI, Télécom Paris, Polytechnique Institute of Paris, France
software engineeringopen source softwaredigital commonscomputer security
AH

Andre Hora

Universidade Federal de Minas Gerais (UFMG)
Software EvolutionSoftware MaintenanceSoftware TestingMining Software Repositories
CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
AM

Audris Mockus

University of Tennessee
Digital ArchaeologySoftware EngineeringVisualizationOptimization
AE

Ahmed E. Hassan

Mustafa Prize Laureate, ACM/IEEE/NSERC Steacie Fellow, ACM Influential/IEEE Distinguished Educator
Mining Software RepositoriesSoftware AnalyticsEmpirical Software EngineeringSoftware