mining software repositories

Designs and implements scalable pipelines and analyses that collect, extract, clean, link, and aggregate artifacts from software hosting and version-control systems (commits, branches, tags, code, issues, pull requests, and metadata) to produce reproducible datasets and metrics. This work includes building API clients and scrapers, selecting and filtering candidate repositories, linking issues to code and tooling changes, and computing measures such as project lifespan, developer activity, issue frequencies, adoption trends, and cross-repository architecture patterns.

miningsoftwarerepositories

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$212K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories

Jan 25, 2025
NH
Nicole Hoess
🏛️ Technical University of Applied Sciences Regensburg | University of Hawaii at Mānoa | Siemens AG

Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.

Data Analysis VariabilityResearch ReliabilitySoftware Engineering

Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.

community-drivenheterogeneous execution environmentsmaintenance

Oops!... I did it again. Conclusion (In-)Stability in Quantitative Empirical Software Engineering: A Large-Scale Analysis

Oct 08, 2025
NH
Nicole Hoess
🏛️ Technical University of Applied Sciences Regensburg | University of Hawaii at Mānoa

This paper investigates validity threats arising from toolchain selection in quantitative empirical software engineering. We formally replicate three high-impact studies by extracting identical project data using four widely adopted mining tools—Git, JIRA, GitHub API, and BIC—and conduct both quantitative and qualitative comparative analyses. Results demonstrate that subtle technical discrepancies across tools—including data modeling assumptions, event definitions, and temporal window handling—propagate and significantly undermine consistency in baseline datasets, statistical outcomes, and ultimately research conclusions. To our knowledge, this is the first systematic study to reveal the critical impact of tool choice on the robustness of empirical findings. We propose a practical framework comprising enhanced tool reusability, improved analytical transparency, and mandatory cross-tool validation. This work advances methodological rigor in software evolution research by highlighting and mitigating tool-induced validity threats.

Analyzes how technical differences affect empirical study outcomesEvaluates tool agreement on data and research conclusionsInvestigates validity threats in software mining tool pipelines

EVOSCAT: Exploring Software Change Dynamics in Large-Scale Historical Datasets

Aug 14, 2025
SS
Souhaila Serbout
🏛️ University of Zurich | Software Institute | USI Lugano

Addressing the challenge of exploring and comparing component evolution dynamics—such as change frequency, aging rate, and quality trends—in large-scale software history data, this paper proposes a scalable, interactive density scatterplot visualization method. The method innovatively integrates flexible temporal axis scaling and alignment, automatic component ranking, and semantics-driven color mapping, enabling efficient analysis of millions of evolution events and customizable, multi-dimensional insights. Leveraging software repository mining and visualization encoding techniques, the system is empirically validated on cross-repository datasets—including OpenAPI specifications and GitHub Actions workflows—demonstrating significant improvements in global identification and comparative analysis of long-term evolutionary patterns across thousands of components. The approach provides a reusable, interactive analytical framework for software evolution research.

Comparing artifact change dynamics over multi-year spansEnabling flexible analysis of historical software metrics trendsVisualizing large-scale software evolution datasets efficiently

This work addresses the challenge of maintaining up-to-date architectural documentation in microservice systems, which is exacerbated by polyglot implementations, multiple repositories, and rapid independent evolution. Existing static refactoring approaches are often limited to single-repository settings or homogeneous technology stacks. To overcome these limitations, we propose a distributed static architecture reconstruction framework that supports multi-language and multi-repository environments. The framework employs pluggable extractor modules for language-specific analysis and introduces mechanisms for cross-repository data propagation and fusion, enabling seamless interoperability with existing static analysis tools. To the best of our knowledge, this is the first framework to enable distributed, collaborative architecture reconstruction, significantly enhancing the scalability and usability of automated documentation generation and maintenance in complex microservice ecosystems.

architecture reconstructionmicroservicemulti-repository

Latest Papers

What's happening recently
View more

This work addresses the critical impact of unresolved issue artifacts—such as bugs and missing documentation—on software quality in open-source projects. To overcome the limitations of existing tools, which lack systematic analysis of issue lifecycles and evolutionary patterns, we propose G-Issue, the first mining tool that integrates issue lifecycle modeling with evolution tracking. Built on a Python API, G-Issue efficiently collects and analyzes issue data from open-source repositories, achieving faster mining performance than mainstream tools while uncovering key patterns in issue evolution. The tool further enables issue prioritization and developer assignment based on evolutionary characteristics, offering both a novel methodology and empirical evidence to support quality management in open-source software development.

issue evolutionissue lifetimeissue-related artifacts

This study addresses the lack of systematic understanding regarding the application domains, maintenance characteristics, and effective design practices of GitHub template repositories. Conducting the first large-scale empirical investigation, the work integrates data mining, statistical analysis, code quality assessment tools—detecting code smells, vulnerabilities, and security hotspots—and an LLM-as-a-judge classification approach to systematically uncover domain distributions, language-specific quality variations, and maintenance patterns. The findings reveal web development as the dominant application domain, with high-quality templates consistently adhering to software engineering best practices and offering comprehensive documentation. Through qualitative evaluation, the study distills actionable design guidelines and identifies common pitfalls, providing practical guidance for developers creating or using template repositories.

empirical studyGitHub template repositoriesmaintenance

This study addresses the lack of systematic understanding regarding the content and long-term evolution of GitHub repositories. It presents the first large-scale empirical analysis of 10,000 real-world open-source repositories, combining static content parsing with time-series modeling to trace the evolution of files, directories, and file extensions over the past decade. The findings reveal that README.md, .gitignore, and LICENSE have become standard components; CI/CD tooling has shifted from diversity toward dominance by GitHub Actions; configuration formats exhibit a clear rise of TOML, YAML, and JSON alongside the decline of XML; and Dockerfiles as well as LLM-related files (e.g., AGENTS.md) have grown significantly. This work provides quantitative evidence for understanding technological shifts and standardization processes in the open-source ecosystem.

empirical studyevolution of open sourceGitHub repositories

Hot Scholars

CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
AE

Ahmed E. Hassan

Mustafa Prize Laureate, ACM/IEEE/NSERC Steacie Fellow, ACM Influential/IEEE Distinguished Educator
Mining Software RepositoriesSoftware AnalyticsEmpirical Software EngineeringSoftware
BA

Bram Adams

Queen's University
software release engineeringsoftware integrationsoftware build systemssoftware modularity
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
YK

Yutaro Kashiwa

Associate Professor@Nara Institute of Science and Technology (NAIST)
Software engineeringMining Software Repository