Score
Designs and builds maps and network visualizations that represent ecosystem actors, system components, data and decision flows, and the boundaries and overlaps across multiple projects or repositories. Analyzes contributor relationships and cross‑ecosystem contribution patterns — including multi-repo contributor networks, central hub projects, and shared infrastructure dependencies — to quantify cross‑ecosystem contributors and locate gaps, overlaps, and assessment or operational boundaries.
This study addresses the challenges of structural understanding and risk governance in large-scale software ecosystems. Focusing on Maven Central, it constructs a Java dependency network comprising 1.3 million nodes and 20.9 million edges. Methodologically, it introduces a novel BFS-based sampling strategy centered on highly connected nodes, leverages the Goblin framework for dependency extraction, and applies topological metrics—including degree centrality, betweenness centrality, PageRank, and connected components—for structural analysis. Results confirm the network exhibits scale-free and small-world properties. Crucially, testing frameworks and general-purpose utility libraries emerge as sparse but high-impact hubs: they enable efficient code reuse while simultaneously serving as primary conduits for systemic security risk propagation. The study thus provides empirical evidence and methodological support for enhancing software ecosystem resilience, optimizing dependency governance, and identifying critical infrastructure components.
This study addresses the lack of systematic, dependency-aware approaches for accurately assessing the ecosystem-wide impact of maintenance activities in open-source software. To bridge this gap, we propose the first impact metric model grounded in dependency propagation, integrating structural centrality with maintenance dynamics. By analyzing 718,750 packages and over two million dependency relationships from PyPI, our method quantifies influence propagation and identifies high-impact packages. Empirical results reveal that merely 0.1% of packages account for approximately 80% of the total ecosystem influence. Furthermore, we demonstrate a significant misalignment between current support mechanisms—such as Tidelift and GitHub Sponsors—and actual package impact, underscoring the effectiveness of our approach in enabling more targeted and scalable resource allocation within open-source ecosystems.
This study quantifies the impact of non-code contributions—such as network position, temporal activity patterns, and code review behavior—on contributor influence within open-source ecosystems. Leveraging 25 years of project data from the Cloud Native Computing Foundation, the work presents the first systematic integration of graph neural networks, temporal network analysis, and multidimensional contribution metrics. The authors develop GPU-accelerated implementations of PageRank and betweenness centrality, alongside a custom LSTM model, to identify five distinct contributor roles. Findings reveal that the top 1% of contributors disproportionately shape structural influence, with “Bridge”-type roles proving critical to network connectivity. The research further demonstrates that network metrics evolve significantly around project milestones and proposes a role-based framework for assessing community health.
Foundational open-source software libraries—such as NumPy and R base—that underpin biomedical research tools are frequently overlooked, with their contributions inadequately quantified and unrecognized in research policy. Method: We construct a cross-ecosystem (PyPI, CRAN, Bioconductor) software dependency graph, leveraging the CZ Software Mentions Dataset to identify empirically observed dependencies in real-world scientific workflows. We propose the first centrality assessment framework tailored to software ecosystems, integrating graph-theoretic metrics—including betweenness and closeness centrality—to systematically quantify the structural importance of foundational libraries. Contribution/Results: Our analysis uncovers several “hidden hero” libraries whose critical role is substantiated by reproducible, data-driven evidence. This work provides an empirical foundation for funding agencies to prioritize investment in essential software infrastructure, thereby strengthening the sustainability and reliability of the biomedical research ecosystem.
This study investigates the mechanisms sustaining cross-community collaboration in open-source software ecosystems and their impact on community survival. Drawing on contribution data from 464 cybersecurity projects and 11,372 contributors between 2001 and 2022, the authors construct a contributor–repository bipartite graph and introduce a “recognition/repetition” relational framework. Empirical analysis combines Louvain community detection, survival analysis, and a residualized hazard model. Findings reveal that cross-community collaboration is highly concentrated within an extremely thin “carrier layer,” with boundary friction diminishing significantly as relational depth increases. The top 50 cross-community contributors account for 54% of such pull requests, whose acceptance rate rises from 42% to 87% and median response time drops from 147 to 49 hours. Despite the stability of this carrier layer, communities formed later face elevated extinction risks, exhibiting pronounced cohort-structured survival patterns.
This study addresses the longstanding lack of a systematic review on breaking changes in software ecosystems, which has led to fragmented understanding. Through a systematic literature review of 97 studies across five major ecosystems, the work proposes a four-dimensional taxonomy and constructs a multidimensional classification framework. It identifies maintenance and design improvements as the primary drivers of breaking changes and exposes trust failures in semantic versioning practices. Integrating qualitative and quantitative approaches, the research encompasses syntactic and behavioral change detection, dependency propagation, and ecosystem governance, synthesizing 43 detection methods and 66 mitigation strategies. While syntactic change detection demonstrates high accuracy, coverage of behavioral changes remains insufficient. The study culminates in actionable practice guidelines and highlights three key research opportunities and challenges, including leveraging large language models for behavioral contract inference.
This study investigates how the core concepts and influence of discontinued projects persist and evolve within the quantum software ecosystem, and how developer contributions migrate across repositories. Employing empirical software engineering methods—including repository activity analysis, contributor network tracing, and cross-project linkage mining—it systematically uncovers, for the first time, the mechanisms underlying knowledge transfer and developer mobility in this emerging domain. The findings reveal that terminated projects can sustain their impact through strategies such as fostering enduring communities, integrating into mature ecosystems, or spawning new artifacts, thereby demonstrating clear evidence of cross-repository conceptual inheritance. These results provide crucial empirical insights into the evolutionary dynamics of software ecosystems in nascent technical fields.