develop open-source software

Design and build open-source software packages, libraries, modules, and reproducible pipelines with modular APIs and performance considerations (including parallel/GPU optimization) and provide accompanying documentation and examples. Create packaging and distribution artifacts and configuration for languages such as Python and R, publish releases to registries, integrate deployments to cloud or edge environments, and manage release processes and community contributions.

developopen-sourcesoftware

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of constructing reproducible software stacks in high-performance computing (HPC) and AI convergence scenarios, where constraints such as lack of root privileges, network isolation, and heterogeneous language environments hinder conventional tooling. For the first time, it systematically applies Nix’s fully isolated build model together with its declarative flake configuration system to HPC-AI hybrid environments, enabling unified management of C/C++ and Python dependencies. By automatically generating Apptainer containers, the approach ensures consistency between local development and deployment on supercomputing systems. The method effectively resolves critical issues including dependency discovery, system library leakage, and cross-project composition, achieving highly reproducible deployments across non-root workstations and production clusters. It demonstrates clear advantages over traditional module systems, Conda environments, and manual containerization, while also highlighting current gaps in Nixpkgs’ coverage of machine learning packages.

dependency managementenvironment isolationHPC-AI software stack

Addressing the challenges of implementing FAIR principles and open science practices in research-oriented Python software engineering—particularly low automation and poor adoption—the paper introduces the first end-to-end automated framework tailored for scientific computing. Methodologically, it integrates Configuration-as-Code, containerized DevOps, and domain-specific academic software engineering best practices within a cloud-native GitHub ecosystem. The framework enables one-click generation of PEP 517/518-compliant buildable package scaffolds, Sphinx-based documentation, pytest test suites, and GitHub Actions CI/CD pipelines. Leveraging Cookiecutter templates and a RESTful API control center, it ensures zero-friction onboarding for both new and legacy projects. Empirical evaluation demonstrates a 90% reduction in project initialization time, achieves 100% baseline test coverage and documentation completeness, and has been validated across multiple prominent open-source scientific libraries; its template repository is widely adopted by the community.

Automates research software engineering for Python applications.Enhances software quality with FAIR and Open Science principles.Integrates with GitHub for automated development workflows.

A first look at License Variants in the PyPI Ecosystem

Jul 19, 2025
WX
Weiwei Xu
🏛️ Peking University | University of Science and Technology Beijing

Open-source license variants—ranging from minimally modified standard licenses to fully custom terms—are pervasive yet poorly understood across ecosystems like PyPI; existing tools fail to reliably detect them, leading to compliance risks and flawed license analysis. This paper presents the first large-scale empirical study characterizing such variants, revealing widespread textual divergence but rare substantive modifications—many of which nonetheless introduce critical license incompatibilities. To address this, we propose LV-Parser, a lightweight license parser leveraging differential analysis and LLM-assisted validation, achieving 0.936 accuracy with 30% lower computational overhead; and LV-Compat, a dependency-aware compatibility checker that improves detection rate by 5.2× and attains 0.98 precision. Together, they form an end-to-end automated pipeline that significantly enhances license identification accuracy and compliance assessment efficacy.

Analyzing license variants in PyPI for compliance complexitiesDetecting license incompatibilities in software dependency networksDeveloping efficient tools for open-source license variant analysis

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

Scientific software often suffers from poor reproducibility and low sharing rates, primarily because researchers lack formal software engineering training—leading to version chaos, uncontrolled code quality, and cumbersome release processes. To address this, scikit-package introduces a progressive software engineering roadmap tailored for domain scientists. It provides standardized packaging tutorials, automated workflow templates (covering build systems, CI/CD pipelines, documentation generation, and package management), and community-maintained pedagogical resources. Its key innovation lies in adapting professional software engineering practices to the cognitive load of non-specialist programmers, enabling systematic progression from script-based functions to production-ready, open-source package releases. Empirical evaluation demonstrates that scikit-package significantly improves the reproducibility and maintainability of scientific code, enhances community sharing efficiency, and lowers barriers to standardized scientific software publication.

Enhancing reproducibility of scientific software through standardized packagingPromoting reusable and maintainable code at varying complexity levelsSimplifying code sharing for non-expert scientists via tutorials and workflows

Latest Papers

What's happening recently
View more

Addressing challenges in FAIR principle implementation—including fragmented data and code lifecycles, lack of executable environments, and high technical barriers—this study proposes a unified open-science platform. The platform uniquely integrates version control, containerized computational environments, and modular project scaffolding to support end-to-end reproducible research, from grant proposal to publication. It interoperates with mainstream scientific toolchains, supports deployment on both local workstations and institutional servers, and provides a lightweight graphical user interface. Empirical validation demonstrates successful re-execution of over a dozen interdisciplinary studies published more than ten years ago, confirming the platform’s robust long-term reproducibility, cross-platform compatibility, and seamless execution across diverse domains. By significantly lowering technical adoption barriers for researchers, the platform enables practical integration of FAIR principles and reproducibility practices into routine scientific workflows.

Bridges disconnected data and code life cyclesEnables FAIR workflows without manual setupUnifies data and software lifecycles for reproducibility

This work addresses the frequent failures in continuous integration (CI) builds of embedded open-source software, which often stem from cross-compilation complexities, board-specific configurations, and toolchain constraints. These failures are compounded by heterogeneous, ephemeral build logs that are difficult to reuse. To tackle this challenge, the authors propose PhantomRun, a framework that enables standardized reproduction of historical failed builds through a build log abstraction layer, metadata standardization, containerized replay environments, and heterogeneous log parsing techniques. PhantomRun is the first system to support large-scale, controllable replay of failed embedded CI builds, offering a unified, machine-readable interface for build artifacts and metadata. Evaluated on 4,628 failed runs, PhantomRun successfully reconstructed 91.8% of the builds, with 98% preserving the original execution outcomes, demonstrating high reproducibility fidelity.

build failurescontinuous integrationcross-compilation

OpenDORS: A dataset of openly referenced open research software

Dec 01, 2025
SD
Stephan Druskat
🏛️ German Aerospace Center (DLR) | Humboldt Universität zu Berlin

Empirical research on scholarly software lacks large-scale, evidence-based foundations. Method: We constructed the largest literature-linked open-source research software dataset to date, comprising 134,352 distinct projects and 134,154 source code repositories, along with their citations in open-access publications. By systematically integrating metadata from open publishing platforms and code hosting services, we extracted structured information—including version history, licenses, programming languages, and functional descriptions—enabling the first fine-grained mapping between research software and its associated scholarly outputs. Contribution/Results: The publicly released dataset includes complete metadata for over 120,000 projects, substantially addressing the scarcity of high-quality empirical data in research software engineering (RSE). It provides a reproducible foundation for assessing software impact, analyzing development practices, and informing evidence-based policy formulation in scholarly software infrastructure.

Creating a dataset of open research software projects referenced in academic literature.Enabling research on software engineering practices in academic software development.Providing metadata on software repositories for large-scale studies of research software.

This study addresses the critical issue of reproducibility in open-source software builds and its implications for software supply chain security. By systematically analyzing source code discrepancies across 85 versions of 28 widely used Java packages between Maven Central and independently built distributions (e.g., from Google or Oracle), the work reveals—for the first time—that dynamic code generation during the build process is the primary cause of irreproducibility, thereby challenging the prevailing assumption that identical source code guarantees consistent rebuilds. Through comprehensive source code comparison, build process analysis, and root-cause attribution, complemented by an in-depth examination of Maven’s extension mechanisms, the authors propose targeted mitigation strategies that substantially enhance build reproducibility and strengthen trust in the software supply chain.

build reproducibilityMavenpackage rebuild

Hot Scholars

DC

Dianne Cook

Professor of Statistics, Monash University
statistical graphicsdata sciencemultivariate data
SC

Sarah C. Lotspeich

Wake Forest University
BiostatisticsEpidemiologyGlobal HealthPublic Health
MA

Mehdi Astaraki

Karolinska Institutet; Stockholm University
Medical Image AnalysisImaging BiomarkersImage ProcessingML/DL
SM

Samuel Muller

Executive Dean and Professor, Faculty of Science and Engineering, Macquarie University
StatisticsModel SelectionVariable SelectionRobustness
HB

Hamed Babaei Giglou

TIB — Leibniz Information Centre for Science and Technology
NLPLLMsReinforcement LearningOntology Engineering