scientific software engineering

Designs, implements, tests, and maintains software that enables computational scientific work, including modular, testable codebases, numerical solvers, experiment scripts, and multi-stage data or conversion pipelines. Builds reproducible releases and automation for experiments by profiling and optimizing performance and memory, managing numerical precision and stability, tuning algorithm hyperparameters, and packaging/tests to support reliable scientific computing.

scientificsoftwareengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$237K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research Scientists

Jan 17, 2025
AC
Alyssia Chen
🏛️ University of Hawai'i at Manoa | University of Nebraska - Lincoln

Poor code readability in scientific software severely hinders cross-team collaboration and research reproducibility—particularly among self-taught researchers, who typically lack formal training in readability best practices, resulting in opaque naming conventions and inadequate documentation. This study employs a mixed-methods approach—including surveys, in-depth interviews, and statistical analysis—across 57 interdisciplinary researchers to empirically investigate current practices. It reveals, for the first time, that in the absence of structured training, researchers heavily rely on informal, ad hoc commenting practices; further, it identifies large language models (LLMs) as an emerging paradigm for enhancing code quality. Results show that 57.9% of participants received no readability-specific instruction, with inconsistent naming and missing documentation identified as the two primary bottlenecks. Based on these findings, we propose a lightweight, human-centered code quality support framework tailored for scientific programmers—addressing a critical gap in the human factors literature on scientific code readability.

Code DocumentationScientific Software ReadabilityVariable Naming

Scientific software often suffers from poor reproducibility and low sharing rates, primarily because researchers lack formal software engineering training—leading to version chaos, uncontrolled code quality, and cumbersome release processes. To address this, scikit-package introduces a progressive software engineering roadmap tailored for domain scientists. It provides standardized packaging tutorials, automated workflow templates (covering build systems, CI/CD pipelines, documentation generation, and package management), and community-maintained pedagogical resources. Its key innovation lies in adapting professional software engineering practices to the cognitive load of non-specialist programmers, enabling systematic progression from script-based functions to production-ready, open-source package releases. Empirical evaluation demonstrates that scikit-package significantly improves the reproducibility and maintainability of scientific code, enhances community sharing efficiency, and lowers barriers to standardized scientific software publication.

Enhancing reproducibility of scientific software through standardized packagingPromoting reusable and maintainable code at varying complexity levelsSimplifying code sharing for non-expert scientists via tutorials and workflows

Computational Reproducibility of R Code Supplements on OSF

May 27, 2025
LS
Lorraine Saju
🏛️ GESIS | Leibniz Institute for the Social Sciences

This study addresses the widespread lack of computational reproducibility in R supplementary code deposited on the Open Science Framework (OSF). A systematic audit of 296 published R code packages revealed that 98.8% incompletely declare dependencies. To address this, we propose the first automated reproducibility auditing framework tailored to the R ecosystem. It combines static source-code analysis—leveraging regular expressions and abstract syntax trees (ASTs)—to accurately infer dependencies, with Docker-based containerized execution and failure diagnostics (e.g., path errors, OS-specific inconsistencies, missing packages) to enable end-to-end environment reconstruction and validation. Experiments successfully executed 25.87% of scripts, identifying undeclared dependencies, hardcoded file paths, and cross-platform compatibility issues as the three primary barriers to reproducibility. The framework enables large-scale, low-cost, and scalable quantitative assessment of computational reproducibility in scholarly research, providing a practical toolchain to enhance transparency and verifiability.

Assessing computational reproducibility of R projectsDeveloping automated pipeline for environment reconstructionIdentifying barriers like undeclared dependencies and file paths

Scientific software selection frequently suffers from non-reproducible benchmarks due to multi-library, multi-metric evaluation and dynamic evolution—such as the introduction of new algorithms or modifications to test cases and evaluation criteria. This paper addresses numerical integration over arbitrary 2D/3D domains with implicit or parameterized boundaries (cut-cell quadrature), proposing the first automated benchmarking framework that systematically integrates CI/CD engineering practices into scientific computing workflows. The framework unifies GitHub Actions, Docker, Python-based scheduling, Jupyter-based report generation, and semantically versioned result archiving. It supports automated configuration, execution, visualization, and historical result comparison. It achieves >90% automation for benchmark tasks and regression detection; reduces integration time for new libraries or algorithms by 70%; and enables precise attribution of performance deviations to specific code commits. The framework significantly enhances reliability, reproducibility, and evolutionary adaptability in scientific software evaluation.

Automating benchmarking of diverse scientific software alternativesManaging expanding parameter spaces in benchmark setupsStreamlining re-evaluation when adding new metrics or cases

Ten simple rules for training scientists to make better software

Feb 07, 2024
KG
K. Gallagher
🏛️ University of Oxford | University of Macau | University of Nottingham

Doctoral students in life sciences commonly lack formal software engineering training, hindering the development of robust, reproducible, and collaborative research software. Method: This study proposes ten pedagogical principles for research software development, establishing the first systematic framework centered on “research software pedagogy”—distinct from generic programming instruction. It integrates software engineering best practices (e.g., Git-based version control, CI/CD pipelines, unit testing, RESTful API design), learning science principles, and authentic research workflows, emphasizing the seamless embedding of automation, documentation, testing, and collaborative practices throughout the research lifecycle. Contribution/Results: The framework delivers a generalizable, plug-and-play pedagogical paradigm. Deployed across multiple Chinese universities’ life sciences PhD programs, it has demonstrably improved software deliverable quality, code reusability, and cross-team collaboration efficiency—bridging critical gaps between computational literacy and rigorous, team-based scientific software practice.

Addressing the lack of formal software development training in research.Enhancing reproducibility and good practices in computational research.Teaching scientists to develop high-quality, sustainable software.

Latest Papers

What's happening recently
View more

Scientific computing notebooks frequently suffer from irreproducibility, poor readability, and limited reusability, posing serious threats to research reliability. This work presents the first large-scale empirical study of 1,510 Jupyter notebooks from 518 code repositories published in Nature in 2024. Through manual reproduction attempts (only 2 successful out of 19), documentation review, code clone detection (≥10 lines, ≥3 instances), and mutation analysis, the study systematically uncovers pervasive issues including chaotic state management, missing dependencies, and excessive code duplication. To address these challenges, the authors propose the first multidimensional quality assessment framework explicitly designed to evaluate reproducibility, readability, and reusability, thereby establishing an empirical foundation and methodological support for improving the quality of scientific code.

code duplicationcode qualityreproducibility

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

Towards Experiment Execution in Support of Community Benchmark Workflows for HPC

Jul 29, 2025
GV
Gregor von Laszewski
🏛️ University of Virginia | Oak Ridge National Laboratory | Hewlett Packard Enterprise Canada | University of Florida | Cummins | San Diego Supercomputer Center | University of California, San Diego

To address low reusability of HPC benchmarks, poor cross-platform portability, and inefficient resource validation, this paper proposes the “benchmark carpentry” paradigm—a lightweight, reusable experimental execution framework. Methodologically, it integrates Cloudmesh’s experiment executor with HPE SmartSim, incorporating standardized workflow templates, AI/ML–simulation coupling mechanisms, and a unified experimental management interface. Its key contribution is the first application of craftsmanship principles to benchmarking process design, enabling automated, cross-domain and cross-architecture benchmark deployment and capability assessment. Evaluated on representative scientific computing workloads—including cloud masking analysis, seismic forecasting, and CFD surrogate modeling—the framework achieves ≥92% workflow reproducibility and reduces average deployment time by 68%, significantly improving resource configuration efficiency. It establishes a scalable, community-driven paradigm for HPC capability validation.

Creating adaptable workflow templates for scientific applicationsDemonstrating HPC compute capability with limited benchmarksImproving experiment management tools for broader workflow adaptability

This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.

computational workflowsportablereproducible

Hot Scholars

JM

Jie M. Zhang

Lecturer (Assistant Professor), King's College London
LLMsSE4MLmachine learning testingmutation testing
AE

Ahmed E. Hassan

Mustafa Prize Laureate, ACM/IEEE/NSERC Steacie Fellow, ACM Influential/IEEE Distinguished Educator
Mining Software RepositoriesSoftware AnalyticsEmpirical Software EngineeringSoftware
MK

Marcos Kalinowski

Professor, Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
Empirical Software EngineeringAI EngineeringAI4SEHuman Aspects in Software Engineering
PA

Paris Avgeriou

Full Professor of Software Engineering, University of Groningen
Software EngineeringSoftware ArchitectureEmpirical Software EngineeringSoftware Maintenance and Evolution
TG

Tom Goldstein

Volpi-Cupal Professor of Computer Science, University of Maryland
Numerical OptimizationMachine LearningDistributed ComputingComputer Vision