manage reproducible experiments

Design, build, and operate configuration-driven experiment management systems and reproducible evaluation pipelines that make computational experiments repeatable, auditable, and rerunnable. This includes authoring registry-based YAML configs and reusable configuration components, pre-registering hypotheses and decision rules, enforcing deterministic run controls (seeds, temperature), versioning code and datasets, cataloging and archiving artifacts, logging configurations and metrics, caching outputs, and providing reproducible harnesses and workflows for evaluation and reruns.

managereproducibleexperiments

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Virtual Laboratory for Managing Computational Experiments

Apr 01, 2025
EA
Eleni Adamidi
🏛️ Athena Research Center

Computational experiments face declining reproducibility and metadata management challenges as scale increases. To address this, we propose a metadata-driven, full-lifecycle management approach for computational experiments and implement it in SCHEMA Lab—a virtual laboratory. Our method introduces an ontology-based metadata model explicitly designed for experimental lifecycles, enabling structured capture of configurations, execution logs, and performance metrics. We further incorporate an experiment lineage graph and semantic grouping mechanisms to support cross-instance traceability and multi-experiment relational analysis. Architecturally, SCHEMA Lab adopts a web-based microservices design, RESTful APIs, and visual workflow orchestration. Empirical evaluation demonstrates a 92% experiment reproduction success rate and over 60% reduction in configuration and audit time. Deployed across multiple HPC and AI research teams, the system significantly enhances scientific reproducibility and collaborative efficiency.

Capturing detailed metadata for experiment configurationsManaging reproducibility in large-scale computational experimentsOrchestrating complex sequences of computational tasks

Addressing challenges in FAIR principle implementation—including fragmented data and code lifecycles, lack of executable environments, and high technical barriers—this study proposes a unified open-science platform. The platform uniquely integrates version control, containerized computational environments, and modular project scaffolding to support end-to-end reproducible research, from grant proposal to publication. It interoperates with mainstream scientific toolchains, supports deployment on both local workstations and institutional servers, and provides a lightweight graphical user interface. Empirical validation demonstrates successful re-execution of over a dozen interdisciplinary studies published more than ten years ago, confirming the platform’s robust long-term reproducibility, cross-platform compatibility, and seamless execution across diverse domains. By significantly lowering technical adoption barriers for researchers, the platform enables practical integration of FAIR principles and reproducibility practices into routine scientific workflows.

Bridges disconnected data and code life cyclesEnables FAIR workflows without manual setupUnifies data and software lifecycles for reproducibility

Let's Talk About It: Making Scientific Computational Reproducibility Easy

Apr 14, 2025
LC
L'azaro Costa
🏛️ University of Porto | INESC TEC

Reproducing scientific computing results remains challenging due to missing experimental documentation, configuration details, and datasets—undermining result credibility. To address this, we propose the first natural language interaction paradigm explicitly designed for computational reproducibility: users describe experiments in plain English; the system automatically parses intent, infers dependencies, constructs a containerized execution environment, and packages code, data, and runtime into a single, cross-platform executable file enabling zero-configuration, one-click re-execution. Our approach integrates natural language understanding (NLU), automated environment provisioning, and lightweight container encapsulation. Evaluated on 18 published computational experiments, it achieves a high rate of fully automated, intervention-free reproduction. A user study demonstrates significantly higher System Usability Scale (SUS) scores compared to leading commercial tools and substantially lower NASA-TLX cognitive workload—validating simultaneous advances in usability and practical utility.

Addressing challenges in computational reproducibility of scientific experimentsImproving usability and reducing workload in computational reproducibility toolsProposing a tool to simplify experiment replication with natural language interaction

Addressing the poor reproducibility and high environmental heterogeneity of interdisciplinary computational experiments, this paper proposes SciRep—a framework that unifies management of code, data, programming languages, dependencies, and execution commands via containerized encapsulation, declarative experiment specifications (YAML/JSON), dependency snapshotting, and deterministic scheduling. It produces lightweight, portable “capsule packages.” SciRep introduces the first domain-agnostic reproducibility packaging paradigm, embodying “configuration-as-documentation” and “execution-as-verification,” and achieves end-to-end reproducible workflows across heterogeneous domains—including medicine, bioinformatics, and computer science—for the first time. It supports multi-language ecosystems (e.g., Python, R, Julia) on Linux/macOS. Empirical evaluation successfully reproduced 16 of 18 published experiments (89%), significantly surpassing the best prior tool’s 61% reproducibility rate; all successfully executed experiments reproduced original results with 100% fidelity.

Addressing reproducibility challenges in computational experiments across scientific domains.Enhancing replicability by ensuring consistent results across different computational environments.Proposing a framework to standardize and automate experiment configuration and execution.

Latest Papers

What's happening recently
View more

This work addresses the widespread irreproducibility of academic Jupyter Notebooks caused by environment drift, missing dependencies, and implicit execution assumptions. The authors propose the first web-oriented, automated reproducibility engineering pipeline that systematically reconstructs and evaluates repository-level execution environments for notebooks hosted on GitHub. By leveraging dependency inference, auto-generated Docker containers, and isolated execution, the pipeline enables large-scale assessment of reproducibility. A novel four-category execution outcome framework is introduced to quantify reproduction fidelity. Evaluation on 443 real-world notebooks shows that containerization resolves 66.7% of dependency-related failures; however, only 46.3% achieve high output fidelity, demonstrating that while containerization is necessary, it is insufficient for bit-for-bit reproducibility. These findings underscore the critical need for systematic reproducibility evaluation in computational research.

computational reproducibilitydependency managementenvironment drift

This work addresses the limitations of existing reproducibility assessment methods, which rely on manual annotations and thus lack scalability and authentic supervision signals reflecting real-world reproduction challenges. The authors propose the first scalable evaluation framework that leverages GitHub user-submitted issues as natural supervision, enabling large-scale assessment of large language model (LLM) agents’ ability to identify paper-to-code reproducibility issues without human annotation. By integrating language understanding with code context analysis, the approach enables non-execution-based detection of reproducibility barriers. Experimental results demonstrate that the best-performing LLM agent identifies at least one semantically relevant reproducibility issue—aligned with those reported by humans—in approximately 90% of the evaluated papers, exhibiting strong performance in both failure detection and semantic localization.

benchmarkingGitHub issuesLLM agents

This study addresses the challenge of limited reproducibility and transparency in software engineering controlled experiments, often stemming from inadequate documentation. While generic preregistration templates—such as those provided by the Open Science Framework (OSF)—exist, they fail to comprehensively address the specific needs of software engineering research. This work presents the first systematic evaluation of the OSF preregistration template’s applicability to software engineering experiments, combining literature analysis, template comparison, and cross-referencing against established software engineering experimental reporting guidelines. The findings reveal that although existing OSF templates partially satisfy methodological requirements, none fully encompass all critical elements, and their customization capabilities are constrained. Based on these insights, the paper advocates for and provides a foundation toward developing a domain-specific, standardized preregistration template tailored to software engineering, thereby filling a critical gap in the field.

controlled experimentsempirical software engineeringregistered reports

This work proposes an AI agent–driven workflow to address the high costs of reproducing large-scale empirical studies, which often stem from discrepancies in computational environments, code, and documentation. The approach decouples scientific reasoning from computational execution: researchers supply standardized diagnostic templates, and the system automatically retrieves and orchestrates reproduction materials within a version-controlled environment. A structured knowledge layer captures failure patterns, enabling adaptive reproduction across heterogeneous studies while ensuring transparency and stability of the analytical pipeline. Evaluated on 92 instrumental variable studies, the method achieves an 87% end-to-end reproduction success rate; when data and code are available, it attains 100% success at both the paper and model levels.

empirical dataexecution bottlenecklarge-scale reanalysis

This work addresses the inefficiency in notebook-based distributed workflows, where minor modifications often trigger full re-execution, severely hindering iterative development and reproducibility. To overcome this limitation, the authors propose NBRewind, a system that, for the first time, enables fine-grained incremental execution and cross-platform portability while preserving reproducibility. NBRewind integrates a dual-kernel architecture—comprising auditing and replay components—with cell-level incremental checkpoints and inter-cell dataflow analysis. It further leverages standardized notebook packaging to facilitate efficient partial re-execution. Evaluation in real-world high-performance computing (HPC) scenarios demonstrates that NBRewind incurs minimal overhead for incremental checkpointing and substantially improves both execution efficiency and cross-site reproducibility.

checkpointingdistributed workflowsiterative development

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
JS

Jimeng Sun

Professor at University of Illinois Urbana-Champaign
AI for healthcareMachine learning for healthcaredeep learning for healthcare
LP

Lap-Pui Chau

The Hong Kong Polytechnic University
Visual Signal Processing
HD

Haodong Duan

Shanghai AI Lab | CUHK | PKU
Computer VisionVideo UnderstandingMultimodal LearningGenerative AI
XC

Xuezhi Cao

Meituan
Data MiningKnowledge GraphLLMs