Score
Designs and implements workflows, tooling, and protocols that enable independent replication and validation of experimental results — including automated replication pipelines, reproducible packaging and reporting, documentation, and statistical reproducibility tests. Builds and evaluates reproducibility assessments and scoring, runs and interprets replication experiments, and produces the code, data, environment specifications, and protocols needed for others to reproduce analyses and training runs.
This work proposes an AI agent–driven workflow to address the high costs of reproducing large-scale empirical studies, which often stem from discrepancies in computational environments, code, and documentation. The approach decouples scientific reasoning from computational execution: researchers supply standardized diagnostic templates, and the system automatically retrieves and orchestrates reproduction materials within a version-controlled environment. A structured knowledge layer captures failure patterns, enabling adaptive reproduction across heterogeneous studies while ensuring transparency and stability of the analytical pipeline. Evaluated on 92 instrumental variable studies, the method achieves an 87% end-to-end reproduction success rate; when data and code are available, it attains 100% success at both the paper and model levels.
Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.
This work addresses the inefficiency and limited scalability of manual evaluation of software engineering reproducibility packages by introducing, for the first time, a multi-agent architecture for assessing reproducibility quality. The authors translate open science guidelines into 31 machine-verifiable reproducibility criteria and integrate rule-based engines with automated scripts to perform both static and dynamic analyses of code, environments, and artifacts. The system generates evidence-based recommendations for improvement and supports human-in-the-loop optimization. Experimental evaluation on five reproducibility packages demonstrates 91.4% execution consistency and 75.4% detection accuracy, while user studies confirm its practical utility and strong potential for adoption.
This study addresses the widespread lack of computational reproducibility in R supplementary code deposited on the Open Science Framework (OSF). A systematic audit of 296 published R code packages revealed that 98.8% incompletely declare dependencies. To address this, we propose the first automated reproducibility auditing framework tailored to the R ecosystem. It combines static source-code analysis—leveraging regular expressions and abstract syntax trees (ASTs)—to accurately infer dependencies, with Docker-based containerized execution and failure diagnostics (e.g., path errors, OS-specific inconsistencies, missing packages) to enable end-to-end environment reconstruction and validation. Experiments successfully executed 25.87% of scripts, identifying undeclared dependencies, hardcoded file paths, and cross-platform compatibility issues as the three primary barriers to reproducibility. The framework enables large-scale, low-cost, and scalable quantitative assessment of computational reproducibility in scholarly research, providing a practical toolchain to enhance transparency and verifiability.
Although top-tier conferences such as ICSE now commonly require authors to submit replication packages, the actual executability and reproducibility of these packages remain largely unassessed. This study presents a large-scale empirical investigation of 100 replication packages from ICSE papers published between 2015 and 2024, involving approximately 650 person-hours of manual execution, debugging, and root-cause analysis. The findings reveal that only 40% of the packages are executable, with just 32.5% running without modification; 82.5% require moderate to substantial changes. Among the executable packages, merely 35% successfully reproduce the original results. This work is the first to expose a significant gap between executability and reproducibility in software engineering replication packages and proposes three actionable guidelines to improve their reliability and utility.
研究使用MLReproMutate软件通过四种变异类别对机器学习研究仓库进行测试,发现现有验证流程常未能检测到影响可复现性的变化。
This study addresses the ambiguity and inconsistency in evaluation criteria for software engineering replication studies, which have led to contradictory interpretations and uncertainty in reported results. Through a systematic review of ten replication studies published between 2021 and 2025, combined with qualitative content analysis, statistical principles, and modeling of measurement uncertainty, this work is the first to uncover the heterogeneity and lack of standardized practices in current evaluation approaches. Building on these insights, the paper proposes a unified evaluation framework that integrates statistical theory, methodological rigor, and measurement theory. Empirical illustration demonstrates that the framework effectively enhances the transparency, consistency, comparability, and reliability of replication studies in software engineering.
This work addresses the limitations of existing reproducibility assessment methods, which rely on manual annotations and thus lack scalability and authentic supervision signals reflecting real-world reproduction challenges. The authors propose the first scalable evaluation framework that leverages GitHub user-submitted issues as natural supervision, enabling large-scale assessment of large language model (LLM) agents’ ability to identify paper-to-code reproducibility issues without human annotation. By integrating language understanding with code context analysis, the approach enables non-execution-based detection of reproducibility barriers. Experimental results demonstrate that the best-performing LLM agent identifies at least one semantically relevant reproducibility issue—aligned with those reported by humans—in approximately 90% of the evaluated papers, exhibiting strong performance in both failure detection and semantic localization.
This work proposes ReproScore, a novel framework that explicitly decouples reproducibility readiness (RRS) from reproducibility outcomes (ROS)—a distinction often conflated in existing tools that mistakenly treat static repository completeness as a proxy for successful execution. RRS comprises 26 fine-grained submetrics assessing code and documentation quality, while ROS is derived through sandboxed execution probes. The two dimensions are integrated into an adaptive composite score (RCS), with customizable metric weights via versioned YAML configurations contributed by the community. Empirical evaluation across 423 GitHub repositories spanning five failure modes demonstrates that environment-related metrics effectively discriminate failure types, yet reveals a near-zero correlation between RRS and actual execution success—exposing a critical “readiness–outcome gap” and thereby validating the necessity and efficacy of the proposed architectural separation.
The scale and complexity of modern scientific research render traditional peer review inadequate for effectively evaluating reproducibility. This work proposes a novel approach that frames reproducibility assessment as a structured reasoning task over scientific literature. It introduces an agent-based reasoning mechanism that leverages large language models to extract structured information, construct directed workflow graphs, and integrate multidimensional scoring for automated evaluation. The method enables consistent, cross-domain, and cross-model reconstruction of computational workflows, achieving accuracies of 61%, 60.71%, and 61.68% on the ReScience C, ReproBench, and GoldStandardDB benchmarks, respectively—substantially outperforming existing techniques.