conduct replication analyses

Designs and implements reproducible analysis pipelines and statistical tests to reproduce published experimental findings and reported effect sizes, including recomputing anomaly returns or analogous metrics and performing out-of-sample replication and backtests. Builds comparative and robustness analyses that compare outcomes across datasets and models, quantify inter-run and inter-annotator consistency and limits of agreement, and run subset and sensitivity tests (for example excluding small entities or splitting pre/post periods) to assess reliability.

conductreplicationanalyses

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an AI agent–driven workflow to address the high costs of reproducing large-scale empirical studies, which often stem from discrepancies in computational environments, code, and documentation. The approach decouples scientific reasoning from computational execution: researchers supply standardized diagnostic templates, and the system automatically retrieves and orchestrates reproduction materials within a version-controlled environment. A structured knowledge layer captures failure patterns, enabling adaptive reproduction across heterogeneous studies while ensuring transparency and stability of the analytical pipeline. Evaluated on 92 instrumental variable studies, the method achieves an 87% end-to-end reproduction success rate; when data and code are available, it attains 100% success at both the paper and model levels.

empirical dataexecution bottlenecklarge-scale reanalysis

A Dataset For Computational Reproducibility

Apr 11, 2025
LC
Lázaro Costa
🏛️ University of Porto | INESC TEC

Scientific computing artifacts—such as analysis scripts and software prototypes—frequently suffer from poor reproducibility due to environmental heterogeneity, dependency drift, and inadequate documentation, thereby undermining research credibility. To address this, we introduce the first cross-disciplinary, structured, and standardized benchmark dataset for computational experiments, encompassing workflows ranging from single-script executions to multi-language, complex pipelines. Our framework uniformly models metadata, standardizes dependency declarations (e.g., requirements.txt, Dockerfiles), encapsulates multi-language execution procedures, and prescribes a rigorous documentation protocol. The dataset comprises dozens of human-validated, fully reproducible experimental cases, enabling objective, comparable, and reproducible evaluation of reproducibility tools. This work fills a critical gap in the field by providing the first systematic, community-grounded benchmark for assessing computational reproducibility, thereby significantly enhancing the rigor, transparency, and comparability of reproducibility research.

Addressing variability in computational environments and softwareEnsuring reproducibility of computational scientific workProviding standardized dataset for evaluating reproducibility tools

Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.

Compares actual experimental processes with textbook methodologies in detailExplores mismatches between proposed replication procedures and researchers' needsInvestigates how experimental researchers conduct experiments in practice

Preclinical reproducibility assessment traditionally relies on costly additional replicate experiments, limiting scalability and efficiency. Method: This study proposes leveraging inherent internal replication—such as across batches, sites, and litters—as a quantifiable resource for reproducibility evaluation. We systematically define six classes of internal replication structures and develop a statistical inference framework integrating mixed-effects modeling, variance decomposition, and multi-site collaborative analysis, augmented by a formal reproducibility hypothesis test. Contribution/Results: Validated on a three-center mouse study, the method significantly enhances statistical robustness and inferential reliability without requiring new experiments. It delivers an immediately deployable, data-driven tool for preclinical reproducibility assessment, enabling a paradigm shift from experiment-driven to data-driven reproducibility evaluation.

Evaluating reproducibility in preclinical experiments using internal replicationProviding a framework for robust statistical inferences in preclinical researchQuantifying internal reproducibility without additional costly replication studies

What is Reproducibility in Artificial Intelligence and Machine Learning Research?

Apr 29, 2024
AD
Abhyuday Desai
🏛️ Ready Tensor, Inc. | Georgetown University

The AI/ML community faces a severe reproducibility crisis, primarily driven by conceptual ambiguity in verification terminology—such as “reproducibility,” “replicability,” and “dependency/independence”—which undermines research credibility and scientific progress. To address this, we propose the first five-dimensional verification taxonomy, systematically defining core concepts—including reproducibility, dependency vs. independent re-executability, and direct vs. conceptual replicability—by clarifying their objectives, prerequisites, and evaluation criteria. Our framework integrates conceptual analysis, terminological standardization, and methodological modeling to yield a structured verification guideline. It enhances experimental rigor in study design, fosters consensus across the research community on verification practices, and significantly improves cross-team result reproducibility and outcome reliability.

Addressing the reproducibility crisis in AI/ML researchClarifying validation terminology in AI/ML reproducibilityProviding a framework for validation study design

Latest Papers

What's happening recently
View more

This work proposes ReproScore, a novel framework that explicitly decouples reproducibility readiness (RRS) from reproducibility outcomes (ROS)—a distinction often conflated in existing tools that mistakenly treat static repository completeness as a proxy for successful execution. RRS comprises 26 fine-grained submetrics assessing code and documentation quality, while ROS is derived through sandboxed execution probes. The two dimensions are integrated into an adaptive composite score (RCS), with customizable metric weights via versioned YAML configurations contributed by the community. Empirical evaluation across 423 GitHub repositories spanning five failure modes demonstrates that environment-related metrics effectively discriminate failure types, yet reveals a near-zero correlation between RRS and actual execution success—exposing a critical “readiness–outcome gap” and thereby validating the necessity and efficacy of the proposed architectural separation.

digital librariesexecutable assessmentreadiness-outcome conflation

This work addresses the limitations of existing reproducibility assessment methods, which rely on manual annotations and thus lack scalability and authentic supervision signals reflecting real-world reproduction challenges. The authors propose the first scalable evaluation framework that leverages GitHub user-submitted issues as natural supervision, enabling large-scale assessment of large language model (LLM) agents’ ability to identify paper-to-code reproducibility issues without human annotation. By integrating language understanding with code context analysis, the approach enables non-execution-based detection of reproducibility barriers. Experimental results demonstrate that the best-performing LLM agent identifies at least one semantically relevant reproducibility issue—aligned with those reported by humans—in approximately 90% of the evaluated papers, exhibiting strong performance in both failure detection and semantic localization.

benchmarkingGitHub issuesLLM agents

This study addresses the critical issue of declining reproducibility in quantum software defect datasets—such as Bugs4Q—due to dependency evolution, which undermines research reliability. The authors present the first systematic evaluation of this reproducibility degradation by reproducing 37 bugs across 21 Qiskit versions through 77,700 executions. Combining root cause analysis, dependency management, and API migration insights, they demonstrate that 93.6% of reproduction failures stem from environmental dependency issues rather than actual bug disappearance. Based on these findings, they propose a novel maintenance paradigm requiring source-level fixes and introduce an enhanced dataset, Bugs4Q-Robust, which boosts the reproduction rate from 16.2% to 78.4% on Qiskit v2.3.1—substantially outperforming conventional version-locking approaches.

Bugs4Qdefect datasetsdependency evolution

Hot Scholars

MD

Matias D. Cattaneo

Princeton University
EconometricsStatisticsMachine LearningData Science
MK

Minwoo Kang

University of California, Berkeley
Language ModelsNatural Language Processing
SC

Sarath Chandar

Associate Professor @ Polytechnique Montreal. Mila. Canada CIFAR AI Chair. Canada Research Chair.
Artificial IntelligenceMachine LearningDeep LearningReinforcement Learning