Score
Designs and executes experimental setups, reimplementations, and debugging workflows to reproduce reported results, bugs, or technical issues described in research papers or system reports. Builds artifacts such as test cases, datasets, code reimplementations, scripts, and analysis logs, and analyzes discrepancies between original and reproduced outcomes to identify causes and required corrections.
This work addresses the frequent neglect of sampling strategy design and generalizability in software engineering research, which often undermines the representativeness of empirical findings. To remedy this, the paper introduces a domain-specific language (DSL) that explicitly models complex sampling workflows over code repositories through composable sampling operators, enabling—for the first time—formal specification and reasoning about the generalizability of sampling strategies. Implemented as a fluent Python API, the DSL is integrated with a statistical metric system to quantitatively assess the external validity of sampled datasets. The authors demonstrate the expressiveness and practical utility of their approach by reconstructing and formalizing the sampling procedures from multiple Mining Software Repositories (MSR) studies, thereby validating the framework’s capacity to capture real-world methodological diversity.
Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.
Automatically reproducing executable bug-fix code pairs from unstructured developer Q&A posts is hindered by ambiguous descriptions and missing dependencies. This work proposes Reprodgen, the first end-to-end automated framework that leverages large language models to jointly model code intent (CI), functional requirements (FR), and structured chains of thought (SCoT) to generate semantically consistent and executable bug-fix code pairs. The approach incorporates an LLM-based iterative review mechanism coupled with real execution validation to ensure correctness. Evaluated on Stack Overflow and GitHub Issues across seven widely used data science libraries, the study introduces the first expert-validated, runnable benchmark of bug-fix pairs. Experimental results demonstrate that Reprodgen reliably reproduces code pairs exhibiting clear behavioral differences between buggy and fixed versions.
Scientific software testing faces unique challenges—including difficult test case design, ambiguous oracle determination, absence of quality assessment standards, and poor applicability of industrial testing tools. Method: We conducted the first large-scale empirical study, combining structured surveys with qualitative analysis and statistical testing across 217 scientific software developers to examine variations in testing practices, tool adoption, and demographic factors. Contribution/Results: We identify three core bottlenecks: test design, result validation, and quality measurement; further reveal widespread lack of awareness of and access to domain-specific testing tools. This work establishes the first empirical evidence of paradigmatic divergence between scientific and conventional software testing, advocating for lightweight, extensible, and computationally aware testing frameworks tailored to scientific computing. Our findings provide foundational evidence and strategic direction for advancing domain-specific testing methodology.
This study addresses the widespread lack of computational reproducibility in R supplementary code deposited on the Open Science Framework (OSF). A systematic audit of 296 published R code packages revealed that 98.8% incompletely declare dependencies. To address this, we propose the first automated reproducibility auditing framework tailored to the R ecosystem. It combines static source-code analysis—leveraging regular expressions and abstract syntax trees (ASTs)—to accurately infer dependencies, with Docker-based containerized execution and failure diagnostics (e.g., path errors, OS-specific inconsistencies, missing packages) to enable end-to-end environment reconstruction and validation. Experiments successfully executed 25.87% of scripts, identifying undeclared dependencies, hardcoded file paths, and cross-platform compatibility issues as the three primary barriers to reproducibility. The framework enables large-scale, low-cost, and scalable quantitative assessment of computational reproducibility in scholarly research, providing a practical toolchain to enhance transparency and verifiability.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
This work addresses the inefficiency and steep learning curve researchers often encounter when trying to map academic papers to their corresponding implementation code. To bridge this gap, the authors propose an automated tool powered by large language models (LLMs) that achieves cross-modal semantic alignment between scholarly texts and source code for the first time. By integrating program analysis techniques, the method automatically identifies code segments that implement specific research ideas described in a paper and generates high-quality traceability mappings. This approach substantially reduces the manual effort required for alignment, enhances the comprehensibility of research software, and improves reproducibility. Preliminary experiments demonstrate the tool’s practicality and effectiveness in real-world scenarios.
Debugging in data-intensive programming faces significant challenges, including fragmented evidence, difficulty in discerning discrepancies between expected and observed behaviors, and the complexity of tracking state evolution across components. Through semi-structured interviews and thematic analysis, this study systematically characterizes practitioners’ debugging practices and, for the first time, identifies three core requirements: cross-artifact evidence alignment, expectation-based comparison mechanisms, and traceable state evolution. Building on these insights, the work constructs a visualization-driven design space tailored to debugging in data-intensive contexts, exposing critical gaps in existing tools and providing a theoretical foundation and clear direction for the development of future debugging aids.
This study addresses the challenge of verifying defects in AI research—particularly errors relying on prior knowledge—when only outputs are available. To this end, it introduces a novel "research contract" mechanism that binds experimental choices, execution obligations, and evidence, thereby establishing verification boundaries under information asymmetry. The proposed approach enables automated verification through deterministic checkers, registration of fault-specific rules, and metadata filtering, while explicitly distinguishing contract-relative verification from scientific truth. Experimental results demonstrate that the system successfully detected all eight registered mutations and identified 104 defects across 144 variant cases, effectively ensuring the reliability of compliance verification.
Scientific computing notebooks frequently suffer from irreproducibility, poor readability, and limited reusability, posing serious threats to research reliability. This work presents the first large-scale empirical study of 1,510 Jupyter notebooks from 518 code repositories published in Nature in 2024. Through manual reproduction attempts (only 2 successful out of 19), documentation review, code clone detection (≥10 lines, ≥3 instances), and mutation analysis, the study systematically uncovers pervasive issues including chaotic state management, missing dependencies, and excessive code duplication. To address these challenges, the authors propose the first multidimensional quality assessment framework explicitly designed to evaluate reproducibility, readability, and reusability, thereby establishing an empirical foundation and methodological support for improving the quality of scientific code.