Score
Designs and executes comparative evaluations of measurement methods, building reproducible benchmarks and analyses that quantify measurement-error properties such as bias, variance, and interval coverage across methods. Produces reproducible analytic code and examples to implement multiple correction approaches side‑by‑side and to identify the data conditions or contexts that favor each method.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
Scientific software selection frequently suffers from non-reproducible benchmarks due to multi-library, multi-metric evaluation and dynamic evolution—such as the introduction of new algorithms or modifications to test cases and evaluation criteria. This paper addresses numerical integration over arbitrary 2D/3D domains with implicit or parameterized boundaries (cut-cell quadrature), proposing the first automated benchmarking framework that systematically integrates CI/CD engineering practices into scientific computing workflows. The framework unifies GitHub Actions, Docker, Python-based scheduling, Jupyter-based report generation, and semantically versioned result archiving. It supports automated configuration, execution, visualization, and historical result comparison. It achieves >90% automation for benchmark tasks and regression detection; reduces integration time for new libraries or algorithms by 70%; and enables precise attribution of performance deviations to specific code commits. The framework significantly enhances reliability, reproducibility, and evolutionary adaptability in scientific software evaluation.
Despite growing adoption of non-targeted analysis (NTA) in food safety, systematic evaluation of liquid/gas chromatography–high-resolution mass spectrometry (LC/GC-HRMS) NTA tools against FAIR principles and the BP4NTA operational pillars—laboratory validation, data/code availability, standardized formats, knowledge integration, and portable implementation—remains lacking. Method: We conducted a longitudinal, systematic audit of 103 NTA tools published between 2004 and 2025, assessing compliance across all six BP4NTA pillars. Contribution/Results: We found a marked increase in openness (56% → 86%) but a paradoxical decline in practical reproducibility (55% → 43%), quantifying for the first time the persistent “discoverable but not runnable” gap. Critical synergistic deficits between Pillar C1 (validation) and C6 (portable implementation) emerged as the primary bottleneck for regulatory-grade reproducibility. This study fills a key gap in food safety NTA tool assessment and proposes a multidimensional audit framework; only 17% of tools satisfy both validation and portability criteria—providing empirical grounding and actionable pathways toward fully reproducible NTA workflows.
This study addresses the critical issue of declining reproducibility in quantum software defect datasets—such as Bugs4Q—due to dependency evolution, which undermines research reliability. The authors present the first systematic evaluation of this reproducibility degradation by reproducing 37 bugs across 21 Qiskit versions through 77,700 executions. Combining root cause analysis, dependency management, and API migration insights, they demonstrate that 93.6% of reproduction failures stem from environmental dependency issues rather than actual bug disappearance. Based on these findings, they propose a novel maintenance paradigm requiring source-level fixes and introduce an enhanced dataset, Bugs4Q-Robust, which boosts the reproduction rate from 16.2% to 78.4% on Qiskit v2.3.1—substantially outperforming conventional version-locking approaches.
This work addresses the frequent neglect of sampling strategy design and generalizability in software engineering research, which often undermines the representativeness of empirical findings. To remedy this, the paper introduces a domain-specific language (DSL) that explicitly models complex sampling workflows over code repositories through composable sampling operators, enabling—for the first time—formal specification and reasoning about the generalizability of sampling strategies. Implemented as a fluent Python API, the DSL is integrated with a statistical metric system to quantitatively assess the external validity of sampled datasets. The authors demonstrate the expressiveness and practical utility of their approach by reconstructing and formalizing the sampling procedures from multiple Mining Software Repositories (MSR) studies, thereby validating the framework’s capacity to capture real-world methodological diversity.
This work addresses the significant limitations of spreadsheet-based analysis in reproducibility, auditability, version control, and automation. It proposes a migration pathway from Excel to research-grade analytical workflows by leveraging Python’s pandas library as a bridge. The study introduces an innovative set of Excel-to-pandas mapping rules, categorizes nine canonical workflow patterns, and compiles a catalog of common failure modes. Seven end-to-end real-world examples demonstrate the approach in practice. By retaining Excel as a familiar interface for input and output while integrating version control, automated refreshing, and seamless incorporation of statistical and machine learning methods, the proposed framework enables governed, reproducible, and auditable tabular data analysis.
This study addresses the ambiguity and inconsistency in evaluation criteria for software engineering replication studies, which have led to contradictory interpretations and uncertainty in reported results. Through a systematic review of ten replication studies published between 2021 and 2025, combined with qualitative content analysis, statistical principles, and modeling of measurement uncertainty, this work is the first to uncover the heterogeneity and lack of standardized practices in current evaluation approaches. Building on these insights, the paper proposes a unified evaluation framework that integrates statistical theory, methodological rigor, and measurement theory. Empirical illustration demonstrates that the framework effectively enhances the transparency, consistency, comparability, and reliability of replication studies in software engineering.
This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.
This study addresses a critical gap in quantum software research: the absence of a systematic auditing mechanism for empirically grounded comparative claims, which has led to a pervasive “instantiation gap” characterized by insufficient evidentiary support. To bridge this gap, the authors propose CLAIMSTAB-QC, the first source-bound auditing framework tailored to empirical comparisons in quantum software. By integrating claim modeling, audit scope delimitation, evidence boundary identification, and directional classification, the framework enables precise validation of comparative assertions against original source materials. An evaluation across 455 claims from 119 papers reveals that only eight claims possessed sufficient matched evidence for direct auditing; among these, two were confirmed, four lacked adequate support, and two were contradicted—highlighting substantial deficiencies in the empirical rigor of current quantum software studies.