Score
Designs, implements, and evaluates software libraries, packages, and applications that encode statistical and mathematical algorithms for data analysis, modeling, inference, and visualization. Work includes ensuring numerical correctness, performance, API usability, testing, documentation, and reproducibility of statistical computations.
This study investigates cross-disciplinary trends in statistical software adoption across economics, political science, and statistics. Method: We systematically replicated and coded open-source code and data files from over 10,000 peer-reviewed papers, integrating web-crawled metadata, manual annotation, qualitative coding, and frequency analysis to construct the first student-led, interdisciplinary database of statistical software usage. Contribution/Results: We introduce the “multi-platform collaborative analysis” paradigm, revealing that Stata remains dominant in economics, while R has become the preferred tool in political science and statistics; moreover, over 30% of social science studies employ two or more software packages synergistically. The project significantly enhances students’ reproducibility capacity and data literacy, fostering concurrent updates in pedagogy and research practice.
Misuse of statistical hypothesis tests severely undermines scientific reliability. This paper proposes a formal verification methodology for statistical programs: preconditions—such as normality, independence, and homoscedasticity—are explicitly encoded as logical assertions in source code; static verification is then performed on OCaml implementations using the Why3 platform to automatically detect missing or conflicting assumptions. The approach innovatively integrates contract-based programming with formal verification, distinguishing between formalizable preconditions (amenable to automated checking) and non-formalizable ones (requiring expert judgment), thereby establishing a human-in-the-loop verification paradigm. Evaluated on canonical statistical tests—including Student’s *t*-test and ANOVA—the method successfully identifies widespread misuses, such as applying the *t*-test to non-normal data or neglecting homoscedasticity checks. Results demonstrate significant improvements in the correctness, auditability, and reproducibility of statistical software.
This study systematically identifies critical barriers in biostatistical software development: information silos across roles and departments lead to redundant implementation; low code readability and reusability hinder maintainability; and the absence of standardized version control and testing practices severely compromises result reproducibility and long-term sustainability. To address these challenges, we propose— for the first time—a lightweight, discipline-aware software engineering framework tailored to biostatistics. It integrates foundational practices—including Git-based version control, unit testing, and documentation standards—while embedding cross-functional collaboration mechanisms, domain-adapted code review protocols, and structured knowledge-sharing strategies. The resulting framework yields a scalable, field-tested capability-building guide that demonstrably improves the reliability and reproducibility of statistical modeling code, enhances team collaboration efficiency, and bridges the practice gap between statisticians and software engineering principles.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
Doctoral students in life sciences commonly lack formal software engineering training, hindering the development of robust, reproducible, and collaborative research software. Method: This study proposes ten pedagogical principles for research software development, establishing the first systematic framework centered on “research software pedagogy”—distinct from generic programming instruction. It integrates software engineering best practices (e.g., Git-based version control, CI/CD pipelines, unit testing, RESTful API design), learning science principles, and authentic research workflows, emphasizing the seamless embedding of automation, documentation, testing, and collaborative practices throughout the research lifecycle. Contribution/Results: The framework delivers a generalizable, plug-and-play pedagogical paradigm. Deployed across multiple Chinese universities’ life sciences PhD programs, it has demonstrably improved software deliverable quality, code reusability, and cross-team collaboration efficiency—bridging critical gaps between computational literacy and rigorous, team-based scientific software practice.
研究设计并实施了一个虚拟统计计算实验室,通过两种R编码教学方式提高统计学入门学生的概念学习和数据科学准备度。
This study addresses the challenge that existing AI code generation tools often fail to ensure fidelity in the software implementation of statistical methods, thereby introducing implementation distortions. To mitigate this issue, the authors propose a novel multi-agent development paradigm built upon Claude Code, incorporating an information isolation mechanism. In this framework, a planning agent generates separate specifications for implementation, simulation, and testing, which are then executed by dedicated agents operating in mutual isolation. This approach pioneers the use of information barriers in AI-assisted programming, eliminating reliance on prior knowledge in code generation while preserving researchers’ full control over methodological decisions. Empirical evaluations demonstrate that the workflow successfully implements probit estimation and integrates with multiple R and Python statistical packages, effectively offloading engineering overhead without compromising implementation accuracy.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
Scientific computing notebooks frequently suffer from irreproducibility, poor readability, and limited reusability, posing serious threats to research reliability. This work presents the first large-scale empirical study of 1,510 Jupyter notebooks from 518 code repositories published in Nature in 2024. Through manual reproduction attempts (only 2 successful out of 19), documentation review, code clone detection (≥10 lines, ≥3 instances), and mutation analysis, the study systematically uncovers pervasive issues including chaotic state management, missing dependencies, and excessive code duplication. To address these challenges, the authors propose the first multidimensional quality assessment framework explicitly designed to evaluate reproducibility, readability, and reusability, thereby establishing an empirical foundation and methodological support for improving the quality of scientific code.
This work addresses the absence of a systematic, traceable, and reproducible framework for reporting the performance of mathematical libraries—a gap that hinders accurate performance evaluation and resource planning for scientific applications on high-performance computing (HPC) systems. To this end, the paper introduces LAAB, the first framework explicitly designed around four core principles: traceability, compatibility, reliability, and accessibility. LAAB establishes an end-to-end reproducible performance evaluation pipeline through standardized benchmarking protocols, comprehensive metadata management, execution environment tracking, and advanced performance analysis techniques. The framework substantially enhances the accuracy and interoperability of mathematical library performance reporting, thereby providing a robust foundation for performance prediction and resource scheduling in scientific computing.