statistical reporting

The practice of recording, summarizing, and presenting evaluation metrics and statistical results in a standardized, reproducible way so tuning and benchmarking can be replicated across frameworks and models. It includes protocols for comparing outputs to baselines and documenting experimental setups and uncertainty.

statisticalreporting

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.

Compares actual experimental processes with textbook methodologies in detailExplores mismatches between proposed replication procedures and researchers' needsInvestigates how experimental researchers conduct experiments in practice

This study addresses the challenge of limited reproducibility and transparency in software engineering controlled experiments, often stemming from inadequate documentation. While generic preregistration templates—such as those provided by the Open Science Framework (OSF)—exist, they fail to comprehensively address the specific needs of software engineering research. This work presents the first systematic evaluation of the OSF preregistration template’s applicability to software engineering experiments, combining literature analysis, template comparison, and cross-referencing against established software engineering experimental reporting guidelines. The findings reveal that although existing OSF templates partially satisfy methodological requirements, none fully encompass all critical elements, and their customization capabilities are constrained. Based on these insights, the paper advocates for and provides a foundation toward developing a domain-specific, standardized preregistration template tailored to software engineering, thereby filling a critical gap in the field.

controlled experimentsempirical software engineeringregistered reports

Preclinical reproducibility assessment traditionally relies on costly additional replicate experiments, limiting scalability and efficiency. Method: This study proposes leveraging inherent internal replication—such as across batches, sites, and litters—as a quantifiable resource for reproducibility evaluation. We systematically define six classes of internal replication structures and develop a statistical inference framework integrating mixed-effects modeling, variance decomposition, and multi-site collaborative analysis, augmented by a formal reproducibility hypothesis test. Contribution/Results: Validated on a three-center mouse study, the method significantly enhances statistical robustness and inferential reliability without requiring new experiments. It delivers an immediately deployable, data-driven tool for preclinical reproducibility assessment, enabling a paradigm shift from experiment-driven to data-driven reproducibility evaluation.

Evaluating reproducibility in preclinical experiments using internal replicationProviding a framework for robust statistical inferences in preclinical researchQuantifying internal reproducibility without additional costly replication studies

This work addresses the lack of systematic and rigorous performance benchmarking methodologies in programming language research, which has undermined the credibility of evaluation results. To remedy this, the paper introduces a closed-loop methodology—Measure-Explain-Test-Improve—that establishes, for the first time, a structured and reproducible workflow for performance assessment in the field. Integrating systematic experimental design, performance metric analysis, result interpretation, and iterative refinement, the approach emphasizes theoretical grounding and practical rigor at every stage. Its key contribution lies in enabling even researchers with limited empirical experience to conduct reliable and methodologically sound performance evaluations, thereby significantly enhancing the scientific validity and reproducibility of performance analysis in programming language research.

benchmarkingperformance evaluationprogramming language research

Latest Papers

What's happening recently
View more

This study addresses the ambiguity and inconsistency in evaluation criteria for software engineering replication studies, which have led to contradictory interpretations and uncertainty in reported results. Through a systematic review of ten replication studies published between 2021 and 2025, combined with qualitative content analysis, statistical principles, and modeling of measurement uncertainty, this work is the first to uncover the heterogeneity and lack of standardized practices in current evaluation approaches. Building on these insights, the paper proposes a unified evaluation framework that integrates statistical theory, methodological rigor, and measurement theory. Empirical illustration demonstrates that the framework effectively enhances the transparency, consistency, comparability, and reliability of replication studies in software engineering.

Empirical StudiesEvaluation CriteriaReplication Assessment

This work proposes an AI agent–driven workflow to address the high costs of reproducing large-scale empirical studies, which often stem from discrepancies in computational environments, code, and documentation. The approach decouples scientific reasoning from computational execution: researchers supply standardized diagnostic templates, and the system automatically retrieves and orchestrates reproduction materials within a version-controlled environment. A structured knowledge layer captures failure patterns, enabling adaptive reproduction across heterogeneous studies while ensuring transparency and stability of the analytical pipeline. Evaluated on 92 instrumental variable studies, the method achieves an 87% end-to-end reproduction success rate; when data and code are available, it attains 100% success at both the paper and model levels.

empirical dataexecution bottlenecklarge-scale reanalysis

Design-oriented visualization research often struggles to meet conventional reproducibility standards due to its inherent subjectivity, contextual dependence, and iterative nature, thereby limiting its transparency and rigor. To address this challenge, this work proposes “traceability” as a viable alternative to traditional reproducibility. It presents the first systematic theoretical framework centered on three core components—recording, reporting, and reading—and introduces tRRRacer, a supporting tool implementing this framework. Through collaborative autoethnography, the authors reflect on practical applications of traceability in design-oriented research, demonstrating its feasibility and yielding actionable principles alongside theoretical insights. This approach offers a novel pathway to enhance the rigor and transparency of such studies without relying on strict reproducibility criteria.

design-oriented researchreproducibilitytraceability

This study addresses the longstanding limitation in networking research caused by the scarcity of reliable and reproducible experimental data. To overcome this challenge, the authors propose employing high-fidelity software models as substitutes for physical devices, enabling the construction of reproducible validation environments through “natural experiments” and establishing a reproducibility-based criterion for experimental reliability. A systematic evaluation framework is developed to conduct full-scale assessments of mainstream network simulation tools. The findings reveal that while most tools technically satisfy reproducibility requirements, their adoption in practice is predominantly influenced by non-technical factors such as popularity and user familiarity. This work introduces a new paradigm and benchmark for reproducible experimentation in networking research.

Experimental DataNetwork ModelingReliability

Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.

benchmark evidencecertifiable fractiondeployment action

Hot Scholars

RD

Richard D. Gill

Emeritus professor of Mathematical Statistics, Leiden University
StatisticsProbabilityMathematicsQuantum foundations
SC

Stephen Casper

PhD student, MIT
AI safetyAI responsibilityred-teamingrobustness
VM

Vasilios Mavroudis

Research Scientist, Alan Turing Institute
Machine LearningSystems SecurityArtificial Intelligence
RB

Rishi Bommasani

CS PhD, Stanford University
Societal Impact of AIAI PolicyAI GovernanceFoundation Models
SK

Sayash Kapoor

CS PhD, Princeton University
ReproducibilityAI agentsSocietal impacts