preregistration practices

Formally specifying experimental protocols, hypotheses, analyses, and multiple-comparison controls before data collection to make claims falsifiable and statistically rigorous. This encompasses designing pre-registered evaluation procedures, replication plans, and robustness checks for empirical claims (e.g., LLM evaluation or survey experiments).

preregistrationpractices

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

StatWhy: Formal Verification Tool for Statistical Hypothesis Testing Programs

May 25, 2024
YK
Yusuke Kawamoto
🏛️ AIST | PRESTO | JST | University of Tsukuba | Kyoto University

Misuse of statistical hypothesis tests severely undermines scientific reliability. This paper proposes a formal verification methodology for statistical programs: preconditions—such as normality, independence, and homoscedasticity—are explicitly encoded as logical assertions in source code; static verification is then performed on OCaml implementations using the Why3 platform to automatically detect missing or conflicting assumptions. The approach innovatively integrates contract-based programming with formal verification, distinguishing between formalizable preconditions (amenable to automated checking) and non-formalizable ones (requiring expert judgment), thereby establishing a human-in-the-loop verification paradigm. Evaluated on canonical statistical tests—including Student’s *t*-test and ANOVA—the method successfully identifies widespread misuses, such as applying the *t*-test to non-normal data or neglecting homoscedasticity checks. Results demonstrate significant improvements in the correctness, auditability, and reproducibility of statistical software.

Automatically check requirements for statistical methods in codeFormally verify correctness of statistical hypothesis testing programsPrevent common errors in statistical program implementation

This study addresses the challenge of error-prone manual verification of tables, figures, and listings (TFLs) in clinical trial reports, which often fails to detect structural or logical inconsistencies. The authors propose PROVE, a novel framework that leverages large language models (LLMs) for semantic parsing and evidence tracing of TFL content, integrated with a programmable rule engine to perform deterministic numerical and logical validation against SDTM/ADaM standards. Designed as a multi-agent architecture, PROVE combines LLM-driven semantic understanding with rule-based checks to enable auditable, configurable automated cross-verification. The approach achieves 100% accuracy under exact label matching; when confronted with linguistic variations, LLM assistance boosts recall from 0.588 to 0.993 and F1 score from 0.735 to 0.996.

clinical trial reportingcross-output consistencyregulatory compliance

This study examines whether preregistration enhances the severity—i.e., the capacity to detect erroneous interpretations—of hypothesis testing, within Popperian falsificationism and Mayo’s error-statistical framework. Method: Through conceptual analysis, philosophical logic, and error-statistical modeling, the paper rigorously evaluates preregistration’s theoretical and practical implications for severity assessment. Contribution/Results: It demonstrates, for the first time, that preregistration fails to strengthen severity in the Popperian paradigm, as it does not alter a theory’s inherent falsifiability. Moreover, under realistic “planning-not-prison” practices—where deviations from preregistered protocols are permitted—the Type I error rate becomes uncontrolled, undermining Mayo’s severity criterion. Consequently, preregistration neither fortifies falsificationist logic nor ensures inferential reliability. These findings challenge the prevailing methodological consensus on preregistration’s epistemic superiority and provide foundational philosophical and statistical grounds for reevaluating scientific practice and metascientific policy.

Deviations in preregistered tests obscure Type I error rate transparencyPreregistration doesn't improve severity assessment when flexible implementations are permittedPreregistration fails to enhance severity evaluation in Popper's theory-centric approach

This study addresses the challenge that data-driven protocol selection in target trial emulation often invalidates statistical inference. To resolve this, the authors propose a two-stage strategy based on sample splitting: the first subsample is used to explore and finalize the target trial protocol, while the second, independent subsample is employed to implement the selected protocol and conduct causal inference. Inspired by the exploratory-to-confirmatory paradigm in clinical trials, this approach decouples protocol specification from inference, thereby preserving the flexibility of scientific exploration while rigorously maintaining nominal coverage guarantees for statistical inference. By integrating sample splitting, target trial emulation, and causal inference, the work provides both theoretical assurance and a practical framework for valid statistical inference in observational studies.

iterative protocol developmentobservational dataselective choices

AI scientist systems face diminished statistical rigor and elevated false discovery risk due to dynamic hypothesis testing. Method: This paper proposes a structured, functional-programming–based assurance framework centered on a novel Research Monad and declarative scaffolding mechanism. Implemented as a Haskell embedded domain-specific language (eDSL) with a monad transformer stack, it enforces online false discovery rate (FDR) correction, strict data isolation, and state consistency constraints *during* LLM-generated code execution—thereby eliminating data leakage and multiple-comparison bias at the architectural level. Contribution/Results: Evaluated across 2,000 simulated hypothesis tests and end-to-end scientific case studies, the framework significantly improves statistical robustness and reproducibility of automated discoveries, establishing a formal, trustworthy foundation for AI-driven scientific research.

Addressing methodological errors like data leakage in hybrid architecturesEnforcing statistical rigor in AI-driven discovery systemsPreventing spurious discoveries in dynamic hypothesis testing

Latest Papers

What's happening recently
View more

This work addresses the lack of standardized, auditable verification mechanisms for autonomous agents in regulated domains, where existing approaches either produce opaque conclusions or non-reproducible private logs. We propose a protocol-layer solution grounded in the four epistemic sources from Indian epistemology—perception, inference, analogy, and testimony—by encapsulating critical outputs as typed ClaimAttestations accompanied by deterministic or conditionally reproducible verify() operations, enabling offline auditability. We define the first cross-vendor reproducible wire format for claim verification and validate the protocol’s correctness and practicality through TLA+ formal modeling (covering 38,563 states with no invariant violations), a Python reference implementation passing 84 test cases, and LLM-based adjudicator experiments. Pilot results indicate that reference implementation quality significantly impacts false positive rates, with differences reaching up to 40 percentage points.

auditabilityautonomous agentsclaim verification

This work addresses the widespread misuse of statistical methods—often stemming from implicit or ambiguous assumptions—which exacerbates the reproducibility crisis in scientific research, particularly in hypothesis testing and meta-analysis where formal verification mechanisms are lacking. To bridge this gap, the authors propose the first formal verification framework tailored for Python-based statistical programs. By developing a Why3-py frontend, they translate dynamically typed, runtime-polymorphic Python code into the WhyML intermediate representation and extend the StatWhy tool to support meta-analysis verification. Integrating program transformation, static analysis, and formal verification techniques, this approach enables, for the first time, automated correctness verification of statistical programs written in Python, effectively uncovering overlooked assumptions and misuses, thereby filling a critical void in the formal verification of statistical software.

formal verificationhypothesis testingmeta-analysis

This study reveals that large language models struggle to effectively assess the veracity of statistical evidence when integrating multi-source information, exhibiting a tendency to rely on superficial stylistic cues in methodological text rather than numerical plausibility when judging source credibility. The work identifies a previously undocumented “cognitive alignment” bias—where models prefer sources with strong analytical register over those with content consistency. Employing interpretable techniques including causal tracing, linear probing (AUC: 0.83–0.92), and component-level attribution, the authors replicate this blind spot across five mainstream models through cross-model and cross-domain experiments. Further analysis localizes the issue to a methodology-register gating mechanism and demonstrates that neither prompt engineering nor post-training interventions adequately mitigate the bias, instead raising concerns about model generalization.

epistemic alignmentlarge language modelsmethodology-register

This work addresses the tendency of current AI-driven scientific systems to generate claims that exceed the scope of supporting evidence. It proposes a “claim calibration” framework that models AI-assisted research as an iterative cycle comprising hypothesis generation, consequence derivation, external validation, belief updating, and claim calibration, emphasizing that scientific assertions must be constrained by evidential warrant. The framework distinguishes four semantic forms of claims, defines the claim-evidence gap and associated epistemic debt, and introduces minimal structural revision as a calibration pathway. Validation is demonstrated through multi-agent collaboration paradigms, AI scientist pipelines, and the AISim-Cal synthetic dynamics example. The study establishes three guiding principles—including “no claim without warrant”—to construct an iterative, reliable evaluation loop for trustworthy AI-enabled scientific research.

AI-assisted researchcalibrationclaim-evidence gap

This study introduces the novel concept of “clinical trial engineering”—the systematic manipulation of statistical analyses to generate misleading clinical trial evidence in support of drug approval, distinct from conventional paper mills. Focusing on 23 studies linked to Iran’s CinnaGen and its subsidiary Orchid Pharmed, the authors applied the INSPECT-SR credibility framework, integrating PubMed literature screening, raw data verification, and co-authorship network analysis to systematically evaluate evidentiary reliability. The investigation uncovered 180 issues spanning nine categories of systemic bias, including incomplete reporting, arithmetic errors, and design flaws. These findings reveal a structural pattern of research manipulation driven by commercial pressures, publication incentives, and permissive regulatory pathways, prompting regulatory agencies to reassess the credibility of the associated clinical evidence.

clinical trial engineeringpublication biasregulatory approval

Hot Scholars

GS

Gerald Schweiger

Professor (Full) Vienna University of Technology
DataScience of ScienceEnergyBuildings