Score
Designs and implements quantitative analyses of observed data to estimate parameters, test hypotheses, and quantify uncertainty to support reproducible inference. This includes selecting and applying appropriate statistical models and tests, checking assumptions and diagnostics, and producing summary estimates, visualizations, and uncertainty measures (e.g., confidence intervals, p-values, effect sizes).
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.
Classical variance change-point detection methods suffer from p-value bias and inflated Type I error due to data reuse in model selection. Existing post-selection inference (PSI) frameworks are restricted to mean-shift detection and do not extend to variance changes. Method: This paper introduces the first PSI framework for variance change-point detection, proposing two general-purpose constructions for post-selection p-values compatible with diverse algorithms (e.g., piecewise constant modeling) and test forms (e.g., constrained likelihood ratio tests). Leveraging conditional inference, convex optimization, and statistical functional theory, the methods rigorously control Type I error conditional on the selected model path and yield uniformly calibrated p-values. Contribution/Results: We establish theoretical validity of the proposed procedures and demonstrate, via extensive simulations and real-data analyses, their improved statistical power and accurate p-value calibration—overcoming a key limitation of PSI in detecting heteroscedastic structural changes.
This paper addresses the lack of a unified framework for assessing systematic error (bias) across causal and descriptive inference. We propose the first cross-paradigm, generalizable bias risk assessment method, integrating modeling assumptions, data-generating mechanisms, and inferential objectives to cover high-risk settings—including randomized controlled trials (RCTs), nonprobability sampling, and statistical extrapolation—beyond traditional medical RCT constraints. Our approach combines qualitative bias mapping, assumption sensitivity analysis, and standardized reporting criteria, mandating explicit documentation of untestable assumptions and model uncertainty. The framework has been adopted as a mandatory reporting requirement by leading journals and funding agencies, thereby enhancing the reliability, interpretability, and external validity of research findings. (132 words)
This study addresses the uncontrolled false positive rates and misleading inferences arising from commonly used simulation methods in shift-share designs. We systematically evaluate prevailing inferential approaches in empirical research through a suite of multilevel simulation experiments. By comparing Monte Carlo analysis with counterfactual data-generating mechanisms, we uncover non-monotonic trade-offs among fidelity, sensitivity, and risk of misdirection across simulation designs. We propose a novel “progressive-fidelity simulation framework,” demonstrating that low-fidelity simulations suffice to expose fundamental inferential flaws, whereas high-fidelity simulations detect subtle, previously overlooked biases—substantially improving detection power. The framework balances interpretability and computational efficiency, offering a reproducible and scalable paradigm for assessing the robustness of causal inference methods.
This study addresses the computational complexity associated with calculating quantiles of the inverse normal distribution, Student’s t-distribution, and outlier rejection criteria in hypothesis testing. To overcome the reliance on table lookups or iterative numerical methods, the paper proposes concise and highly accurate analytical approximations formulated as closed-form expressions. These approximations significantly reduce computational overhead while maintaining precision sufficient for practical statistical applications. The resulting method offers substantial gains in computational efficiency, making it particularly well-suited for resource-constrained environments or scenarios requiring rapid statistical inference. By bridging theoretical rigor with practical utility, the approach delivers both methodological insight and real-world applicability.
This study addresses the common reliance on unrealistic assumptions about average treatment effects in experimental and observational research designs. It proposes a novel paradigm that shifts focus from directly positing average effects to modeling the full distribution of individual treatment effects, from which more plausible assumptions about average effects can be derived. By integrating distributional modeling with cross-disciplinary case studies, the approach demonstrates its validity and utility across diverse fields—including medicine, economics, and psychology—offering researchers a principled, heterogeneity-aware framework for specifying effect sizes grounded in empirical realism rather than idealized assumptions.
Traditional statistical inference relies on finite-dimensional modeling assumptions, which are often inadequate for uncertainty quantification in nonparametric function estimation. This work proposes a retrospective inference paradigm that centers on a point estimate derived from observed data and generates “estimation clones” to emulate repeated sampling. Inference is then conducted via the empirical distribution of these clones, without imposing probabilistic assumptions on the true parameter. The approach delivers robust uncertainty quantification in nonparametric regression while naturally accommodating parametric models, where it recovers classical inferential results. By integrating with smooth spline ANOVA models, the framework achieves both flexibility and theoretical rigor, offering a practical pathway for valid inference in nonparametric settings.
This study addresses the substantial bias often introduced in meta-analyses when estimating standard deviations solely from the five-number summary—specifically, the minimum, maximum, and median—due to insufficient information, which can compromise inferential reliability. To mitigate this issue, the authors propose a novel estimation method based on a scaled Beta distribution that incorporates data shape characteristics to improve accuracy. A comprehensive sensitivity analysis is systematically conducted to quantify estimation uncertainty. Through extensive simulation studies and real-data applications, the proposed approach demonstrates markedly superior performance over conventional estimators across a variety of underlying distributions. Additionally, the authors provide an interactive web tool to facilitate practical implementation, enabling researchers to readily assess and correct potential bias in standard deviation estimates, thereby enhancing the robustness of meta-analytic findings.