statistical analysis

Designs and implements quantitative analyses of observed data to estimate parameters, test hypotheses, and quantify uncertainty to support reproducible inference. This includes selecting and applying appropriate statistical models and tests, checking assumptions and diagnostics, and producing summary estimates, visualizations, and uncertainty measures (e.g., confidence intervals, p-values, effect sizes).

statisticalanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.

analysis reasoningassumptionsdata analysis

Quantifying Uncertainty: All We Need is the Bootstrap?

Mar 29, 2024
UZ
Urvsa Zrimvsek
🏛️ University of Ljubljana

This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.

Assessing bootstrap's potential to simplify statistical education and practiceComparing double bootstrap performance against traditional confidence interval techniquesEvaluating bootstrap as universal alternative for uncertainty quantification methods

Post-selection inference for quantifying uncertainty in changes in variance

May 24, 2024
RC
Rachel Carrington
🏛️ Lancaster University

Classical variance change-point detection methods suffer from p-value bias and inflated Type I error due to data reuse in model selection. Existing post-selection inference (PSI) frameworks are restricted to mean-shift detection and do not extend to variance changes. Method: This paper introduces the first PSI framework for variance change-point detection, proposing two general-purpose constructions for post-selection p-values compatible with diverse algorithms (e.g., piecewise constant modeling) and test forms (e.g., constrained likelihood ratio tests). Leveraging conditional inference, convex optimization, and statistical functional theory, the methods rigorously control Type I error conditional on the selected model path and yield uniformly calibrated p-values. Contribution/Results: We establish theoretical validity of the proposed procedures and demonstrate, via extensive simulations and real-data analyses, their improved statistical power and accurate p-value calibration—overcoming a key limitation of PSI in detecting heteroscedastic structural changes.

Avoiding bias in testing post-selection changepointsExtending post-selection inference to variance changesQuantifying uncertainty in detected variance changepoints

Towards a unified approach to formal risk of bias assessments for causal and descriptive inference

Aug 22, 2023
OP
O. Pescott
🏛️ UK Centre for Ecology & Hydrology | University of Newcastle

This paper addresses the lack of a unified framework for assessing systematic error (bias) across causal and descriptive inference. We propose the first cross-paradigm, generalizable bias risk assessment method, integrating modeling assumptions, data-generating mechanisms, and inferential objectives to cover high-risk settings—including randomized controlled trials (RCTs), nonprobability sampling, and statistical extrapolation—beyond traditional medical RCT constraints. Our approach combines qualitative bias mapping, assumption sensitivity analysis, and standardized reporting criteria, mandating explicit documentation of untestable assumptions and model uncertainty. The framework has been adopted as a mandatory reporting requirement by leading journals and funding agencies, thereby enhancing the reliability, interpretability, and external validity of research findings. (132 words)

Addressing invisible uncertainty and systematic errors in statistical modelsMandating bias reporting to clarify research limitations and applicabilityUnified framework for assessing bias in causal and descriptive inference

Assessing Inference Methods

Dec 18, 2019
BF
Bruno Ferman
🏛️ Sao Paulo School of Economics - FGV

This study addresses the uncontrolled false positive rates and misleading inferences arising from commonly used simulation methods in shift-share designs. We systematically evaluate prevailing inferential approaches in empirical research through a suite of multilevel simulation experiments. By comparing Monte Carlo analysis with counterfactual data-generating mechanisms, we uncover non-monotonic trade-offs among fidelity, sensitivity, and risk of misdirection across simulation designs. We propose a novel “progressive-fidelity simulation framework,” demonstrating that low-fidelity simulations suffice to expose fundamental inferential flaws, whereas high-fidelity simulations detect subtle, previously overlooked biases—substantially improving detection power. The framework balances interpretability and computational efficiency, offering a reproducible and scalable paradigm for assessing the robustness of causal inference methods.

Analyzing trade-offs in simulation-based inference assessmentsEvaluating reliability of inference methods for false-positive controlProposing alternatives to misleading shift-share design evaluations

Latest Papers

What's happening recently
View more

This study addresses the computational complexity associated with calculating quantiles of the inverse normal distribution, Student’s t-distribution, and outlier rejection criteria in hypothesis testing. To overcome the reliance on table lookups or iterative numerical methods, the paper proposes concise and highly accurate analytical approximations formulated as closed-form expressions. These approximations significantly reduce computational overhead while maintaining precision sufficient for practical statistical applications. The resulting method offers substantial gains in computational efficiency, making it particularly well-suited for resource-constrained environments or scenarios requiring rapid statistical inference. By bridging theoretical rigor with practical utility, the approach delivers both methodological insight and real-world applicability.

computational simplificationhypothesis testingoutlier rejection

This study addresses the common reliance on unrealistic assumptions about average treatment effects in experimental and observational research designs. It proposes a novel paradigm that shifts focus from directly positing average effects to modeling the full distribution of individual treatment effects, from which more plausible assumptions about average effects can be derived. By integrating distributional modeling with cross-disciplinary case studies, the approach demonstrates its validity and utility across diverse fields—including medicine, economics, and psychology—offering researchers a principled, heterogeneity-aware framework for specifying effect sizes grounded in empirical realism rather than idealized assumptions.

average treatment effecteffect sizeexperimental design

Traditional statistical inference relies on finite-dimensional modeling assumptions, which are often inadequate for uncertainty quantification in nonparametric function estimation. This work proposes a retrospective inference paradigm that centers on a point estimate derived from observed data and generates “estimation clones” to emulate repeated sampling. Inference is then conducted via the empirical distribution of these clones, without imposing probabilistic assumptions on the true parameter. The approach delivers robust uncertainty quantification in nonparametric regression while naturally accommodating parametric models, where it recovers classical inferential results. By integrating with smooth spline ANOVA models, the framework achieves both flexibility and theoretical rigor, offering a practical pathway for valid inference in nonparametric settings.

model assumptionsnonparametric estimationretrospective inference

This study addresses the substantial bias often introduced in meta-analyses when estimating standard deviations solely from the five-number summary—specifically, the minimum, maximum, and median—due to insufficient information, which can compromise inferential reliability. To mitigate this issue, the authors propose a novel estimation method based on a scaled Beta distribution that incorporates data shape characteristics to improve accuracy. A comprehensive sensitivity analysis is systematically conducted to quantify estimation uncertainty. Through extensive simulation studies and real-data applications, the proposed approach demonstrates markedly superior performance over conventional estimators across a variety of underlying distributions. Additionally, the authors provide an interactive web tool to facilitate practical implementation, enabling researchers to readily assess and correct potential bias in standard deviation estimates, thereby enhancing the robustness of meta-analytic findings.

data shapemeta-analysissensitivity analysis

Hot Scholars

YK

Youngrae Kim

University of Southern California
Machine LearningComputer VisionDomain Adaptation