Score
Design and implement statistical estimators and software that infer a time-varying reproduction number r_t from incidence time series using penalized-likelihood approaches (e.g., spline bases and roughness/regularization penalties). This includes providing frequentist uncertainty quantification (confidence intervals, profile-likelihood or bootstrap), selecting/tuning smoothing penalties, and engineering computational optimizations for fast, near real-time execution (including software implementations such as convrt).
Accurate real-time estimation of the effective reproduction number $R_t$ is hindered by reporting delays and the computational inefficiency and prior sensitivity of conventional Bayesian approaches. This work proposes ConvRt, a fast frequentist method that decouples smooth inference of $R_t$ from future forecasting by deconvolving the latent infection process and employing spline basis functions with penalized likelihood for smooth modeling. By avoiding reliance on Bayesian priors, ConvRt directly infers the dynamic trajectory of $R_t$ from observed data. Evaluations on both simulated and real-world epidemic data demonstrate that ConvRt achieves higher point estimation accuracy, provides reliable uncertainty quantification, and offers substantially improved computational speed compared to existing methods.
This work addresses the lack of reliable uncertainty quantification in online trend estimation for nonstationary time series by proposing a general online bootstrap method applicable to trend estimators expressed as time-window-weighted sample means—such as exponential smoothing and moving averages. Built upon asymptotic theory, the method provides the first uniform-in-time coverage guarantees for trend inference under nonstationarity, enabling adaptive anomaly detection and A/B testing in streaming data settings. Empirical evaluations demonstrate that the framework achieves well-calibrated uncertainty estimates and scales effectively across diverse nonstationary scenarios, offering a practical and real-time solution for accurate uncertainty quantification in large-scale online time series analysis.
This study addresses the issue that neglecting within-curve dependence of residuals in function-on-function regression leads to severely inadequate confidence interval coverage and estimation bias. To overcome this, it proposes a dependent residual inference framework integrating REML fitting with leverage-adjusted curve-clustered sandwich covariance (CL2) estimation, incorporating Satterthwaite degrees-of-freedom correction, and compares block neighborhood cross-validation (NCV) smoothing strategies to optimize estimation accuracy. The work reveals the undersmoothing tendency of REML and establishes distinct optimal strategies for shape description versus statistical inference. Empirically, CL2 elevates interval coverage to near-nominal levels, while NCV reduces coefficient surface errors by two- to fourfold, providing evidence-based recommendations for default parameters in the R package refund.
Real-time estimation of the time-varying effective reproduction number (Rₜ) during the COVID-19 pandemic suffers from the absence of ground-truth labels and heavy reliance on empirical hyperparameter tuning. Method: We propose a ground-truth-free, data-driven framework based on a time-varying autoregressive Poisson (TV-AR-Poisson) model, integrating penalized likelihood estimation with a novel Stein’s Unbiased Risk Estimate (SURE) criterion. We extend Stein’s lemma to discrete-time Poisson settings and design a weekly-scale model to jointly account for infection stochasticity and reporting delay robustness. Contribution/Results: We theoretically establish the asymptotic unbiasedness of the SURE estimator. Extensive experiments on synthetic data validate its statistical efficacy, while application to real-world COVID-19 incidence data across multiple countries yields highly consistent, low-latency Rₜ trajectories. The framework significantly enhances model adaptability and practical deployability without requiring manual calibration or external validation labels.
Standard likelihood-based inference for ARMA models frequently converges to local optima, resulting in biased parameter estimates and undercoverage of confidence intervals; existing approaches lack a general-purpose remedy. This paper proposes a structural-aware stochastic initialization optimization algorithm and a profile likelihood confidence interval method. We introduce the first initialization strategy explicitly tailored to the geometric structure of the ARMA likelihood surface. Furthermore, we rigorously establish and generalize the superiority of the profile likelihood approach—demonstrating its substantially improved coverage accuracy relative to conventional Fisher information–based methods. Through extensive Monte Carlo simulations and empirical analyses, our methods markedly enhance estimation consistency and restore confidence interval coverage toward nominal levels. The proposed framework effectively mitigates long-standing inferential biases in both industrial applications and scientific research.
研究使用wild bootstrap和Efron's bootstrap方法解决Cox回归Lasso变量选择后的可靠系数推断问题。
This study addresses the challenges of estimating causal effects of time-varying interventions on rare survival outcomes in large-scale longitudinal observational studies, where high computational costs and severe class imbalance often hinder reliable inference. The authors propose a subsampling and inverse probability reweighting framework tailored for longitudinal survival data, which integrates seamlessly with existing causal estimators—such as g-formula–based ICE—while preserving estimator consistency and substantially reducing computational burden. This approach represents the first application of a subsampling strategy that jointly optimizes computational efficiency and statistical consistency in the context of causal inference for longitudinal rare events, effectively mitigating model instability induced by outcome imbalance. Simulations and an empirical analysis using electronic health records to assess the impact of social-behavioral factors on suicide risk demonstrate that the method markedly improves computational efficiency while enhancing both the stability and accuracy of causal estimates.
This study addresses the challenge of conducting inference in settings where the marginal distributions of multivariate responses are difficult to specify accurately—such as in longitudinal data or high-dimensional heteroscedastic regression—under the assumption that the conditional mean model is correctly specified. The authors develop a √n-consistent estimator via a penalized estimating equation and propose a hypothesis testing procedure focused on low-dimensional subvectors of parameters. A key innovation is the introduction of a cross-fitted covariance calibration mechanism, which substantially reduces the sensitivity of the test to misspecification of the nuisance covariance function. The resulting test statistic follows an asymptotic chi-squared distribution, achieving robustness by controlling Type I error while significantly enhancing statistical power for efficient and reliable inference.
Traditional statistical inference relies on finite-dimensional modeling assumptions, which are often inadequate for uncertainty quantification in nonparametric function estimation. This work proposes a retrospective inference paradigm that centers on a point estimate derived from observed data and generates “estimation clones” to emulate repeated sampling. Inference is then conducted via the empirical distribution of these clones, without imposing probabilistic assumptions on the true parameter. The approach delivers robust uncertainty quantification in nonparametric regression while naturally accommodating parametric models, where it recovers classical inferential results. By integrating with smooth spline ANOVA models, the framework achieves both flexibility and theoretical rigor, offering a practical pathway for valid inference in nonparametric settings.
This study addresses the challenge that data-driven variable selection methods—such as Lasso and its adaptive variants—undermine the validity of classical statistical inference in Cox proportional hazards models, particularly leading to inflated false positive rates in right-censored survival data. The authors systematically evaluate the post-selection inference performance of sample splitting, exact post-selection inference, and debiased Lasso approaches. For the first time, these methods are comprehensively compared within a simulation framework designed to closely mimic real-world biomedical scenarios, and their practical utility is further validated using publicly available datasets. The findings elucidate the trade-offs among these methods in controlling Type I error rates and achieving estimation accuracy, thereby offering reliable and practical inference strategies for high-dimensional survival data analysis.