Score
Design, fit, and evaluate regression models that combine fixed effects and random effects to represent population-level relationships together with group- or series-specific deviations (for example, intercepts and slopes) in hierarchical or repeated-measures data. Use these models to pool information across groups, estimate population-level trends and group-specific parameters, and quantify uncertainty and variance components.
To address the challenges of jointly selecting fixed and random effects in multilevel functional regression—and the inability of existing methods to capture cluster-specific heterogeneity—this paper proposes MuFuMES, the first sparse simultaneous selection framework tailored for multilevel functional mixed-effects models. MuFuMES integrates spline basis expansion, a spike-and-slab group Lasso prior, and an expectation-conditional maximization (ECM) algorithm to enable efficient maximum a posteriori (MAP) estimation. Simulation studies demonstrate near-zero false positive and false negative rates. Applied to accelerometer data from NHANES 2011–2012, MuFuMES successfully identifies age- and race-specific diurnal activity pattern heterogeneities, uncovering biologically interpretable, hierarchical covariate effect structures.
To address inadequate calibration of both local (e.g., group-level random effects) and global parameters in multilevel Bayesian models under complex survey designs, this paper proposes an improved survey-weighted pseudo-posterior framework with an automated post-processing pipeline—enabling, for the first time, consistent uncertainty calibration for both parameter types within hierarchical structures. The method integrates mixed-effects modeling, weighted likelihood construction, and posterior reweighting calibration, ensuring asymptotic consistency theoretically and seamless compatibility with mainstream Bayesian software computationally. Simulation studies and empirical analysis using the National Survey on Drug Use and Health (NSDUH) demonstrate that the proposed approach substantially improves estimation accuracy and reliability of uncertainty quantification, particularly reducing bias in random-effects inference. The corresponding algorithms are publicly available and integrated into the R package `csSampling`, offering a scalable, reproducible solution for hierarchical Bayesian analysis under complex sampling.
Existing data integration methods struggle to accommodate complex survey designs and typically assume that multiple data sources originate from the same population, rendering them unsuitable for non-probability samples. This work proposes a model-assisted calibration framework that extends such integration to multiple probability survey samples—a first in the literature—accepting either individual-level data or aggregated summary statistics as input. The approach guarantees design-consistent estimation without requiring correct specification of the outcome model and naturally accommodates complex sampling designs. Coupled with Taylor linearization for variance estimation, the method substantially enhances the efficiency of regression analysis while preserving validity for finite-population inference. Simulation studies and empirical analyses using NHANES and NHIS data demonstrate consistent efficiency gains across diverse scenarios.
Accurate estimation and inference for the average treatment effect (ATE) in multi-covariate randomized controlled trials (RCTs) remain challenging, particularly under high-dimensional covariates and small sample sizes. Method: Building on the Neyman finite-population framework, we propose a bias-corrected regression-adjustment estimator with cross-fitting and introduce, for the first time, an HC3-type heteroskedasticity-robust standard error tailored to high-dimensional settings. We rigorously derive first- and second-order stochastic expansions of the random component of regression-adjustment estimators, identifying the source of higher-order bias in conventional inference; leveraging this insight, we design a cross-fitting procedure to eliminate bias and extend HC3 standard errors to stratified experimental designs. Results: Simulations and reanalysis of Angrist et al. (2009)’s education RCT demonstrate that our method substantially improves estimation accuracy, confidence interval coverage, and statistical power—especially in small-sample and high-dimensional scenarios.
This paper addresses the pervasive estimation bias in nonlinear network models with individual fixed effects—such as bilateral binary link formation models. We propose a systematic bias-correction method grounded in Jackknife resampling, the first to extend this technique to network fixed-effect frameworks. Our approach accommodates directed and undirected networks, non-binary outcomes, and higher-order dependence structures (e.g., triangles, quadruples). It delivers unbiased estimation of average treatment effects and counterfactual predictions, and is compatible with causal inference settings such as gravity models. Empirical application to country-level import-export data demonstrates substantially improved parameter consistency. Moreover, the method robustly identifies key network dependence features—including reciprocity and transitivity—without restrictive parametric assumptions. By integrating rigorous asymptotics with computational tractability, our framework provides a general, scalable, and theoretically grounded bias-correction tool for network data analysis.
Traditional heterogeneity analyses rely on pre-specified subgroups, limiting their ability to uncover complex effect modification mechanisms, while existing data-driven approaches often lack interpretability. This study proposes a novel paradigm grounded in causal transportability, treating population composition as a continuous variable to model the relationship between effect modifier distributions and the overall exposure effect. The framework enables estimation of intervention effects across diverse populations and identification of key vulnerability features. It integrates causal transportability, effect modifier selection—combining prior knowledge with data-adaptive strategies—and effect surface modeling, facilitating both effect attribution and ranking of modifier importance. Empirical application to child stunting and drought exposure successfully identifies critical modifiers, and an open-source Shiny interactive tool is provided to support broad adoption.
This study addresses the challenge that existing statistical methods struggle to effectively estimate average treatment effects in experiments involving both randomly assigned treatments and fixed covariates. The authors develop a unified theoretical framework that, for the first time, establishes a general estimating equation theory for misspecified linear regression models with mixed regressors—combining random treatment indicators and fixed covariates—and extends this framework to clustered data settings. By integrating estimating equations, misspecification-robust analysis, and causal inference techniques, the proposed approach yields valid causal interpretations of regression coefficients and their standard errors even under model misspecification. This methodology is broadly applicable to practical experimental designs, including completely randomized trials.
Longitudinal data often exhibit multiple sources of heterogeneity, including divergent mean trajectories, increasing residual variance over time, and occasional outlying measurements. Conventional homogeneous models may yield inefficient parameter estimates and inflated variance assessments in such settings. This work proposes a novel Bayesian mixture model that, for the first time, incorporates covariate-driven binary indicator variables within a unified Bayesian framework to jointly model these three forms of heterogeneity via logistic regression. Inference is carried out using Markov chain Monte Carlo (MCMC) methods, and the approach facilitates posterior-probability-based model selection to evaluate the necessity of each heterogeneous component. Simulation studies demonstrate that the proposed method accurately identifies underlying heterogeneity structures and yields efficient fixed-effect estimates. Its practical utility is further corroborated through application to DHEAS hormone data from the Study of Women’s Health Across the Nation (SWAN).
Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.
This study addresses the susceptibility of parametric models to misspecification and the lack of hypothesis testing with Type I error control in nonparametric Gaussian processes (GPs). We propose a hypothesis testing framework based on additive GPs with restricted pairwise interactions. Methodologically, we establish contraction rate theory for additive GPs to effectively mitigate dimensionality dependence. Technically, by integrating additive GPs, pairwise interaction modeling, and statistical inference, the framework enables controlled testing of nonlinear associations and interaction effects. Simulations demonstrate that the proposed approach substantially improves statistical power compared to standard GPs. An application in environmental health successfully identifies nonlinear effects of metal exposures and their demographic modifications, validating the practical utility of the framework.