Score
Fit generalized estimating equation (GEE) models to estimate population‑averaged associations from correlated or clustered observations by specifying a working correlation structure and robust (sandwich) standard errors. Apply the Mundlak correction by including cluster/group means of time-varying predictors to separate within‑ and between‑cluster effects and to control for cluster‑level confounds.
This study addresses the substantial bias in generalized estimating equations (GEE) when the number of independent clusters is small. Viewing GEE as an M-estimator for clustered data, the authors derive a corrected estimating equation that achieves first-order bias reduction while accounting for the dependence of the working covariance on mean parameters. Innovatively, they formulate the bias-corrected GEE estimator under a pairwise odds-ratio parameterization, circumventing the stringent compatibility constraints imposed by traditional correlation-based parameterizations on marginal means and thereby enhancing suitability for small-sample settings. Six new estimators implemented in the R package geer demonstrate markedly reduced bias across various scenarios, while preserving efficiency and confidence interval coverage comparable to standard GEE. The practical utility of the proposed approach is further validated through application to clinical trial data analysis.
Conventional generalized estimating equations (GEE) suffer from low estimation efficiency in small-to-moderate samples when analyzing pure-tone audiometry data, where complex within-cluster correlation structures—particularly interaural dependence—violate standard working correlation assumptions. Method: This paper introduces a novel second-order GEE framework that jointly estimates regression coefficients and correlation structure parameters—specifically modeling interaural correlation parametrically—thereby improving statistical efficiency for ear-level covariate effects, especially under moderate-to-strong within-cluster dependence. Contribution/Results: Simulation studies and analysis of real data from the Conservation of Hearing Study demonstrate that the proposed method achieves superior efficiency gains for ear-level exposure effect estimation compared to independence-, exchangeable-, and unstructured-GEE approaches. It robustly identifies a statistically significant association between dietary adherence and hearing loss. The core innovation lies in establishing an estimable and interpretable correlation structure modeling framework, advancing precise analysis of high-dimensional repeated-measures data in auditory epidemiology.
This study addresses the instability of conventional generalized estimating equations (GEE) in separation scenarios—such as small samples, sparse data, or rare events—where non-convergence or extreme estimates commonly occur, particularly under non-independent working correlation structures. To overcome these limitations, the authors propose a penalized GEE framework that integrates Jeffreys’ prior penalty with a marginalized odds ratio parameterization, effectively circumventing the breakdown of traditional correlation parameters at extreme probabilities and guaranteeing finite estimates. The approach accommodates multiple link functions, including logit, probit, clog-log, and cauchit, and introduces both a one-step algorithm (OPGEE) and a hybrid variant (HPGEE) to enhance computational efficiency. Simulation studies and an analysis of a respiratory disease trial demonstrate that the proposed method substantially outperforms standard GEE under separation while maintaining comparable performance in regular settings. The methodology is implemented in the R package geer.
This study addresses bias in marginal population parameter estimation for non-normal, bivariate correlated data—particularly in longitudinal settings. We systematically compare the performance of generalized joint regression models (GJRM), generalized linear mixed models (GLMM), and generalized estimating equations (GEE). Through Monte Carlo simulations and analysis of real-world physician visit data, we first demonstrate that GLMM yields substantial bias in marginal parameter estimates under non-identity link functions and skewed response distributions. In contrast, GJRM—when the copula is correctly specified—exhibits unbiased marginal estimation, robust standard errors, and superior model fit. GJRM maintains consistency of marginal estimates and validity of statistical inference across diverse non-normal distributions (e.g., Poisson, negative binomial, Beta), outperforming GLMM, GEE, and generalized linear models (GLM). These findings establish GJRM as a more reliable methodological choice for analyzing such complex correlated data.
This study addresses the limitations of traditional fixed-effects models, which rely on strong assumptions of linear additivity and independence, thereby struggling to accommodate group heterogeneity and within-group dependence and leading to biased cross-group comparisons. To overcome these issues, the authors propose a Graph Neural Network–based Generalized Mundlak Estimator (GME-GNN) that dispenses with conventional intercept terms and instead incorporates group-level balancing statistics to control for between-group confounding. By leveraging the message-passing mechanism of graph neural networks, the method adaptively learns nonlinear representations to flexibly capture intra-group interaction structures. The estimator is theoretically shown to possess double robustness and asymptotic normality. Both simulation experiments and empirical analyses demonstrate its superior performance over existing approaches in bias reduction and cross-group inference.
本文针对具有信息性集群大小的集群随机试验中治疗效果估计问题,提出了一种简单且稳健的方法:通过饱和集群大小模型与g-计算来准确估计个体和集群平均治疗效果。
This study addresses the sensitivity of causal effect estimation to model misspecification in longitudinal cluster-randomized and quasi-experimental designs. Within an M-estimation framework, it demonstrates that fixed-effects models yield consistent and asymptotically normal estimates of nonparametrically defined treatment effects, provided the treatment effect structure is correctly specified—even when other model components are arbitrarily misspecified. The work establishes, for the first time, that fixed-effects models are valid for estimating superpopulation marginal effects and reveals their robustness to partial misspecification of the treatment effect structure across diverse longitudinal settings. Through theoretical analysis, simulations, and reanalyses of empirical data, the paper further shows that fixed-effects models outperform mixed-effects models in robustness and reliability when time-invariant confounding exists at the cluster or individual level.
本文提出了一种有限混合广义估计方程(MixGEE)方法,用于解决多变量相关结果的聚类问题,特别适用于生态学生物区域划分。
This work addresses key limitations of traditional Bayesian profile regression in handling hierarchical or longitudinal data, its inability to effectively model interactions between latent profile clusters and covariates, and its reduced statistical efficiency under highly correlated covariate structures. To overcome these challenges, the authors integrate generalized linear mixed models (GLMMs) into the Bayesian profile regression framework, incorporating random effects to accommodate multilevel and longitudinal designs. The latent clusters derived from profile clustering are explicitly included as predictors in the outcome model, enabling direct modeling of cluster–covariate interactions. This approach substantially extends the applicability of profile regression and enhances both modeling efficiency and predictive performance for complex dependency structures. An open-source R package implementing this method combines Bayesian inference, GLMMs, and profile clustering, supporting continuous or binary outcomes and mixed-type covariates, with broad utility in epidemiology, social sciences, and clinical research.
This study addresses the bias in causal effect estimation arising from the coexistence of unmeasured cluster-level confounding and treatment effect heterogeneity in observational clustered data. To tackle this dual challenge, the authors propose an intra-group g-computation approach: clusters are first stratified by observed treatment prevalence, within-stratum g-computation is implemented using random-effects models, and estimates across strata are then aggregated to correct for bias. This method innovatively embeds random-effects modeling within the g-computation framework, effectively mitigating both sources of bias. Simulation studies demonstrate that the proposed estimator achieves the lowest root mean squared error when unmeasured confounding and heterogeneity co-occur. Applied to data from Bangladesh, the method reveals that adolescent pregnancy is associated with an average reduction of 0.12 in child height-for-age Z-scores (95% CI: [–0.18, –0.06]).