fit gee models

Fit generalized estimating equation (GEE) models to estimate population‑averaged associations from correlated or clustered observations by specifying a working correlation structure and robust (sandwich) standard errors. Apply the Mundlak correction by including cluster/group means of time-varying predictors to separate within‑ and between‑cluster effects and to control for cluster‑level confounds.

fitgeemodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the substantial bias in generalized estimating equations (GEE) when the number of independent clusters is small. Viewing GEE as an M-estimator for clustered data, the authors derive a corrected estimating equation that achieves first-order bias reduction while accounting for the dependence of the working covariance on mean parameters. Innovatively, they formulate the bias-corrected GEE estimator under a pairwise odds-ratio parameterization, circumventing the stringent compatibility constraints imposed by traditional correlation-based parameterizations on marginal means and thereby enhancing suitability for small-sample settings. Six new estimators implemented in the R package geer demonstrate markedly reduced bias across various scenarios, while preserving efficiency and confidence interval coverage comparable to standard GEE. The practical utility of the proposed approach is further validated through application to clinical trial data analysis.

bias reductioncorrelated datageneralized estimating equations

Generalized Estimating Equations for Hearing Loss Data with Specified Correlation Structures

Jun 28, 2023
ZW
Zhuoran Wei
🏛️ Harvard T.H. Chan School of Public Health | East China Normal University | Harvard Medical School

Conventional generalized estimating equations (GEE) suffer from low estimation efficiency in small-to-moderate samples when analyzing pure-tone audiometry data, where complex within-cluster correlation structures—particularly interaural dependence—violate standard working correlation assumptions. Method: This paper introduces a novel second-order GEE framework that jointly estimates regression coefficients and correlation structure parameters—specifically modeling interaural correlation parametrically—thereby improving statistical efficiency for ear-level covariate effects, especially under moderate-to-strong within-cluster dependence. Contribution/Results: Simulation studies and analysis of real data from the Conservation of Hearing Study demonstrate that the proposed method achieves superior efficiency gains for ear-level exposure effect estimation compared to independence-, exchangeable-, and unstructured-GEE approaches. It robustly identifies a statistically significant association between dietary adherence and hearing loss. The core innovation lies in establishing an estimable and interpretable correlation structure modeling framework, advancing precise analysis of high-dimensional repeated-measures data in auditory epidemiology.

Assessing dietary impact on hearing via advanced GEEImproving GEE efficiency for within-cluster covariatesModeling complex correlation in hearing loss data

This study addresses the instability of conventional generalized estimating equations (GEE) in separation scenarios—such as small samples, sparse data, or rare events—where non-convergence or extreme estimates commonly occur, particularly under non-independent working correlation structures. To overcome these limitations, the authors propose a penalized GEE framework that integrates Jeffreys’ prior penalty with a marginalized odds ratio parameterization, effectively circumventing the breakdown of traditional correlation parameters at extreme probabilities and guaranteeing finite estimates. The approach accommodates multiple link functions, including logit, probit, clog-log, and cauchit, and introduces both a one-step algorithm (OPGEE) and a hybrid variant (HPGEE) to enhance computational efficiency. Simulation studies and an analysis of a respiratory disease trial demonstrate that the proposed method substantially outperforms standard GEE under separation while maintaining comparable performance in regular settings. The methodology is implemented in the R package geer.

correlated binary datafinite estimationgeneralized estimating equations

This study addresses bias in marginal population parameter estimation for non-normal, bivariate correlated data—particularly in longitudinal settings. We systematically compare the performance of generalized joint regression models (GJRM), generalized linear mixed models (GLMM), and generalized estimating equations (GEE). Through Monte Carlo simulations and analysis of real-world physician visit data, we first demonstrate that GLMM yields substantial bias in marginal parameter estimates under non-identity link functions and skewed response distributions. In contrast, GJRM—when the copula is correctly specified—exhibits unbiased marginal estimation, robust standard errors, and superior model fit. GJRM maintains consistency of marginal estimates and validity of statistical inference across diverse non-normal distributions (e.g., Poisson, negative binomial, Beta), outperforming GLMM, GEE, and generalized linear models (GLM). These findings establish GJRM as a more reliable methodological choice for analyzing such complex correlated data.

Compares copula-based GJRM with GLMM and GEE for bivariate correlated dataDemonstrates GJRM's superior accuracy and flexibility in modeling skewed distributionsIdentifies GLMM bias in non-normal data with non-identity link functions

This study addresses the limitations of traditional fixed-effects models, which rely on strong assumptions of linear additivity and independence, thereby struggling to accommodate group heterogeneity and within-group dependence and leading to biased cross-group comparisons. To overcome these issues, the authors propose a Graph Neural Network–based Generalized Mundlak Estimator (GME-GNN) that dispenses with conventional intercept terms and instead incorporates group-level balancing statistics to control for between-group confounding. By leveraging the message-passing mechanism of graph neural networks, the method adaptively learns nonlinear representations to flexibly capture intra-group interaction structures. The estimator is theoretically shown to possess double robustness and asymptotic normality. Both simulation experiments and empirical analyses demonstrate its superior performance over existing approaches in bias reduction and cross-group inference.

cross-group comparisongroup heterogeneityMundlak estimator

Latest Papers

What's happening recently
View more

This study addresses the sensitivity of causal effect estimation to model misspecification in longitudinal cluster-randomized and quasi-experimental designs. Within an M-estimation framework, it demonstrates that fixed-effects models yield consistent and asymptotically normal estimates of nonparametrically defined treatment effects, provided the treatment effect structure is correctly specified—even when other model components are arbitrarily misspecified. The work establishes, for the first time, that fixed-effects models are valid for estimating superpopulation marginal effects and reveals their robustness to partial misspecification of the treatment effect structure across diverse longitudinal settings. Through theoretical analysis, simulations, and reanalyses of empirical data, the paper further shows that fixed-effects models outperform mixed-effects models in robustness and reliability when time-invariant confounding exists at the cluster or individual level.

causal inferencefixed-effects modelslongitudinal cluster trials

This work addresses key limitations of traditional Bayesian profile regression in handling hierarchical or longitudinal data, its inability to effectively model interactions between latent profile clusters and covariates, and its reduced statistical efficiency under highly correlated covariate structures. To overcome these challenges, the authors integrate generalized linear mixed models (GLMMs) into the Bayesian profile regression framework, incorporating random effects to accommodate multilevel and longitudinal designs. The latent clusters derived from profile clustering are explicitly included as predictors in the outcome model, enabling direct modeling of cluster–covariate interactions. This approach substantially extends the applicability of profile regression and enhances both modeling efficiency and predictive performance for complex dependency structures. An open-source R package implementing this method combines Bayesian inference, GLMMs, and profile clustering, supporting continuous or binary outcomes and mixed-type covariates, with broad utility in epidemiology, social sciences, and clinical research.

Bayesian profile regressioncovariate interactionsGeneralised Linear Mixed Models

This study addresses the bias in causal effect estimation arising from the coexistence of unmeasured cluster-level confounding and treatment effect heterogeneity in observational clustered data. To tackle this dual challenge, the authors propose an intra-group g-computation approach: clusters are first stratified by observed treatment prevalence, within-stratum g-computation is implemented using random-effects models, and estimates across strata are then aggregated to correct for bias. This method innovatively embeds random-effects modeling within the g-computation framework, effectively mitigating both sources of bias. Simulation studies demonstrate that the proposed estimator achieves the lowest root mean squared error when unmeasured confounding and heterogeneity co-occur. Applied to data from Bangladesh, the method reveals that adolescent pregnancy is associated with an average reduction of 0.12 in child height-for-age Z-scores (95% CI: [–0.18, –0.06]).

causal effect estimationhierarchical dataobservational studies

Hot Scholars

NL

Nikola Ljubešić

Researcher at Jožef Stefan Institute
natural language processingcomputational linguisticscomputational social science
BW

Bingkai Wang

University of Michigan
Clinical trialscausal inferencestatistics
JZ

Jia Zhou

Chongqing University
Human-Computer InteractionOlder Adults and ICTHuman Factors and Ergonomics
GL

Gabriel Loewinger

Machine Learning Research Scientist at National Institute of Mental Health
statisticsmachine learningapplied optimizationneuroscience