multilevel regression analysis

Designs and fits multilevel (hierarchical) linear or linear mixed‑effects models that estimate relationships between predictors and an outcome while explicitly modelling nested or clustered data via level‑specific random intercepts and slopes. Uses these models to partition variance across levels, test and compare fixed effects between groups or contexts, handle repeated measures, and identify strongest predictors.

multilevelregressionanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Comparing multilevel and fixed effect approaches in the generalized linear model setting

Nov 04, 2024
HB
He Bai
🏛️ University of Massachusetts Amherst | Reed College | Grinnell College

This paper addresses bias and standard error misspecification in multilevel models (MLMs) for estimating treatment effects under generalized linear models (GLMs) when group-level confounding is present. We demonstrate that, unlike in linear models, MLMs are not equivalent to regularized fixed-effects (FE) estimators in GLMs, and their default standard errors typically underestimate within-group dependence-induced variability. To remedy this, we propose bias-corrected MLM (bcMLM), which reduces estimation bias relative to FE, as confirmed through simulations and empirical analysis. Furthermore, we show that cluster bootstrap inference substantially outperforms default standard errors in accounting for clustering structure. Our work provides a theoretically grounded, computationally feasible estimation framework for causal inference in nonlinear GLMs—filling critical gaps in both the theoretical understanding of bias mechanisms and the practical implementation of bias correction for MLMs in generalized settings.

Addressing bias in treatment coefficient estimates for GLMsComparing multilevel and fixed effect models in GLM settingsEvaluating standard error accuracy in non-linear GLM contexts

Assessing the impact of variance heterogeneity and misspecification in mixed-effects location-scale models

May 23, 2025
VJ
Vincent Jeanselme
🏛️ University of Cambridge | Columbia University | UCL

This study investigates the statistical performance of linear mixed models (LMMs) and mixed-effects location-scale models (MELSMs) under heteroscedasticity and model misspecification. Using large-scale longitudinal Monte Carlo simulations, we systematically evaluate estimation bias and confidence interval coverage. Key contributions are: (1) When heteroscedasticity is ignored, LMMs severely overestimate random-effect standard deviations and yield invalid confidence intervals; (2) We first demonstrate an asymmetric impact of misspecification in MELSMs: misspecification of the location component substantially biases scale-parameter estimates, whereas misspecification of the scale component does not affect consistency of location-parameter estimators; (3) We establish that MELSMs maintain robustness for location inference even under scale-structure misspecification, thereby providing a more reliable foundation for modeling non-normal outcomes, joint models, and survival analyses.

Assessing impact of model misspecification on variance estimatesEvaluating MELSM performance under misspecified scale assumptionsExamining bias in LMMs due to heteroscedasticity violations

This study addresses the bias and uncertainty in parameter estimation arising from missing not at random (MNAR) data in linear multilevel models. The authors propose a sensitivity analysis approach that jointly models the outcome variable and the dropout mechanism, incorporating an interpretable sensitivity parameter. Under assumptions weaker than missing at random, this method enables partial identification of model parameters and constructs corresponding uncertainty intervals. To the best of our knowledge, this is the first integration of sensitivity analysis with multilevel modeling. The validity of the proposed approach is demonstrated through simulation studies, and its practical utility is illustrated in an empirical analysis examining the effects of loneliness and physical activity on memory trajectories, yielding robust and more interpretable inferences.

missing not at randommultilevel modelspartial identification

Variable Selection for Fixed and Random Effects in Multilevel Functional Mixed Effects Models

May 08, 2025
RG
Rahul Ghosal
🏛️ University of South Carolina | Harvard University

To address the challenges of jointly selecting fixed and random effects in multilevel functional regression—and the inability of existing methods to capture cluster-specific heterogeneity—this paper proposes MuFuMES, the first sparse simultaneous selection framework tailored for multilevel functional mixed-effects models. MuFuMES integrates spline basis expansion, a spike-and-slab group Lasso prior, and an expectation-conditional maximization (ECM) algorithm to enable efficient maximum a posteriori (MAP) estimation. Simulation studies demonstrate near-zero false positive and false negative rates. Applied to accelerometer data from NHANES 2011–2012, MuFuMES successfully identifies age- and race-specific diurnal activity pattern heterogeneities, uncovering biologically interpretable, hierarchical covariate effect structures.

Addressing high-dimensional multilevel functional data with cluster-specific effectsIdentifying age and race-specific heterogeneity in physical activity patternsSimultaneous selection of fixed and random effects in multilevel functional regression

Uniform inference in linear mixed models

Jul 25, 2025
KO
Karl Oskar Ekvall
🏛️ University of Florida | Karolinska Institutet

In linear mixed models, conventional asymptotic inference fails when the random-effects covariance matrix approaches or lies on the parameter boundary—e.g., zero variances or correlation coefficients tending to ±1. To address this, we propose a unified, finite-sample distributional approximation method. Our approach quantifies the deviation between linear combinations of score functions and standard normality, integrating uniform approximation theory with high-dimensional statistical techniques to construct computationally tractable confidence regions. The method accommodates both cluster-independent and crossed random-effects structures and is valid under high-dimensional asymptotics—where the number of parameters grows jointly with the number of random effects. We establish theoretical guarantees: the resulting confidence regions achieve near-nominal coverage in finite samples and remain robust under boundary scenarios. Extensive simulations confirm superior small-sample performance and computational feasibility.

Address singular covariance matrices in random effects inferenceDevelop uniform parameter approximations for linear mixed modelsProvide finite-sample bounds for near-boundary variance cases

Latest Papers

What's happening recently
View more

Conventional approaches to deciding whether to use multilevel models based on point estimates of the intraclass correlation coefficient (ICC) ignore sampling uncertainty and lack a foundation in statistical inference. This study proposes the "Negligible Effect Significance Test" (NEST), which introduces equivalence testing into ICC assessment for the first time. By constructing an ICC pivotal quantity from the F-statistic and incorporating design effect considerations, NEST defines a threshold for a “negligible” ICC that quantifies an acceptable level of variance inflation. The method provides a statistically rigorous decision rule for determining when multilevel modeling is warranted, accompanied by an R implementation to help researchers evaluate whether clustering structures merit explicit modeling—thereby avoiding misjudgments arising from neglecting sampling variability.

equivalence testingintraclass correlation coefficientmultilevel modeling

Traditional epidemiological approaches struggle to disentangle hierarchical variation and structural disparities in cardiovascular disease (CVD) mortality across multiscale geographic units. This study proposes a reproducible multilevel statistical inference framework that integrates Gaussian, Poisson, and population-offset Poisson models to separate demographic effects from structural risk factors within a county-nested hierarchy. Fixed effects—including year, sex, race, PM2.5, and O₃—along with county-level random intercepts are incorporated, using data from Ohio and Pennsylvania between 1999 and 2020. Findings reveal that Pennsylvania experienced a steeper decline in age-standardized CVD mortality, though several CVD subtypes plateaued or rebounded after 2010. Black populations exhibited significantly elevated risk, and PM2.5 showed stronger associations with ischemic and hypertensive heart diseases. The inclusion of a population offset notably reduced unexplained variance, demonstrating the framework’s utility as a generalizable tool for environmental health assessment and health equity research.

cardiovascular mortalitycontextual heterogeneityhierarchical variation

This work proposes an efficient fitting approach for generalized linear mixed-effects models in which the response variable follows a non-Gaussian distribution and its mean is nonlinearly related to the linear predictor through a link function. Built within the lme4 framework, the method jointly estimates fixed effects, conditional modes of random effects, and their variance–covariance structure via penalized iteratively reweighted least squares. High-accuracy approximations to the likelihood are achieved using either Laplace approximation or adaptive Gauss–Hermite quadrature. The approach accommodates user-specified exponential family distributions and link functions, supporting common non-Gaussian responses such as binomial and Poisson outcomes. By balancing computational efficiency with modeling flexibility, the proposed method substantially extends the applicability of generalized linear mixed models while maintaining robust performance.

Generalized Linear Mixed ModelsLink functionsMaximum likelihood estimation

This study addresses the severe undercoverage of conventional confidence intervals in linear mixed models when variance parameters approach boundary values. We propose a novel method based on a modified profile score statistic derived from the restricted likelihood, which constructs confidence intervals by estimating nuisance parameters within an expanded parameter space and inverting the test statistic. This approach enables accurate inference in challenging boundary scenarios, such as near-zero random-effect variances or strong correlations. Simulation studies demonstrate that the proposed method maintains coverage rates stably near the nominal 95% level while reducing interval widths by 1.7 to 3.1 times compared to general-purpose inference procedures, alongside a 47-fold improvement in computational efficiency. The practical utility of this framework is further validated through its successful application to a longitudinal autism study.

boundary parametersconfidence intervalscoverage probability

Hot Scholars

AD

Adrian Dobra

Professor of Statistics, University of Washington
Graphical modelsBayesian StatisticsMultivariate Statistics
PT

Petter Törnberg

University of Amsterdam
Digital platformsAIHeterodox Computational Social Sciencecomplex systems
HQ

Haodong Qi

Malmo University
DemographyComputational Social ScienceEconomicsInternational Migration
HP

Hans-Peter Piepho

Biostatistics Unit, Institute of Crop Science, University of Hohenheim, Germany
BiometryBiometricsBiostatisticsStatistics