longitudinal data analysis

Designing studies and applying statistical methods for repeated-measures or time-series data to detect within-subject changes and trends over time, account for temporal dependence, and use model outputs for tracking outcomes across extended follow-up.

longitudinaldataanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

This study addresses the challenge of estimating time-varying treatment effects in mobile health micro-randomized trials, where repeated interventions influence both longitudinal outcomes and recurrent events. The authors propose an innovative joint longitudinal-survival modeling framework that, for the first time, incorporates time-varying treatment effects from micro-randomized trials into a unified joint model. By leveraging Bayesian inference, the approach flexibly captures dynamic associations among repeated treatments, multidimensional longitudinal markers, and recurrent event times, while accommodating diverse treatment effect mechanisms. Model selection is guided by information criteria, and the performance of the survival submodel is evaluated using calibration plots. Simulation studies and an analysis of a substance use micro-randomized trial demonstrate that the proposed method achieves excellent model fit and accurate estimation of treatment effects.

joint modellongitudinal outcomesmobile health

This study addresses the challenge of modeling time-varying correlations among multiple longitudinal variables that dynamically depend on individual-level covariates—a feature inadequately captured by existing methods. The authors propose the TiVAC model, which operates within a bivariate Gaussian framework and employs penalized splines in a semiparametric formulation to flexibly characterize the smooth, covariate-dependent evolution of correlations over time. Estimation is efficiently performed via penalized maximum likelihood using a Newton–Raphson algorithm. TiVAC is the first method to enable flexible modeling of covariate-driven time-varying correlation effects while providing simultaneous confidence bands for formal inference. Simulation studies demonstrate its superior performance across diverse scenarios. Application to data from 291 bipolar disorder patients reveals age-dependent heterogeneity in how sex and use of neurologic medications modulate the correlation between depression and anxiety symptoms.

correlation heterogeneitycovariate-dependent correlationmultivariate longitudinal studies

Existing latent growth curve models (LGCMs) struggle to disentangle the baseline trait effect from the dynamic time-varying effect of time-varying covariates (TVCs) on longitudinal outcomes, particularly failing to model how TVCs explain variability in random intercepts and slopes. To address this, we propose a systematic decomposition framework that partitions each TVC into: (1) a baseline effect predicting random intercepts and slopes; and (2) three distinct time-varying effects—interval-level slope, interval-level change, and deviation-from-baseline change. This is the first approach to explicitly model TVC contributions to growth parameter variability within the LGCM framework, supporting both linear and common nonlinear growth trajectories. Implementation is provided for OpenMx and Mplus 8. Simulation and empirical studies demonstrate unbiased parameter estimates, high precision, and confidence interval coverage near nominal levels. All code is publicly available to ensure reproducibility and facilitate extension.

Addresses limitation of latent growth curve models with time-varying covariatesDecomposes time-varying covariate effect into baseline and temporal componentsEnables unbiased estimation of covariate impact on longitudinal outcomes

This study addresses the unclear mechanisms through which time-varying covariates (TVCs) influence latent class trajectory heterogeneity in nonlinear growth mixture models (GMMs). We propose a novel TVC decoupling framework that decomposes each TVC into two distinct components: a baseline trait—capturing its effect on initial growth factors—and a time-specific state—directly influencing observed outcomes—while allowing heterogeneous effects across latent classes. Methodologically, we integrate GMM with an extended mixture-of-experts (MoE) architecture and the proposed TVC decomposition, validated via Monte Carlo simulations and empirical longitudinal analysis. Simulation results confirm unbiased parameter estimation and nominal coverage of confidence intervals. Empirically, we identify significant between-class heterogeneity in both baseline and dynamic effects of reading ability on mathematics achievement trajectories. To our knowledge, this is the first work to systematically disentangle baseline versus dynamic TVC effects and model their cross-class heterogeneity within nonlinear GMMs, thereby enhancing theoretical precision and empirical interpretability in attributing trajectory heterogeneity.

Examines heterogeneity in trait and state effects across latent subgroups using simulations and real data.Extends growth mixture models to incorporate time-varying covariate effects on nonlinear trajectories.Proposes decomposing time-varying covariates into trait and state features for analysis.

Latest Papers

What's happening recently
View more

This study addresses the challenge of estimating the causal effect of longitudinal treatment strategies on survival outcomes using electronic health records, where monitoring frequency of covariates varies across patients, variable types, and time, and may itself carry information about underlying health status—potentially biasing conventional causal inference methods. For the first time, monitoring indicators are formally treated as time-varying confounders. The authors integrate inverse probability weighting, G-computation, and longitudinal targeted maximum likelihood estimation (TMLE) to develop a unified framework for causal effect estimation under informative monitoring. This approach substantially reduces bias arising from ignoring the monitoring mechanism and extends the applicability of both static and dynamic treatment strategies. Simulations confirm its validity, and an application to real-world ICU data demonstrates its ability to accurately assess the impact of different mechanical ventilation initiation strategies on mortality.

causal inferenceelectronic health recordsinformative monitoring

A Statistical Framework for Understanding Causal Effects that Vary by Treatment Initiation Time in EHR-based Studies

Dec 22, 2025
LB
Luke Benz
🏛️ Harvard T.H. Chan School of Public Health | Harvard Pilgrim Health Care Institute | Kaiser Permanente Washington Health Research Institute | Kaiser Permanente Southern California | University of California San Francisco | University of Pennsylvania

In real-world EHR studies, heterogeneous patient enrollment times and evolving treatment technologies cause observed treatment effects to vary with calendar time—yet it remains challenging to disentangle whether such variation reflects genuine changes in causal efficacy or merely shifts in patient population composition. This paper proposes an identification framework for calendar-time-specific average treatment effects (t-ATE). We introduce a novel doubly robust time-varying effect projection coupled with marginal structural model (MSM) selection, and develop a standardized-analysis-based covariate shift attribution metric to decouple population drift from true therapeutic change. Applied to a Kaiser Permanente cohort (2005–2011) of severely obese patients undergoing bariatric surgery, our method precisely quantifies the time-varying comparative effectiveness of two surgical procedures versus non-surgical management. Results indicate that approximately 40% of the observed temporal variation in treatment effects stems from baseline covariate drift—not intrinsic efficacy decay.

Describes how and why treatment effects vary by initiation timeEstimates calendar-time specific average treatment effects in EHR studiesQuantifies covariate shift role in explaining observed effect changes

This study addresses the frequent challenges of convergence difficulties and inferential bias in mixed-effects models for repeated measures (MMRM) under small-sample settings, multiple time points, and uncertainty in covariance structure selection. The authors propose a novel modeling strategy that integrates an empirically bias-corrected covariance estimator with a heterogeneous autoregressive covariance structure. Through extensive Monte Carlo simulations and analysis of a diabetes clinical trial dataset, they systematically evaluate various MMRM configurations. Their findings indicate that the proposed approach achieves high convergence rates and near-nominal confidence interval coverage in moderate-to-large samples. In small-sample scenarios, a parsimonious model incorporating baseline covariates and treatment–time interaction terms substantially improves both convergence and estimation efficiency, offering a robust and practical modeling framework for longitudinal analysis in clinical trials.

longitudinal dataMMRMmodel convergence

This study addresses the bias arising in existing varying-coefficient models when the dependence between observation times and longitudinal outcomes is ignored. To handle informative observation times, the authors propose a novel approach that models the observation mechanism via a proportional intensity model and, for the first time, integrates inverse intensity weighting with sieve estimation within a weighted least squares framework to obtain closed-form estimates of coefficient functions. Theoretical analysis establishes that the resulting estimator is consistent, achieves the optimal convergence rate, and is asymptotically normal, thereby enabling pointwise confidence interval construction. Simulation studies demonstrate that the proposed method substantially outperforms conventional unweighted approaches under informative observation schemes, and its practical utility is illustrated through an application to neuroimaging data from Alzheimer’s disease patients.

biased estimationinformative observation timesinvalid inference

This study addresses the limitation of traditional regression adjustment methods, which can only estimate average treatment effects and fail to capture the dynamic evolution of treatment effects over time. The authors propose a novel longitudinal treatment effect estimation framework that, for the first time, incorporates time-varying covariate transition mechanisms into regression adjustment. By modeling post-treatment covariate trajectories through transition kernels, the method enables a fine-grained characterization of effect heterogeneity across time. The proposed estimator is theoretically shown to achieve the semiparametric efficiency bound and possesses asymptotic normality. Both simulation studies and empirical analysis using A/B test data from a Japanese streaming platform demonstrate that the approach substantially reduces estimation variance and enhances statistical inference efficiency.

dynamic trajectoriesintermediate outcomeslongitudinal treatment effects

Hot Scholars

HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
EC

Erjia Cui

Division of Biostatistics and Health Data Science, University of Minnesota
Biostatistics
KS

Koustuv Saha

University of Illinois Urbana-Champaign
Computational Social ScienceSocial ComputingHuman-Centered Machine LearningWellbeing
XS

Xu Shi

University of Michigan
Electronic Health RecordCausal InferenceNegative ControlMachine Translation
LP

Layla Parast

University of Texas at Austin
Surrogate markersrisk predictionsurvival analysiscausal inference