🤖 AI Summary
This study addresses the challenge of non-random missingness—such as dropout driven by negative affect—in high-dimensional experience sampling method (ESM) data, which can introduce bias in conventional approaches and standard machine learning models. The authors propose a novel neural network architecture that, for the first time, extends generalized linear mixed-effects models into a deep learning framework, jointly modeling fixed and random effects to flexibly capture both the mean structure and within-subject correlations in longitudinal data. By integrating variational autoencoders with Bayesian data augmentation, the method enables semi-parametric modeling and robust inference under general distributional assumptions and arbitrary missingness mechanisms. Empirical evaluations on the GrowIt! study and simulation experiments demonstrate its potential, though further improvements in model stability are needed to enhance practical performance.
📝 Abstract
The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. Our work is motivated by such data collected by the GrowIt! app, which was released to investigate daily emotions among adolescents during the COVID-19 pandemic. Current procedures to analyse ESM data face various challenges. While standard statistical techniques may not scale well to a high-dimensional setting, machine learning procedures can give biased results due to selection bias introduced by missingness. In our motivating dataset, adolescents dropped out due to previous strong feelings of negative emotions. Hence, the implied missing data are of the missing-at-random type that standard machine learning procedures cannot accommodate. We develop a novel neural network architecture that generalises mixed effects models to deep learning to overcome these challenges. It allows semi-parametric and flexible modelling of data's mean and correlation structure through fixed and random effects. For estimation, we use an adaptation of variational auto-encoders and a Bayesian data augmentation algorithm. Through this approach, the model can accommodate longitudinal outcomes following generic distributions, scale well to high-dimensional settings and provide valid inference when data are missing-at-random. We applied the Deep Generalised Mixed Model to the GrowIt! study and various simulations. The results show potential for the Deep Generalised Mixed Model, yet suboptimal performance due to model instability.