ArXiv

Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data

Authors
Nina van Gerwen, Dimitris Rizopoulos, Manon Hillegers...
Categories
stat.ML, cs.LG
arXiv
https://arxiv.org/abs/2608.05930v1
PDF
https://arxiv.org/pdf/2608.05930v1

Brief

The paper develops the Deep Generalised Mixed Model (DGMM) to analyse hierarchical ESM longitudinal data (motivated by the GrowIt! adolescent COVID-19 study) where dropout induces missing-at-random patterns. DGMM extends mixed-effects modelling using neural networks, an adapted variational auto-encoder, and Bayesian data augmentation to scale to high-dimensional, generic-distribution outcomes. Results on GrowIt! and simulations show potential but limited reliability due to model instability. Summary is based on the abstract.

Why it matters

Van Gerwen et al. (Nina van Gerwen, Dimitris Rizopoulos, Manon Hillegers, Loes Keijsers, Sten Willemsen; arXiv:2608.05930v1, 2026-08-06) introduce the Deep Generalised Mixed Model (DGMM), a neural-network generalisation of mixed-effects models that combines a variational auto-encoder adaptation with Bayesian data augmentation to give valid inference under missing-at-random (MAR) dropout.

Key details

  • DGMM targets high-dimensional experience sampling method (ESM) longitudinal data (applied to the GrowIt! adolescent COVID-19 dataset) and supports generic outcome distributions; experiments and simulations show promise but authors report suboptimal performance caused by model instability.
Source evidence

Abstract

The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. Our work is motivated by such data collected by the GrowIt! app, which was released to investigate daily emotions among adolescents during the COVID-19 pandemic. Current procedures to analyse ESM data face various challenges. While standard statistical techniques may not scale well to a high-dimensional setting, machine learning procedures can give biased results due to selection bias introduced by missingness. In our motivating dataset, adolescents dropped out due to previous strong feelings of negative emotions. Hence, the implied missing data are of the missing-at-random type that standard machine learning procedures cannot accommodate. We develop a novel neural network architecture that generalises mixed effects models to deep learning to overcome these challenges. It allows semi-parametric and flexible modelling of data's mean and correlation structure through fixed and random effects. For estimation, we use an adaptation of variational auto-encoders and a Bayesian data augmentation algorithm. Through this approach, the model can accommodate longitudinal outcomes following generic distributions, scale well to high-dimensional settings and provide valid inference when data are missing-at-random. We applied the Deep Generalised Mixed Model to the GrowIt! study and various simulations. The results show potential for the Deep Generalised Mixed Model, yet suboptimal performance due to model instability.