🤖 AI Summary
This study addresses the unexpectedly suboptimal performance of complex text models in predicting adult depressive symptoms from childhood essays. To investigate this, we systematically compared six covariate-based baselines, such as logistic regression, against various Transformer architectures and zero-shot large language models for long-horizon depression prediction, with performance evaluated via ROC curves. Our findings reveal that traditional statistical baselines (AUC=0.737) significantly outperform the best-performing Transformer model (AUC=0.670), and that incorporating textual features fails to enhance predictive performance. By challenging the prevailing assumption regarding the inherent superiority of advanced NLP techniques, this work demonstrates that simple covariates offer greater robustness and practical utility for ultra-long-term mental health forecasting.
📝 Abstract
Natural language processing (NLP) models can detect depression-related language in text written near the time symptoms are measured, but whether pretrained transformers can predict depressive symptoms from text written twelve years earlier is largely untested. In the National Child Development Study, a British birth cohort, we predict probable depressive symptoms at age 23 from essays the same people wrote at age 11. Our baseline, a logistic regression on six childhood covariates, outperforms every text model that sees only the essay: seven fine-tuned transformers, a bag-of-words model, frozen embeddings and four zero-shot large language models. Its area under the receiver operating characteristic curve (AUC-ROC) is 0.737 against 0.670 for the best transformer on the primary seed, and no added text score detectably raises the baseline's AUC-ROC. None of the five domain-pretrained transformers detectably beats its general-domain control after Bonferroni correction. For long-horizon prediction, the baseline remains the model to beat.