🤖 AI Summary
Depression prediction is hindered by scarce high-quality labeled data, limited model interpretability, and poor clinical applicability. To address these challenges, we propose the Score-guided Token Probability Summation (SToPS) module, integrated with a domain-adapted pretrained language model. Our approach leverages a novel, real-world autobiographical narrative corpus comprising 3,699 samples—curated from ecological momentary assessment (EMA) and publicly available clinical interview transcripts. STOPS enables interpretable, high-confidence depression detection via fine-grained, token-level probability guidance. Experimental results demonstrate substantial improvements in predictive reliability: an overall AUC of 0.789 and 0.904 on high-confidence subsets. Cross-scenario validation confirms robust generalizability. This work pioneers the integration of token-level probabilistic guidance into depression screening, uniquely balancing predictive performance, model transparency, and clinical credibility—establishing a new paradigm for early, scalable, and trustworthy depression risk assessment.
📝 Abstract
Advances in large language models (LLMs) have enabled a wide range of applications. However, depression prediction is hindered by the lack of large-scale, high-quality, and rigorously annotated datasets. This study introduces DepressLLM, trained and evaluated on a novel corpus of 3,699 autobiographical narratives reflecting both happiness and distress. DepressLLM provides interpretable depression predictions and, via its Score-guided Token Probability Summation (SToPS) module, delivers both improved classification performance and reliable confidence estimates, achieving an AUC of 0.789, which rises to 0.904 on samples with confidence $geq$ 0.95. To validate its robustness to heterogeneous data, we evaluated DepressLLM on in-house datasets, including an Ecological Momentary Assessment (EMA) corpus of daily stress and mood recordings, and on public clinical interview data. Finally, a psychiatric review of high-confidence misclassifications highlighted key model and data limitations that suggest directions for future refinements. These findings demonstrate that interpretable AI can enable earlier diagnosis of depression and underscore the promise of medical AI in psychiatry.