🤖 AI Summary
This study addresses automatic depression detection in clinical interview settings through a context-aware multimodal (audio + text) fusion approach. Methodologically, it introduces: (i) a novel topic-modeling–based text data augmentation strategy leveraging BERTopic and LDA; (ii) deep 1D convolutional neural networks for acoustic feature modeling and Transformer architectures for semantic textual representation; and (iii) a cross-modal alignment and adaptive fusion mechanism. Experimental results demonstrate state-of-the-art (SOTA) performance for both unimodal modalities—audio (+3.2% accuracy gain) and text—while the multimodal system achieves performance on par with the best contemporary systems. The proposed framework offers an interpretable, robust paradigm for clinical speech-text analysis under low-resource and high-noise conditions, advancing practical applicability in real-world mental health assessment.
📝 Abstract
In this study, we focus on automated approaches to detect depression from clinical interviews using machine learning approached, which the models are trained on multi-modal data. Differentiating from successful machine learning approaches such as context-aware analysis through feature engineering and end-to-end deep neural networks to depression detection utilizing the Distress Analysis Interview Corpus, we propose a novel method that incorporates a data augmentation procedure based on topic modelling using transformer and deep 1D convolutional neural network (CNN) for acoustic feature modeling. The simulation results demonstrate the effectiveness of the proposed method for training multi-modal deep learning models. Our deep 1D CNN and transformer models achieve the state-of-the-art performance for the audio and text modalities respectively, while our multi-modal results are comparable with the state-of-the-art depression detection systems.