Context-aware Deep Learning for Multi-modal Depression Detection

📅 2019-05-01
🏛️ IEEE International Conference on Acoustics, Speech, and Signal Processing
📈 Citations: 76
✨ Influential: 7
📄 PDF
🤖 AI Summary
This study addresses automatic depression detection in clinical interview settings through a context-aware multimodal (audio + text) fusion approach. Methodologically, it introduces: (i) a novel topic-modeling–based text data augmentation strategy leveraging BERTopic and LDA; (ii) deep 1D convolutional neural networks for acoustic feature modeling and Transformer architectures for semantic textual representation; and (iii) a cross-modal alignment and adaptive fusion mechanism. Experimental results demonstrate state-of-the-art (SOTA) performance for both unimodal modalities—audio (+3.2% accuracy gain) and text—while the multimodal system achieves performance on par with the best contemporary systems. The proposed framework offers an interpretable, robust paradigm for clinical speech-text analysis under low-resource and high-noise conditions, advancing practical applicability in real-world mental health assessment.

Technology Category

Machine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
In this study, we focus on automated approaches to detect depression from clinical interviews using machine learning approached, which the models are trained on multi-modal data. Differentiating from successful machine learning approaches such as context-aware analysis through feature engineering and end-to-end deep neural networks to depression detection utilizing the Distress Analysis Interview Corpus, we propose a novel method that incorporates a data augmentation procedure based on topic modelling using transformer and deep 1D convolutional neural network (CNN) for acoustic feature modeling. The simulation results demonstrate the effectiveness of the proposed method for training multi-modal deep learning models. Our deep 1D CNN and transformer models achieve the state-of-the-art performance for the audio and text modalities respectively, while our multi-modal results are comparable with the state-of-the-art depression detection systems.
Problem

Research questions and friction points this paper is trying to address.

Depression Detection
Speech Analysis
Text Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Depression Recognition
Pre-trained Transformer Models
1D Convolutional Neural Networks
🔎 Similar Papers
No similar papers found.
Nanyang Technological University | Institute of Infocomm Research | UBTech
G
Genevieve Lam
School of Computer Science and Engineering, Nanyang Technological University, Singapore
D
Dongyan Huang
Institute of Infocomm Research, A*STAR, Singapore; UBTech, Shenzhen City, 518055, P. R. China
Weisi Lin
Weisi Lin
President's Chair Professor in Computer Science, CCDS, Nanyang Technological Unversity
Perception-inspired signal modelingperceptual multimedia quality evaluationvideo compressionimage processing & analysis