Voice Biomarkers for Depression and Anxiety

📅 2026-05-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing approaches to speech-based depression and anxiety detection, which rely on handcrafted paralinguistic features and fail to fully exploit biomarkers embedded in raw audio. To overcome this, the authors propose an end-to-end deep learning framework that automatically extracts content-agnostic acoustic biomarkers directly from raw speech signals and integrates lexical features to enhance predictive performance. The model is trained and evaluated on a large-scale, representative dataset comprising approximately 5,000 participants, achieving both sensitivity and specificity of 71%. This work represents the first successful demonstration of automatic extraction of content-agnostic vocal biomarkers without manual feature engineering, and the best-performing model has been made publicly available on Hugging Face.
📝 Abstract
Current approaches to detecting depression and anxiety from speech primarily rely on machine learning techniques that utilize hand-engineered paralinguistic features and related acoustic descriptors derived from time- and frequency-domain representations of speech signals. Applying deep learning methods directly to raw speech signals has the potential to produce biomarker representations with substantially greater predictive power. However, these approaches typically require large volumes of carefully annotated data to learn robust and clinically meaningful representations of the underlying biomarkers. In this paper, we describe our efforts toward developing a deep learning model trained on a large-scale proprietary dataset comprising ~65,000 utterances collected from more than 23,000 subjects representative of relevant United States demographics. We present the techniques employed and analyze their impact on model performance. Our results demonstrate that the proposed models can extract content-agnostic biomarker information, which, when combined with lexical features extracted from audio, yields improved predictive performance in production settings. Our models are evaluated on ~5000 unique subjects and achieve performance of 71% in terms of sensitivity and specificity. To foster further research in mental health assessment from speech, we release the best-performing model described in this paper on HuggingFace.
Problem

Research questions and friction points this paper is trying to address.

Voice Biomarkers
Depression
Anxiety
Speech-based Detection
Mental Health Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

voice biomarkers
deep learning
raw speech signals
content-agnostic representation
mental health assessment
🔎 Similar Papers
No similar papers found.
O
Oleksii Abramenko
Kintsugi Mindful Wellness, Inc.
N
Noah D. Stein
Kintsugi Mindful Wellness, Inc.
Colin Vaz
Colin Vaz
Ph.D. Candidate, University of Southern California
signal processingspeech processingmachine learning