CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings

πŸ“… Unknown Date
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses a critical limitation in existing text-based AI systems for cognitive behavioral therapy (CBT)β€”their inability to capture discrepancies between verbal content and vocal expression, which are essential for assessing psychological distress. To bridge this gap, the authors introduce CBT-Audio, the first publicly available audio dataset specifically curated for CBT sessions, comprising 1,802 patient utterances annotated by experts with distress intensity labels. They systematically evaluate ten open-source audio language models across three modalities: text-only, audio-only, and multimodal fusion. Results demonstrate that eight model families achieve superior performance when incorporating audio, with four showing statistically significant gains. Case analyses further reveal that multimodal fusion excels in scenarios where speech content and prosody conflict, underscoring the irreplaceable value of vocal cues in accurate distress assessment.
πŸ“ Abstract
Cognitive behavioural therapy is widely used to help patients understand and manage psychological distress. It is often delivered through spoken conversation, where therapists attend not only to what patients say, but also to how they say it, because these cues can help therapists decide how to respond and adapt treatment. Progress in building AI systems for CBT remains largely limited to text, partly because most available datasets are text based and shareable spoken CBT data are scarce under ethical and privacy constraints. This creates a blind spot because text based models and evaluations cannot capture the mismatch between the transcript and the patient's voice, even though therapists often rely on this mismatch to understand patient distress. We introduce CBT-Audio, a dataset for evaluating patient distress estimation from spoken CBT sessions with audio language models. CBT-Audio contains 1,802 patient turns from 96 publicly available CBT recordings, with turn-level distress labels validated on an experts-annotated subset. We evaluate 10 open source audio language models under three input conditions, where models receive only patient audio, only the transcript, or both audio and transcript. Our results show that audio can provide useful information beyond text, especially when combined with transcripts. Adding audio to transcript input improves distress estimation over using the transcript alone in 8 of 10 model families, with significant gains in 4, and case studies show the clearest benefit when verbal content and vocal delivery diverge. CBT-Audio makes spoken patient behaviour measurable for AI evaluation in CBT-related tasks and supports future work on audio language models for mental health interaction.
Problem

Research questions and friction points this paper is trying to address.

Cognitive Behavioural Therapy
Distress Estimation
Audio Language Models
Spoken CBT Data
Voice-Text Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio language models
distress estimation
cognitive behavioural therapy
multimodal evaluation
CBT-Audio dataset
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Q
Qixuan Hu
School of Computer Science, Faculty of Engineering, University of Sydney, Australia
Shuchang Ye
Shuchang Ye
The University of Sydney
Multi-modal Learning
X
Xumou Zhang
School of Computer Science, Faculty of Engineering, University of Sydney, Australia
A
Anastasia Serafimovska
School of Psychology, Faculty of Science, University of Sydney, Australia
A
Anastasia Suraev
School of Psychology, Faculty of Science, University of Sydney, Australia
A
Amit Saha
School of Computer Science, Faculty of Engineering, University of Sydney, Australia
P
Ping-hsiu Lin
CHeBA (Centre for Healthy Brain Ageing), School of Clinical Medicine, Discipline of Psychiatry & Mental Health, The University of New South Wales, Australia
S
Sydney Su
Usman Naseem
Usman Naseem
Lecturer (Asst. Prof.) @Macquarie University
Natural Language ProcessingLLM AlignmentNLP for Social GoodTrust and Safety
Adam G. Dunn
Adam G. Dunn
University of Sydney
Public Health InformaticsDigital Public HealthClinical Research InformaticsMisinformation
Jinman Kim
Jinman Kim
School of Computer Science, University of Sydney
Medical VisualisationMedical Image SegmentationTelehealth