Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether acoustic factors introduce systematic confounds and latent robustness vulnerabilities in self-supervised speech models for Alzheimer’s disease (AD) prediction, even when high accuracy is achieved. Leveraging three self-supervised backbone networks, the authors intervene on both input and representation spaces through controlled noise and reverberation perturbations. By combining layer-wise linear decoding with geometric alignment analysis, they establish the causal influence of acoustic factors on AD predictions. The findings demonstrate that acoustic factors lacking significant inter-group differences can nonetheless alter predictive outcomes, with noise perturbations exhibiting the strongest and most directionally consistent effects. Accordingly, this work proposes an intervention-based robustness evaluation criterion that challenges conventional assessment paradigms, offering new evidence to enhance the clinical trustworthiness of speech-based AD detection systems.
📝 Abstract
Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.
Problem

Research questions and friction points this paper is trying to address.

Alzheimer's disease assessment
speech-based assessment
self-supervised learning
acoustic robustness
recording factors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised learning
Intervention-based robustness
Alzheimer's disease assessment
Geometric alignment analysis
Acoustic representations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Serli Kopar
Hertie Institute for AI in Brain Health, Germany; Tübingen AI Center, Germany
Alkis Koudounas
Alkis Koudounas
2x Applied Research Intern @ Amazon AGI | Ph.D. Student @ Politecnico di Torino
Speech and Language ProcessingMultimodalResponsible AI
R
Roshan P. Rane
Hertie Institute for AI in Brain Health, Germany; Tübingen AI Center, Germany
S
Sam Gijsen
Hertie Institute for AI in Brain Health, Germany; Tübingen AI Center, Germany
P
Paula A. Perez-Toro
Friedrich-Alexander-University Erlangen-Nürnberg, Germany
K
Kerstin Ritter
Hertie Institute for AI in Brain Health, Germany; Tübingen AI Center, Germany