Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring

πŸ“… 2026-09-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited generalizability of speech-based Alzheimer’s disease screening models across languages, tasks, and recording protocols. Through multi-corpus evaluation, it reveals an underlying feature-direction conflict mechanism. Taking XLM-R as the textual baseline, this work introduces a cross-corpus evidence anchoring mechanism alongside the GroupDRO algorithm, integrating interpretable acoustic-linguistic features via balanced and anchor-dominant fusion strategies to enable feature transferability auditing and worst-case domain robustness optimization. Experimental results demonstrate that the balanced fusion strategy achieves an average AUC of 0.785, while the anchor-dominant strategy elevates the worst-domain AUC to 0.615. Both approaches significantly outperform existing baselines, effectively narrowing the gap toward real-world clinical deployment.
πŸ“ Abstract
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across corpora, with pause, silence, and speech rate showing high protocol sensitivity. Furthermore, while the XLM-R text baseline achieves strong average performance, its Area Under the ROC Curve (AUC) drops to 0.520 on the weakest held-out domain. A standard GroupDRO baseline reaches a 0.766 mean speaker AUC and a 0.504 worst-domain AUC under the same protocol. To address this, we propose a fusion method that integrates XLM-R text baseline scores with evidence anchors selected during training. Balanced fusion achieves a 0.785 mean speaker AUC, while anchor-heavy fusion raises the worst-case speaker AUC to 0.615. This work highlights the need to audit feature transferability and report worst-case domain robustness in cognitive speech screening.
Problem

Research questions and friction points this paper is trying to address.

Alzheimer's disease screening
cross-corpus generalization
deployment gap
speech-based detection
domain robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Corpus Generalization
Evidence Anchoring
Alzheimer's Speech Screening
Domain Robustness
Feature Transferability
πŸ”Ž Similar Papers
Z
Zijian Lu
Nanjing University of Posts and Telecommunications, Nanjing, China
S
Sizhe Liu
University of Science and Technology of China, Hefei, China
Y
Yin Zhang
University of Science and Technology of China, Hefei, China
J
Jixuan Deng
University of Science and Technology of China, Hefei, China
X
Xinrong Lin
Shouyi Technology, Hefei, China
X
Xinchen Yuan
University of Science and Technology of China, Hefei, China
C
Chicheng Jin
University of Science and Technology of China, Hefei, China
Y
Yiping Zuo
Nanjing University of Posts and Telecommunications, Nanjing, China
Yuanchao Li
Yuanchao Li
University of Edinburgh
speech technologiesspoken language processingaffective computingdigital healthHCI