π€ AI Summary
This study addresses the limited generalizability of speech-based Alzheimerβs disease screening models across languages, tasks, and recording protocols. Through multi-corpus evaluation, it reveals an underlying feature-direction conflict mechanism. Taking XLM-R as the textual baseline, this work introduces a cross-corpus evidence anchoring mechanism alongside the GroupDRO algorithm, integrating interpretable acoustic-linguistic features via balanced and anchor-dominant fusion strategies to enable feature transferability auditing and worst-case domain robustness optimization. Experimental results demonstrate that the balanced fusion strategy achieves an average AUC of 0.785, while the anchor-dominant strategy elevates the worst-domain AUC to 0.615. Both approaches significantly outperform existing baselines, effectively narrowing the gap toward real-world clinical deployment.
π Abstract
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across corpora, with pause, silence, and speech rate showing high protocol sensitivity. Furthermore, while the XLM-R text baseline achieves strong average performance, its Area Under the ROC Curve (AUC) drops to 0.520 on the weakest held-out domain. A standard GroupDRO baseline reaches a 0.766 mean speaker AUC and a 0.504 worst-domain AUC under the same protocol. To address this, we propose a fusion method that integrates XLM-R text baseline scores with evidence anchors selected during training. Balanced fusion achieves a 0.785 mean speaker AUC, while anchor-heavy fusion raises the worst-case speaker AUC to 0.615. This work highlights the need to audit feature transferability and report worst-case domain robustness in cognitive speech screening.