🤖 AI Summary
This study addresses the resource-intensive nature of conventional Alzheimer’s disease screening and the loss of critical prosodic cues inherent in automatic speech recognition (ASR) pipelines. We propose a bilingual speech large language model framework that bypasses intermediate ASR by directly processing raw audio to learn joint acoustic-semantic representations. Through multi-task learning, the model simultaneously performs cognitive state classification and generates clinically interpretable explanations. Demonstrating zero-shot cross-task generalization, it achieves state-of-the-art accuracy and AUROC across six datasets. Furthermore, clinician evaluations confirm that the generated explanations exhibit high clinical relevance and consistency. Ultimately, this work enables efficient, interpretable, and automated cognitive screening.
📝 Abstract
Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system directly processes raw speech to learn joint acoustic-semantic representations, preserving critical prosodic cues often lost in transcription. Utilising our newly collected PUTH-AD dataset alongside multiple open-source corpora, we implemented a multi-task learning objective that simultaneously performs cognitive status classification and generates clinician-understandable natural language explanations. Our system achieved the highest average accuracy and AUROC across six dataset/task conditions, comparing three representative baselines. The system demonstrated cross-task transfer to held-out PUTH-AD task subsets, maintaining classification accuracy on an entirely unseen cognitive task without task-specific fine-tuning. Furthermore, clinician evaluation confirms that the generated explanations are both clinically relevant and largely consistent with the underlying speech evidence, supporting their potential utility in clinical interpretation. This study provides a scalable, objective, and explainable framework for speech-based cognitive screening, combining cognitive status classification with natural language explanations that clinicians can assess and verify, bridging the gap between advanced AI and clinical utility.