🤖 AI Summary
为解决病理语音可理解性评估的准确性和可解释性问题,提出基于发音反转的ART-NAD方法,使用声道约束变量代替原有特征,提高了解释性。
📝 Abstract
Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-supervised features that are hard to interpret, providing only frame-level explanations. We propose ART-NAD, a reference-audio intelligibility metric that replaces the \texttt{wav2vec2} features of NAD with vocal-tract constriction variables (tract variables, TVs) predicted from audio by a speaker-independent acoustic-to-articulatory inversion model trained on the same \texttt{wav2vec2} features. ART-NAD is computed as the multivariate Dynamic Time Warping distance between the nine-channel quasi-TV trajectories of the test and one or more references. Across 20 reference-audio protocols spanning six pathological-speech datasets and five languages, ART-NAD with silence trimming (ART-NAD-FA) reaches the same average speaker-level Pearson correlation as NAD-FA (both $r=0.71$) on the same self-supervised backbone, with no significant per-protocol difference (Wilcoxon $p=0.18$), and is the strongest reference-audio metric on 6 of the 20 protocols. Beside the score itself, each TV channel visualizes which constriction deviates from the reference over time, providing interpretable information as to where articulation breaks down.