XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出XSQ-AST框架,结合SQ-AST模型、WhisperX音素对齐及多种显著性方法,无需重新训练模型即可定位合成语音中的局部伪影。
📝 Abstract
Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple saliency methods to produce temporally localised artifact diagnostics without model retraining. Saliency maps are projected onto continuous distributions via kernel density estimation and onto phoneme boundaries via phoneme-discretised saliency maps. A 40-participant listening test validated the framework across five perceptual dimensions. Attention Rollout, Attention Flow and an adapted GradCAM produced temporal distributions that correlated with listener highlights, with different methods best suited to different artifact types. An AUC-ROC analysis confirmed discrimination above chance.
Problem

Research questions and friction points this paper is trying to address.

synthetic speech
artifacts
localisation
global quality scores
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explainable Audio Spectrogram Transformer
Synthetic Speech Artifacts
Saliency Methods
Phoneme Alignment
Temporal Localization
🔎 Similar Papers
No similar papers found.
B
Ben Heritage
AudioLab, School of Physics, Engineering and Technology, University of York, United Kingdom
L
Luca Resti
AudioLab, School of Physics, Engineering and Technology, University of York, United Kingdom
M
Mónica Villanueva Aylagas
SEED – Electronic Arts (EA), Sweden
T
Timothy Mehlenbacher
Electronic Arts (EA), United States
K
Konrad Tollmar
SEED – Electronic Arts (EA), Sweden
J
James Alfred Walker
Department of Computer Science, University of York, United Kingdom