DuRe-ST: Dual-Relation Spectro-Temporal Modeling for Speech Deepfake Detection

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing speech deepfake detection methods that overlook the covariation between spectral and temporal representations. To this end, we propose a dual-relational modeling framework that constructs a joint covariance affinity graph. By integrating polynomial graph filtering with a graph attention mechanism via normalization, the approach synergistically exploits the complementary relationships induced by covariance and attention to effectively capture high-order dependencies and enhance feature representations. Experimental results demonstrate that the proposed model introduces only minimal additional parameters while achieving substantial performance gains. Specifically, it reduces the equal error rate (EER) by 25.9% on the ASVspoof benchmark and by 28.4% in cross-dataset generalization scenarios, significantly outperforming current state-of-the-art methods.
📝 Abstract
Previous speech deepfake detectors can adaptively capture spectro-temporal dependencies through graph attention, yet they largely overlook the co-variation between spectral and temporal representations. To address this gap, we construct a normalized affinity graph from their joint covariance and apply polynomial graph filtering to capture higher-order covariance-induced dependencies. We first develop Cov-ST to isolate the contribution of covariance-based relational modeling. Although it improves detection performance, its sensitivity to the polynomial order suggests limited robustness when covariance relations are modeled alone. We therefore propose DuRe-ST, which jointly exploits covariance-induced and graph-attention-induced relations to capture complementary second-order co-variation and adaptive spectro-temporal dependencies. Experiments show that DuRe-ST achieves an average relative EER reduction of 25.9% over XLSR-AASIST on the ASVspoof benchmarks and 28.4% across four cross-dataset benchmarks with only 4-8k additional trainable back-end parameters. It further outperforms the strongest publicly available comparison models by 2.2-13.2% in relative EER on four benchmarks, while remaining smaller than the publicly available models considered.
Problem

Research questions and friction points this paper is trying to address.

Speech Deepfake Detection
Spectro-Temporal Modeling
Co-variation
ASVspoof
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Deepfake Detection
Spectro-Temporal Modeling
Covariance Graph
Polynomial Graph Filtering
Graph Attention
🔎 Similar Papers
S
Shaole Li
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR
S
Siqing Qin
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR
Y
Youzhi Tu
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR
Kong Aik Lee
Kong Aik Lee
The Hong Kong Polytechnic University, Hong Kong
Speaker and Spoken Language RecognitionSpeech ProcessingDigital Signal ProcessingSubband