On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the robustness of Quantum Support Vector Machines (QSVMs) for audio deepfake detection under low-resource and cross-corpus scenarios, where conventional performance typically degrades. Methodologically, features are extracted using a frozen wav2vec 2.0 encoder and reduced to four dimensions via PCA to accommodate the constraints of near-term quantum hardware. QSVMs are then systematically compared against classical SVMs and multilayer perceptrons (MLPs). The findings reveal that the inductive bias inherent in quantum kernels confers distinct advantages under distribution shifts: in severe domain-shift scenarios, QSVMs achieve an AUC of 76%, substantially outperforming MLPs, which exhibit near-random behavior. However, no consistent gains are observed in near-domain transfer settings. Overall, this work empirically demonstrates the potential of QSVMs to address extreme domain shifts in audio deepfake detection.
📝 Abstract
Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited training data. We compare a Quantum Support Vector Machine (QSVM), a classical support vector machine (SVM), and a multilayer perceptron (MLP), all trained on frozen wav2vec 2.0 embeddings using a strict budget of 200 training samples. To match the qubit budget of near-term quantum hardware, embeddings are reduced to four dimensions using principal component analysis, and all models use the same reduced features. Experiments on ASVspoof 2019, ASVspoof 5, the ADD 2023 Challenge, and the In-the-Wild dataset show that under severe domain shift from ASVspoof 2019 to ADD 2023, the MLP degrades to near-random performance, with an area under the curve of approximately 50% and an equal error rate of 50.0%. In contrast, the QSVM maintains meaningful discrimination, achieving an area under the curve of 76.0% and an equal error rate of 27.0%. This advantage is not consistent across transfer directions. When trained on ADD 2023, the QSVM falls below chance on two of three transfers, while the MLP performs better. These results suggest that quantum kernel methods can be competitive under severe cross-corpus shifts and strict low-resource constraints, but do not provide a consistent advantage under near-domain transfer. We interpret these findings as an empirical characterization of quantum kernel inductive bias under distribution shift, rather than evidence of quantum advantage, since the four-qubit kernel can be simulated exactly on classical hardware.
Problem

Research questions and friction points this paper is trying to address.

Audio Deepfake Detection
Cross-Corpus
Low-Resource
Quantum Kernel
Distribution Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantum Kernel Methods
Cross-Corpus Deepfake Detection
Low-Resource Learning
Domain Shift Robustness
wav2vec 2.0 Embeddings
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.