Quantum Vision Theory Applied to Audio Classification for Deepfake Speech Detection

📅 2026-04-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of accuracy and robustness in deepfake audio detection by introducing, for the first time, quantum vision (QV) theory—inspired by quantum wave-particle duality—into the audio domain. The proposed paradigm transforms time-frequency features such as STFT, Mel-spectrograms, and MFCCs into information waves via a QV module, followed by authenticity classification using QV-enhanced CNN and Vision Transformer (QV-ViT) architectures. Evaluated on the ASVspoof dataset, the QV-CNN model achieves 94.20% accuracy and a 9.04% equal error rate (EER) with MFCC inputs, while its Mel-spectrogram variant attains 94.57% accuracy, significantly outperforming current state-of-the-art methods.

Technology Category

Machine Learning: Quantum Machine LearningComputer Vision: Diffusion Models for VisionNatural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)

Application Category

Web Mining and Content Analysis: Web data provenance, reliability, and authenticitySecurity and Privacy: Data transparency and provenanceSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
We propose Quantum Vision (QV) theory as a new perspective for deep learning-based audio classification, applied to deepfake speech detection. Inspired by particle-wave duality in quantum physics, QV theory is based on the idea that data can be represented not only in its observable, collapsed form, but also as information waves. In conventional deep learning, models are trained directly on these collapsed representations, such as images. In QV theory, inputs are first transformed into information waves using a QV block, and then fed into deep learning models for classification. QV-based models improve performance in image classification compared to their non-QV counterparts. What if QV theory is applied speech spectrograms for audio classification tasks? This is the motivation and novelty of the proposed approach. In this work, Short-Time Fourier Transform (STFT), Mel-spectrograms, and Mel-Frequency Cepstral Coefficients (MFCC) of speech signals are converted into information waves using the proposed QV block and used to train QV-based Convolutional Neural Networks (QV-CNN) and QV-based Vision Transformers (QV-ViT). Extensive experiments are conducted on the ASVSpoof dataset for deepfake speech classification. The results show that QV-CNN and QV-ViT consistently outperform standard CNN and ViT models, achieving higher classification accuracy and improved robustness in distinguishing genuine and spoofed speech. Moreover, the QV-CNN model using MFCC features achieves the best overall performance on the ASVspoof dataset, with an accuracy of 94.20% and an EER of 9.04%, while the QV-CNN with Mel-spectrograms attains the highest accuracy of 94.57%. These findings demonstrate that QV theory is an effective and promising approach for audio deepfake detection and opens new directions for quantum-inspired learning in audio perception tasks.
Problem

Research questions and friction points this paper is trying to address.

deepfake speech detection
audio classification
speech spoofing
quantum-inspired learning
ASVSpoof
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantum Vision
deepfake speech detection
information waves
QV-CNN
audio classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Khalid Zaman
Graduate School of Advanced Science and Technology, Japan Advanced Institute of Science and Technology, Nomi, 923-1292, Ishikawa, Japan
M
Melike Sah
Computer Engineering Department, Cyprus International University, Nicosia, 99258, North Cyprus via Mersin 10, Turkiye
A
Anuwat Chaiwongyen
Department of Management Information Systems, Thammasat University, Khlong Luang, Pathum Thani, 12121, Thailand
Cem Direkoglu
Cem Direkoglu
Associate Professor at Middle East Technical University - Northern Cyprus Campus
Computer VisionImage AnalysisVideo AnalysisSignal ProcessingMachine Learning