Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of existing self-supervised anti-spoofing models under unseen attacks, where the semantic processing in audio language models (ALMs) often attenuates critical acoustic cues. Building upon the Voxtral framework, this work investigates the conflict mechanism between semantic and acoustic representations within ALMs and proposes Spooftral, a novel model employing an instruction-guided detection paradigm based on label sequence likelihood. By integrating DoRA-based low-rank adaptation with instruction tuning, the proposed approach effectively mitigates feature conflicts and enhances detection robustness. Experimental results demonstrate that Spooftral achieves an equal error rate of 4.25% on the ASVspoof 5 evaluation set, establishing a new paradigm for cross-domain speech spoofing detection.
📝 Abstract
Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines the Voxtral audio-language model (ALM) framework for spoofing detection, as a step toward combining CM capabilities within the ALM framework. We analyze how Voxtral captures spoofing cues through audio-text processing and propose an instruction-guided approach that uses label-sequence likelihoods to evaluate bonafide and spoofed speech. Experiments on the ASVspoof databases show that without task-specific adaptation, the LLM layers emphasize semantic representations, reducing the separability of spoof-discriminative acoustic cues compared to the Whisper-based audio encoder. Consequently, spoofing-related information becomes less separable after language-model processing. We also applied lightweight adaptation using weight-decomposed low-rank adaptation (DoRA) to the Voxtral model and propose the Spooftral model, achieving an equal error rate (EER) of 4.25% on the ASVspoof5 evaluation set.
Problem

Research questions and friction points this paper is trying to address.

Speech Spoofing Detection
Audio-Language Model
Self-Supervised Learning
Countermeasures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Language Model
Speech Spoofing Detection
Instruction-Guided Approach
Weight-Decomposed Low-Rank Adaptation (DoRA)
Spooftral
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Avishai Weizman
School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel
Y
Yehuda Ben-Shimol
School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel
Itshak Lapidot
Itshak Lapidot
Afeka Tel-Aviv Academic College of Engineering
Speaker verificationSpeaker diarizationAnti-spoofingML for Biomedical Spectroscopy