UZH-CL at ArA-DF 2026: Prompt-Tuned Foundation Models and Track-Adaptive Score Fusion for Arabic Speech Deepfake Detection

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generalization challenges in deepfake speech detection under low-resource Arabic dialect diversity and unknown acoustic channels. We propose a wavelet-based prompt tuning method built upon a frozen W2V-BERT-2.0 backbone, integrating cross-layer attention aggregation with attentive statistics pooling to construct the detection system. Furthermore, we introduce a differentiated multi-model track-adaptive score fusion strategy specifically designed to enhance dialectal generalization and acoustic robustness. The proposed system achieves an equal error rate (EER) of 1.96% on Track 1 (ranking 6th) and 1.04% on Track 2 (ranking 3rd) in the official evaluation. These results represent substantial reductions in error rates compared to the baseline systems, demonstrating the effectiveness of the proposed approach for robust deepfake speech detection.
📝 Abstract
Detecting synthetic and voice-converted speech remains difficult for low-resource languages with dialectal diversity, where systems must generalize across regional dialects and unseen acoustic channels. We present the UZH-CL submission to the ArA-DF 2026 Shared Task on Arabic speech deepfake detection, covering Track~1 (dialect generalization) and Track~2 (acoustic robustness). We freeze a W2V-BERT-2.0 backbone and adapt it with \emph{Wavelet Prompt Tuning}, updating under 1% of parameters, and aggregate multi-layer representations with cross-layer attention and a general attentive-statistics pooling head rather than a specialized graph backend. Complementary detectors are obtained by varying adaptation strategy, training data, augmentation, and encoder family. We find that the two shift types require different fusion regimes: a broad multi-window ensemble for dialect generalization, and a compact, channel-matched, center-crop ensemble for acoustic robustness. Official evaluation yields 1.96% EER on Track~1 (6th place) and 1.04% EER on Track~2 (3rd place), corresponding to 87% and 96% relative reductions over the XLS-R+AASIST baselines.
Problem

Research questions and friction points this paper is trying to address.

Arabic speech deepfake detection
dialect generalization
acoustic robustness
low-resource languages
synthetic speech detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wavelet Prompt Tuning
Cross-layer Attention
Track-Adaptive Score Fusion
Arabic Deepfake Detection
Parameter-Efficient Adaptation
🔎 Similar Papers