Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the significant degradation in generalization of existing audio deepfake detection models when faced with shifts in generators, corpora, or recording conditions. To enhance cross-domain robustness, the authors propose a reconstruction probing method based on a frozen Diffusion Transformer (DiT). Trained exclusively on genuine speech, the DiT generates reconstruction residuals under multi-rate masking. These residuals are then integrated into WavLM representations via a novel audio-anchored fusion mechanism, which applies them as scalar-gated additive corrections in a non-competitive manner, thereby avoiding gating-induced attenuation. The approach achieves EERs of 6.54% and 13.84% on ASVspoof 2021 Eval and In-the-Wild Full partitions, respectively, outperforming a specialized WavLM-ResNet18 baseline across three subsets on average and substantially improving cross-domain detection robustness.
📝 Abstract
Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions at masking ratios 0.5, 0.75, and 0.9 yield explicit multi-ratio residual maps. Because these residuals are domain sensitive, our audio-anchored detector passes the projected frozen-WavLM auditory representation into the fusion sum without gate-based attenuation and uses residuals only as a scalar-gated additive correction. The pre-specified seed-42 run obtains 6.5442% EER / 0.18456 min-DCF on ASVspoof 5 Eval and 13.8372% / 0.36921 on ITW Full; three-seed means are 6.8885 (0.3308)% and 15.3328 (2.0719)%. The latter is below a separately optimized WavLM-ResNet18 reference under both supervision settings. Auxiliary supervision raises dynamic competitive fusion from 18.4007% to 25.2968% mean ITW EER, worsening all three seeds. The results support reconstruction residuals as complementary evidence and motivate a non-competitive auditory path for ASVspoof 5-to-ITW transfer, without claiming a componentwise causal ablation of anchoring alone.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake detection
cross-domain
domain shift
reconstruction residuals
robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Transformer
reconstruction residuals
audio-anchored fusion
cross-domain deepfake detection
frozen representation
🔎 Similar Papers
No similar papers found.
H
Haotian Mo
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Jie Liu
Jie Liu
Unknown affiliation
Numerical Method for Partial Differential Equations
Siqi Shen
Siqi Shen
Xiamen University
Reinforcement Learning3D Vision
S
Songzhu Mei
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
X
Xinhai Chen
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Xiangyang Wang
Xiangyang Wang
Shenzhen Institues of Advanced Technology (SIAT), Chinese Academy of Science
ExoskeletonBioroboticsWearable robot
Y
Yigui Feng
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
Shuai Li
Shuai Li
Massachusetts Institute of Technology
Optical imaging
G
Gencheng Liu
College of Computer Science and Technology, National University of Defense Technology, Changsha, China
K
Keqi Yang
Hunan Zhongke Youxin Technology Co., Ltd., Changsha, China
Qinglin Wang
Qinglin Wang
National University of Defense Technology
Parallel algorithmsHigh Performance ComputingDeep LearningMachine LearningGPU