Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of missing cross-modal correspondence and impeded information transfer in video deepfake detection by proposing the CTCLD framework. This method innovatively introduces a conditional latent variable denoising mechanism to bridge heterogeneous modal distributions, enabling bidirectional audio-visual translation and smooth cross-domain alignment. Furthermore, grounded in Bayesian decomposition theory, it precisely captures subtle inconsistencies across modalities. Experimental results demonstrate that CTCLD achieves comprehensive domain alignment, significantly enhancing both the performance and robustness of video deepfake detection.
📝 Abstract
The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.
Problem

Research questions and friction points this paper is trying to address.

video deepfake detection
cross-modal translation
multimodal detection
audio-visual correspondence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-modal Translation
Conditional Latent Denoising
Video Deepfake Detection
Bayesian Foundation
Domain Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xinzhe Li
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR
Y
Youzhi Tu
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR
Kong Aik Lee
Kong Aik Lee
The Hong Kong Polytechnic University, Hong Kong
Speaker and Spoken Language RecognitionSpeech ProcessingDigital Signal ProcessingSubband