🤖 AI Summary
This study addresses the challenges of missing cross-modal correspondence and impeded information transfer in video deepfake detection by proposing the CTCLD framework. This method innovatively introduces a conditional latent variable denoising mechanism to bridge heterogeneous modal distributions, enabling bidirectional audio-visual translation and smooth cross-domain alignment. Furthermore, grounded in Bayesian decomposition theory, it precisely captures subtle inconsistencies across modalities. Experimental results demonstrate that CTCLD achieves comprehensive domain alignment, significantly enhancing both the performance and robustness of video deepfake detection.
📝 Abstract
The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.