Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting high-fidelity lip-synced deepfake videos, which existing methods struggle to identify due to their neglect of the biological coupling between lip movements and head poses. To overcome this limitation, this work proposes LipDA, a novel framework that pioneers the use of lip-head biocoupling violations as forensic cues. By integrating multimodal feature contrastive learning with temporal dynamics modeling, LipDA quantifies motion inconsistencies and captures generator-specific model fingerprints, thereby unifying deepfake detection and source attribution within a single architecture. Extensive evaluations on two benchmark datasets demonstrate that the proposed framework achieves a detection AUC exceeding 97% and an attribution accuracy of 97.5%, significantly outperforming state-of-the-art techniques in both tasks.
📝 Abstract
Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97\% AUC in detection and 97.5\% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at https://github.com/AnsonShe/LipDA.
Problem

Research questions and friction points this paper is trying to address.

LipSync forgery detection
source attribution
deepfake defense
lip-head inconsistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

LipSync Detection
Source Attribution
Head-Lip Inconsistency
Audio-Visual Synchronization
Deepfake Defense
🔎 Similar Papers
No similar papers found.