DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of existing detection models and the inherent difficulty for a single foundation model to simultaneously capture global semantics and local details. To this end, we propose a hierarchical multi-granularity detection framework. Specifically, by freezing the CLIP and DINOv3 backbones, we design a dual-branch complementary fusion architecture coupled with parameter-efficient adaptation modules to achieve synergistic representation learning of both global semantic and local features. Extensive experiments demonstrate that the proposed framework yields superior performance across multiple benchmarks, significantly enhancing detection capabilities in cross-dataset scenarios and against unseen forgery techniques.
📝 Abstract
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.
Problem

Research questions and friction points this paper is trying to address.

Deepfake Detection
Foundation Models
Generalization
Cross-manipulation
Complementary Fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foundation Models
Dual-Branch Fusion
Deepfake Detection
Parameter-Efficient Adaptation
Multi-granular Representation
🔎 Similar Papers
No similar papers found.
F
Fengming Gu
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, 100049, China; State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China
M
Mingjie He
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China; University of Chinese Academy of Sciences, Beijing, 100049, China
Z
Zonghui Guo
Faculty of Information Science and Engineering, Ocean University of China, Qingdao, 266404, China
J
Jie Zhang
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, 100190, China; University of Chinese Academy of Sciences, Beijing, 100049, China
Shiguang Shan
Shiguang Shan
Professor of Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine LearningFace Recognition