🤖 AI Summary
This study addresses the limited generalization of existing detection models and the inherent difficulty for a single foundation model to simultaneously capture global semantics and local details. To this end, we propose a hierarchical multi-granularity detection framework. Specifically, by freezing the CLIP and DINOv3 backbones, we design a dual-branch complementary fusion architecture coupled with parameter-efficient adaptation modules to achieve synergistic representation learning of both global semantic and local features. Extensive experiments demonstrate that the proposed framework yields superior performance across multiple benchmarks, significantly enhancing detection capabilities in cross-dataset scenarios and against unseen forgery techniques.
📝 Abstract
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.