DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the growing diversity of physical and digital face spoofing attacks by proposing a lightweight, vision-only unified anti-spoofing framework that eschews complex multimodal fusion and textual encoders. Built upon a DINO backbone, the method introduces hierarchical Visual Prompt Tuning (VPT) coupled with a Prompt Routing Network (PRN) to dynamically inject input-adaptive prompts, effectively disentangling diverse spoofing artifacts. Through architectural refinements alone, the approach achieves state-of-the-art performance without relying on multimodal cues. Evaluated on the UniAttackData benchmark, it significantly outperforms current leading vision-language models in both accuracy and efficiency, demonstrating that a purely visual design—when enhanced with structured prompt-based adaptation—can surpass more complex multimodal counterparts in face anti-spoofing tasks.
📝 Abstract
With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physical and digital threats. While existing Vision-Language Models (VLMs) demonstrate high generalization in this context, they heavily rely on complex multimodal fusion and external text encoders. In this paper, we propose DINO-VPT, a lightweight, vision-only framework leveraging hierarchical visual prompt tuning. By dynamically injecting prompts conditioned on input features via a Prompt Routing Network (PRN), our method effectively disentangles diverse spoofing artifacts without requiring multimodal fusion. Evaluations on the UniAttackData benchmark demonstrate that DINO-VPT achieves higher accuracy than state-of-the-art VLM-based methods. Our results indicate that a properly structured vision-only architecture can achieve state-of-the-art performance in unified FAS without the need for multimodal supervision.
Problem

Research questions and friction points this paper is trying to address.

Face Anti-Spoofing
Physical-Digital Spoofing
Unified Detection
Vision-only Model
Spoofing Artifacts
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual prompt tuning
face anti-spoofing
vision-only framework
Prompt Routing Network
hierarchical prompting
🔎 Similar Papers
No similar papers found.