Off-the-shelf Vision Models Benefit Image Manipulation Localization

📅 2026-04-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Image manipulation localization has long struggled to leverage semantic priors from general-purpose vision models. This work proposes ReVi, a trainable adapter that, for the first time, effectively reveals and exploits such priors. By freezing off-the-shelf vision backbones—such as those from generative or segmentation models—and fine-tuning only lightweight adapters, ReVi decouples semantic redundancy from manipulation-specific cues through a robust principal component analysis–inspired mechanism, thereby enhancing the latter. Without requiring retraining of the backbone, the method achieves significant performance gains across multiple benchmarks, demonstrating the feasibility of building efficient and scalable frameworks for image tampering localization.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessIntelligent Robots: ManipulationMachine Learning: Adversarial Learning & Robustness

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSecurity and Privacy: Privacy-enhancing technologies
📝 Abstract
Image manipulation localization (IML) and general vision tasks are typically treated as two separate research directions due to the fundamental differences between manipulation-specific and semantic features. In this paper, however, we bridge this gap by introducing a fresh perspective: these two directions are intrinsically connected, and general semantic priors can benefit IML. Building on this insight, we propose a novel trainable adapter (named ReVi) that repurposes existing off-the-shelf general-purpose vision models (e.g., image generation and segmentation networks) for IML. Inspired by robust principal component analysis, the adapter disentangles semantic redundancy from manipulation-specific information embedded in these models and selectively enhances the latter. Unlike existing IML methods that require extensive model redesign and full retraining, our method relies on the off-the-shelf vision models with frozen parameters and only fine-tunes the proposed adapter. The experimental results demonstrate the superiority of our method, showing the potential for scalable IML frameworks.
Problem

Research questions and friction points this paper is trying to address.

Image Manipulation Localization
off-the-shelf vision models
semantic priors
model adaptation
manipulation detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

image manipulation localization
off-the-shelf vision models
trainable adapter
semantic disentanglement
robust principal component analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.