A Lightweight and Effective Image Tampering Localization Network with Vision Mamba

📅 2025-02-14
📈 Citations: 0
Influential: 0
📄 PDF

career value

222K/year
🤖 AI Summary
Existing image forgery localization methods face a fundamental trade-off between global modeling capability (limited by CNNs) and computational efficiency (hampered by Transformer’s quadratic complexity). Method: We propose ForMa—the first lightweight localization network for blind image forensics that integrates Vision Mamba, a linear-complexity state-space model, to achieve efficient global dependency modeling. It introduces a parameter-free pixel-rearrangement upsampling decoder and a novel noise-augmented decoding mechanism to enhance sensitivity to subtle tampering cues. Multi-scale Vision Mamba features are fused to capture hierarchical long-range contextual information. Contributions/Results: (1) First application of Vision Mamba to image forgery localization; (2) Novel noise-injection decoding and parameter-free upsampling strategy. ForMa achieves state-of-the-art generalization and robustness across 10 standard benchmarks with the lowest computational overhead. Code is publicly available.

Technology Category

Application Category

📝 Abstract
Current image tampering localization methods primarily rely on Convolutional Neural Networks (CNNs) and Transformers. While CNNs suffer from limited local receptive fields, Transformers offer global context modeling at the expense of quadratic computational complexity. Recently, the state space model Mamba has emerged as a competitive alternative, enabling linear-complexity global dependency modeling. Inspired by it, we propose a lightweight and effective FORensic network based on vision MAmba (ForMa) for blind image tampering localization. Firstly, ForMa captures multi-scale global features that achieves efficient global dependency modeling through linear complexity. Then the pixel-wise localization map is generated by a lightweight decoder, which employs a parameter-free pixel shuffle layer for upsampling. Additionally, a noise-assisted decoding strategy is proposed to integrate complementary manipulation traces from tampered images, boosting decoder sensitivity to forgery cues. Experimental results on 10 standard datasets demonstrate that ForMa achieves state-of-the-art generalization ability and robustness, while maintaining the lowest computational complexity. Code is available at https://github.com/multimediaFor/ForMa.
Problem

Research questions and friction points this paper is trying to address.

Lightweight image tampering localization
Global dependency modeling with linear complexity
Enhanced decoder sensitivity to forgery cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Mamba for global modeling
Linear-complexity feature capture
Noise-assisted decoding strategy
Kun Guo
Kun Guo
School of Psychology, Sport Science & Wellbeing, University of Lincoln
Social attentionvisual perceptioncognitive neuroscience
G
Gang Cao
School of Computer and Cyber Sciences, Communication University of China, Beijing 100024, China, and also with the State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing 100024, China
Z
Zijie Lou
School of Computer and Cyber Sciences, Communication University of China, Beijing 100024, China, and also with the State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing 100024, China
X
Xianglin Huang
School of Computer and Cyber Sciences, Communication University of China, Beijing 100024, China, and also with the State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing 100024, China
J
Jiaoyun Liu
School of Information Engineering, Changsha Medical University, Changsha 410219, China