From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited perception of tampering traces in multimodal large language models (MLLMs) and the constrained generalization of external expert models by proposing an expert knowledge internalization framework. Through text- and image-focused strategies alongside a forensics-oriented universal representation alignment loss, the method enables progressive knowledge transfer within a two-stage training paradigm. This approach overcomes the dual mismatches in spatial precision and perceptual granularity, deeply integrating forensic capabilities into the MLLM itself. Consequently, the proposed framework achieves efficient inference without relying on external experts, attains state-of-the-art performance across multiple benchmarks, and significantly enhances cross-domain generalization.
📝 Abstract
Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.
Problem

Research questions and friction points this paper is trying to address.

Tampered Text Detection
Multimodal Large Language Models
Expert Knowledge Internalization
Spatial Precision Mismatch
Perceptual Granularity Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Expert Knowledge Internalization
Tampered Text Detection
Multimodal Large Language Models
Forensic-General Representation Alignment
Double Mismatch
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Kaiqing Lin
Kaiqing Lin
Shenzhen University
Multimedia ForensicsMultimedia SecuritySteganalysis
Songze Li
Songze Li
Shanghai AI Laboratory; Fudan University
Computer Vision
S
Shen Chen
Tencent Youtu Lab, Shanghai, China
Y
Yunfei Guo
Tencent Youtu Lab, Shanghai, China
X
Xiaoye Qiu
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen Key Laboratory of Media Security, and SZU-AFS Joint Innovation Center for AI Technology, Shenzhen University, Shenzhen 518060, China
Haodong Li
Haodong Li
UC San Diego. Prev: HKUST, ZJU, Tencent.
3DVGenerative ModelsAgents
Taiping Yao
Taiping Yao
Tencent
face anti-spoofing;deepfake;adversial attack
B
Bo Wang
Tencent Youtu Lab, Shanghai, China
Y
Youchang Xiao
Tencent Youtu Lab, Shanghai, China
Bin Li
Bin Li
Nanjing University of Science and Technology
Algorithmic Game TheoryMechanism DesignAuction Theory
S
Shouhong Ding
Tencent Youtu Lab, Shanghai, China