VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

📅 2026-01-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing image editing models in handling multilingual, text-dense, and structurally complex visual documents—particularly Chinese-language materials—and the absence of dedicated evaluation benchmarks. To bridge this gap, we introduce the first high-quality, manually annotated bilingual (Chinese–English) benchmark for visual document editing, encompassing diverse and intricate layouts such as academic papers, posters, and exam sheets. We propose a decoupled evaluation framework grounded in OCR parsing that enables fine-grained assessment of editing fidelity across multiple dimensions, including textual content accuracy, layout preservation, and linguistic consistency. Experimental results demonstrate that our benchmark effectively exposes the shortcomings of state-of-the-art models, with automated evaluation scores showing strong correlation with human judgments.

Technology Category

Application Category

📝 Abstract
In recent years, multimodal image editing models have achieved substantial progress, enabling users to manipulate visual content through natural language in a flexible and interactive manner. Nevertheless, an important yet insufficiently explored research direction remains visual document image editing, which involves modifying textual content within images while faithfully preserving the original text style and background context. Existing approaches, including AnyText, GlyphControl, and TextCtrl, predominantly focus on English-language scenarios and documents with relatively sparse textual layouts, thereby failing to adequately address dense, structurally complex documents or non-Latin scripts such as Chinese. To bridge this gap, we propose \textbf{V}isual \textbf{D}oc \textbf{E}dit Bench(VDE Bench), a rigorously human-annotated and evaluated benchmark specifically designed to assess image editing models on multilingual and complex visual document editing tasks. The benchmark comprises a high-quality dataset encompassing densely textual documents in both English and Chinese, including academic papers, posters, presentation slides, examination materials, and newspapers. Furthermore, we introduce a decoupled evaluation framework that systematically quantifies editing performance at the OCR parsing level, enabling fine-grained assessment of text modification accuracy. Based on this benchmark, we conduct a comprehensive evaluation of representative state-of-the-art image editing models. Manual verification demonstrates a strong consistency between human judgments and automated evaluation metrics. VDE Bench constitutes the first systematic benchmark for evaluating image editing models on multilingual and densely textual visual documents.
Problem

Research questions and friction points this paper is trying to address.

visual document editing
multilingual text editing
dense textual layout
non-Latin scripts
image editing models
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual document editing
multilingual benchmark
dense text layout
decoupled evaluation framework
OCR-based assessment
H
Hongzhu Yi
University of Chinese Academy of Sciences
Y
Yujia Yang
University of Chinese Academy of Sciences
Y
Yuanxiang Wang
University of Chinese Academy of Sciences
Z
Zhenyu Guan
University of Chinese Academy of Sciences
J
Jiahuan Chen
University of Chinese Academy of Sciences
Chenxi Bao
Chenxi Bao
MBZUAI
Music GenerationInteractive Music DesignComputer Music
T
Tiankun Yang
University of Chinese Academy of Sciences
Yixuan Yuan
Yixuan Yuan
Associate Professor in Chinese University of Hong Kong
Medical image analysisAI in healthcareBrain data analysisEndoscopy
T
Tianyu Zong
University of Chinese Academy of Sciences
X
Xinming Wang
Institute of Automation, Chinese Academy of Sciences
Tao Yu
Tao Yu
Institute of Automation, Chinese Academy of Sciences
MLLM
R
Ruiwen Tao
Tencent
H
Haijin Liang
Tencent
J
Jin Ma
Tencent
J
Jinwen Luo
Tencent
Y
Yeshani Xinyu Zuo
Tencent
Jungang Xu
Jungang Xu
Professor, University of Chinese Academy of Sciences
Multimodal Data Processing