When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of Vision-Language Models (VLMs) to contextual misguidance in scene text recognition, wherein legible text is erroneously rewritten into contextually plausible alternatives. We introduce the SceneFaith benchmark and propose a Literal/Canonical/Other taxonomy to systematically quantify this "contextual rewriting" behavior. Through generative image construction, comparative evaluations across 15 VLMs, and ablation studies, we reveal the underlying mechanisms driving this phenomenon. Results demonstrate that all evaluated models exhibit contextual rewriting, with error rates ranging from 8.45% to 58.51%. Furthermore, removing background context or enhancing visual clarity significantly improves recognition accuracy. These findings expose a critical bottleneck in how VLMs balance visual evidence against contextual priors during scene text recognition, highlighting the need for more robust alignment between visual perception and linguistic reasoning.
📝 Abstract
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45\% to 58.51\%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Scene Text Recognition
Misleading Context
Contextual Rewriting
Model Faithfulness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Scene Text Recognition
Contextual Bias
Benchmark
Misleading Context
Y
Yuxing Cheng
School of Artificial Intelligence, Jilin University
Y
Yuan Wu
School of Artificial Intelligence, Jilin University
Yi Chang
Yi Chang
Jilin University
Information RetrievalData MiningNatural Language ProcessingMachine LearningArtificial Intelligence