🤖 AI Summary
This work addresses the challenges of redundant noise and ambiguous localization in dense or small-scale scene text detection with multimodal large language models (MLLMs), which stem from conventional multi-visual-token mechanisms. To overcome these limitations, we propose SPaTS, a single-patch text spotting framework that introduces a novel single visual token routing strategy, assigning each text instance a unique anchor token while leveraging global image features for geometric refinement. The framework incorporates an unsupervised reinforcement learning module (SPaSO) to optimally select tokens, complemented by Directional Embedding Alignment (DEA) and Patch-Enhanced Decoding (PED) to enhance localization robustness and accuracy. Extensive experiments demonstrate that SPaTS significantly outperforms state-of-the-art closed-source and OCR-specialized MLLMs across multiple benchmarks, achieving superior end-to-end scene text detection performance.
📝 Abstract
Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code will be released soon.