MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the recognition failure of arbitrary-shaped text in natural scenes caused by detection and localization errors. To this end, we propose a two-stage end-to-end recognition framework that integrates multi-encoder segmentation with vision-language models. Methodologically, six visual encoders, including CLIP, are kept frozen and combined with FPN and PSENet, leveraging feature orthogonality to prevent representation homogenization. Furthermore, Qwen3-VL is incorporated and fine-tuned via LoRA to eliminate background interference. Evaluated on the CTW1500 dataset without synthetic data, the proposed framework achieves state-of-the-art performance, yielding a detection F-measure of 91.99% and an end-to-end H-mean of 85.86%.
📝 Abstract
Scene text spotting remains challenging for arbitrarily shaped text instances such as curved signs and dense multi-oriented characters in natural images, where tightly coupled architectures propagate localization errors directly into recognition failures. We present a two-stage pipeline that combines multi-encoder segmentation with vision-language model recognition to address this problem. In the detection stage, six frozen vision encoders (CLIP, DINOv2, SigLIP, EVA-CLIP, SAM, and ConvNeXt) extract complementary features spanning semantic, spatial, and texture spectra, which are fused through a trainable hierarchical Feature Pyramid Network with channel attention and decoded via a deep-supervision Progressive Scale Expansion network to generate precise instance-level text masks. By keeping the encoders frozen, their independently learned feature spaces remain orthogonal during fusion, preventing the feature homogenization that degrades boundary precision in single-backbone detectors. The detection stage produces tight polygon masks that conform to the actual shape of curved and arbitrarily oriented text, rather than axis-aligned rectangles that inevitably include background content. In the recognition stage, these polygon-masked crops isolate the target text from surrounding clutter, allowing a Qwen3-VL-8B-Instruct model, fine-tuned via Low-Rank Adaptation on polygon-cropped scene text, to focus purely on reading the text without interference from neighbouring words or background noise. Without any synthetic pretraining data, our method achieves 91.99% detection F-measure and 85.86% end-to-end H-mean on CTW1500, setting a new state of the art and achieving strong performance on Total-Text and ICDAR 2015 without any synthetic training data. Code is available at https://github.com/doubleblind-afk/MEVL-STP
Problem

Research questions and friction points this paper is trying to address.

Scene text spotting
Arbitrarily shaped text
Text detection
Text recognition
Natural images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scene Text Spotting
Multi-Encoder Fusion
Vision Language Model
Arbitrarily Shaped Text
Low-Rank Adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aman Anand
Rajiv Gandhi Institute of Petroleum Technology, India
P
Partha Pratim Roy
Indian Institute of Technology (ISM) Dhanbad, India
Shivakumara Palaiahnakote
Shivakumara Palaiahnakote
University of Salford
Artificial Intelligence & Image ProcessingInformation SystemsVideo Text Processing