FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between visual compression ratio and performance under fixed-resolution settings in long-context reasoning. To overcome this limitation, we propose an adaptive multi-resolution framework that integrates low-DPI global views with selective enhancement of critical regions, thereby transcending conventional fixed-resolution constraints. Methodologically, the approach enables dynamic information localization and enhancement through REL-CoT data construction, multi-resolution supervised fine-tuning (REL-SFT), and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that the proposed method preserves multimodal capabilities while achieving a score of 87.4 on RULER v1 and a 13.91-point improvement on MRCR. Furthermore, it yields a 2.79× inference speedup, significantly outperforming baseline models.
📝 Abstract
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Problem

Research questions and friction points this paper is trying to address.

visual text compression
long-context reasoning
compression-performance trade-off
adaptive resolution
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Text Compression
Adaptive Resolution
Reasoning-Evidence Localization
Supervised Fine-Tuning
Group Relative Policy Optimization
🔎 Similar Papers
No similar papers found.