Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational cost of high-resolution tokens in vision-language models by introducing foveal compression, which for the first time incorporates the human foveal mechanism into token compression. The proposed method interleaves native- and compressed-resolution tokens, employing a self-distillation merger for feature alignment and a lightweight selector to dynamically determine high-fidelity preservation strategies for individual spatial units under a fixed budget. Experiments demonstrate that this approach matches baseline performance under extremely tight budgets and outperforms random allocation under moderate budgets. Furthermore, the analysis reveals a complementary bottleneck between regional selection and compression fidelity, indicating that localized high-fidelity preservation is not universally optimal.
📝 Abstract
Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Visual Token Compression
Foveated Compression
Token Budget
Region Selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foveated Compression
Visual Tokens
Self-Distillation
Vision-Language Models
Token Efficiency