🤖 AI Summary
This study addresses the prohibitive computational cost of high-resolution tokens in vision-language models by introducing foveal compression, which for the first time incorporates the human foveal mechanism into token compression. The proposed method interleaves native- and compressed-resolution tokens, employing a self-distillation merger for feature alignment and a lightweight selector to dynamically determine high-fidelity preservation strategies for individual spatial units under a fixed budget. Experiments demonstrate that this approach matches baseline performance under extremely tight budgets and outperforms random allocation under moderate budgets. Furthermore, the analysis reveals a complementary bottleneck between regional selection and compression fidelity, indicating that localized high-fidelity preservation is not universally optimal.
📝 Abstract
Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.