CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the proliferation of video visual tokens with increasing resolution and duration, as well as representation mismatches between training and deployment. To this end, it proposes a codec-native visual encoder that employs segment-wise attention to enable single-pass forward encoding of long videos. The method utilizes learnable abstract tokens to compress contextual information while preserving fine-grained evidence, combined with an alternating layer design and contrastive pretraining to support unified image-text-video understanding. Experimental results demonstrate that, after pretraining on 565M image-text pairs and 6.4M videos, the proposed encoder matches or surpasses OneVision-Encoder performance using only 400 visual tokens. This substantially reduces downstream computational overhead while effectively balancing efficient inference with robust comprehension capabilities.
📝 Abstract
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git
Problem

Research questions and friction points this paper is trying to address.

vision-language models
visual token compression
long-video understanding
scaling bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Codec-Native Vision Encoder
Native Token Compression
Segmented Attention
Abstract Tokens
Unified Image-Video Understanding
🔎 Similar Papers
No similar papers found.