Video Encoders Built on Image Representations

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between frame interaction timing and visual token cost in video encoding by proposing a novel paradigm that decouples the process into three stages: frame representation, token allocation, and temporal interaction. Methodologically, it adopts a strategy of preserving independent frame evidence before jointly allocating tokens and deferring temporal context integration, thereby preventing premature feature mixing. Architecturally, the proposed compact encoder combines a frozen image encoder, a question-aware selector, and a lightweight residual refinement module. Extensive experiments across thirteen benchmarks demonstrate that this approach matches full-image performance while utilizing only approximately 30% of the tokens, achieving substantial inference acceleration.
📝 Abstract
The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve independently computed frame representations but incur a much larger visual-token cost when all image tokens are forwarded. We ask a basic question: whether a compact video encoder can instead be built on image representations. To answer this question, we separate three operations that are often coupled: per-frame representation, cross-frame token allocation, and temporal interaction. A frozen image encoder first produces frame-specific candidates. A question-aware selector then allocates a fixed token budget across frames using relevance, diversity, and cross-frame correspondence, after which a lightweight learned refiner reads neighboring-frame context and writes residual updates only to the retained anchors. This preserves source positions and keeps the visual output at the fixed budget. Across 13 benchmarks and three vision-language backbones, the resulting pathway matches full-image aggregate performance while using only about 28%-35% of its visual tokens. Specifically, on Qwen3-VL-8B, it achieves a 13-benchmark macro-average of 62.75 with 1,535 visual tokens, compared with 62.58 for the full Image pathway at 4,424 tokens and 59.49 for native Conv3D at 2,212 tokens. On Qwen3-VL-32B, it reaches a 13-benchmark macro-average of 66.28, compared with 66.09 for Image, while providing a 2.16x end-to-end speedup. These results show that compact video encoding does not require early temporal mixing: frame-specific evidence can be preserved first, allocated jointly, and temporally contextualized after selection.
Problem

Research questions and friction points this paper is trying to address.

video encoder
image representations
visual token efficiency
vision-language models
compact encoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Encoder
Image Representations
Token Allocation
Temporal Interaction
Vision-Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jusheng Zhang
Stanford University
W
Wenhao Wang
Vast Intelligence Lab
L
Longqi Cai
Google DeepMind
Liangzhe Yuan
Liangzhe Yuan
Google DeepMind
Multimodal Foundation ModelMachine PerceptionRobotics
Y
Yuxiao Wang
Google DeepMind
Ming-Hsuan Yang
Ming-Hsuan Yang
University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence