RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing video-language models that encode frames independently, thereby exhausting token budgets for long videos or discarding inter-frame dynamics. To overcome this, we propose RESUME, a stateful codec representation framework that initializes latent states via anchor frames and recursively updates them through motion vectors and residual signals, transforming predicted frames into state sequences for large language models. The core innovation lies in constraining the codec recursion within the representation layer, propagating temporal trajectory information through accumulated states rather than independent token sets, alongside a shared readout mechanism. Experiments demonstrate that RESUME significantly outperforms baselines on temporal benchmarks, effectively enhancing temporal reasoning capabilities while maintaining competitive performance on general visual question answering and validating its order sensitivity.
📝 Abstract
Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.
Problem

Research questions and friction points this paper is trying to address.

Video Language Modeling
Temporal Reasoning
Token Efficiency
Codec Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

stateful codec representation
recurrent state updates
video language modeling
temporal reasoning
motion and residual signals
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13