Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the progressive degradation of temporal representations during inference in large video models, where temporal information diminishes across network layers and ultimately vanishes, resulting in weak temporal reasoning capabilities. By elucidating this decay mechanism and monitoring layer-wise temporal divergence vectors to precisely identify intermediate-layer peaks, this work proposes a training-free temporal activation injection strategy that reintroduces critical temporal information into subsequent layers to preserve representational integrity. As the first training-free inference enhancement approach of its kind, the proposed method significantly improves temporal reasoning performance across three mainstream architectures and four benchmarks, while introducing negligible impact on non-temporal tasks.
📝 Abstract
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $τ_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $τ_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
Problem

Research questions and friction points this paper is trying to address.

VideoLLMs
temporal reasoning
representation fading
temporal divergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

VideoLLMs
Temporal Reasoning
Training-Free Inference
Temporal Activation Injection
Representation Decay
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Youngwoo Shin
Korea Advanced Institute of Science and Technology (KAIST)
Y
Yusung Ro
Korea Advanced Institute of Science and Technology (KAIST)
M
Minseo Kim
Korea Advanced Institute of Science and Technology (KAIST)
Junmo Kim
Junmo Kim
School of Electrical Engineering, KAIST
Statistical Signal ProcessingImage ProcessingComputer VisionMachine LearningInformation Theory