CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video temporal localization methods predominantly rely on reward signals based solely on prediction correctness, often neglecting the alignment between predicted timestamps and boundary-level visual evidence, which limits localization accuracy. To address this issue, this work proposes a boundary evidence alignment mechanism that explicitly models visual evidence corresponding to temporal boundaries within a reinforcement learning framework by introducing boundary-specific visual evidence tokens and a structured generation strategy. Furthermore, a capability-aware gating mechanism is devised to dynamically modulate supervision strength, enabling adaptive optimization during training. This approach effectively mitigates the misalignment between visual evidence and temporal boundaries, leading to significant performance gains across multiple public benchmarks and demonstrating the efficacy of the proposed alignment mechanism.
📝 Abstract
Large vision-language models (LVLMs) have achieved substantial performance gains in Video Temporal Grounding (VTG) through reinforcement learning (RL). However, existing methods primarily rely on outcome correctness rewards that evaluate only the final predicted intervals, leaving boundary-related visual evidence and its correspondence with timestamp predictions insufficiently constrained. In this paper, we delve into timestamp prediction and its underlying boundary-level visual evidence, showing prevalent misalignment between visual evidence and predicted timestamps across widely used benchmarks. To address this issue, we propose Competence-Aware Visual Boundary Evidence Alignment (CAVE), which augments localization optimization with boundary-specific visual evidence rewards to mitigate evidence-timestamp misalignment. Specifically, to explicitly represent the boundary-specific visual evidence, CAVE introduces boundary-specific evidence tokens and initializes their structured generation and distinct boundary semantics through a lightweight supervised warm-up. During RL, the visual boundary evidence alignment reward reinforces the visual attention of special evidence tokens within the ground-truth boundaries, thereby promoting alignment between visual evidence and temporal boundaries. Moreover, performance-aware gating for evidence supervision is designed to adaptively retain evidence guidance for poorly localized groups while reducing it once localization becomes sufficiently accurate to avoid over-constraining fine-grained boundary refinement. Extensive experiments on several public VTG benchmarks demonstrate the effectiveness of our method.
Problem

Research questions and friction points this paper is trying to address.

Video Temporal Grounding
Visual Boundary Evidence
Evidence-Timestamp Misalignment
Boundary Alignment
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Temporal Grounding
Visual Boundary Evidence Alignment
Reinforcement Learning
Boundary-Specific Evidence Tokens
Competence-Aware Supervision
🔎 Similar Papers
No similar papers found.