🤖 AI Summary
This study addresses the limitation of existing generative video temporal localization models, which lack explicit confidence scores and thus struggle to effectively filter candidate segments. To overcome this, we propose introducing a lightweight confidence head during the decoding stage, combined with offline verification and reinforcement learning for set-level optimization, thereby decoupling candidate generation from acceptance. This approach outputs continuous confidence scores without requiring external verifiers, supporting ranking, threshold-based filtering, and rejection mechanisms while remaining compatible with ground-truth-anchored supervision. Experimental results demonstrate that the proposed framework achieves significant improvements in Recall@0.5 on the OMTG-Bench benchmark. Furthermore, it enables flexible adjustment of precision-recall trade-offs for diverse downstream applications.
📝 Abstract
Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro Recall@0.5 from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.