Grounding with Confidence: Controllable Generative Video Temporal Grounding

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing generative video temporal localization models, which lack explicit confidence scores and thus struggle to effectively filter candidate segments. To overcome this, we propose introducing a lightweight confidence head during the decoding stage, combined with offline verification and reinforcement learning for set-level optimization, thereby decoupling candidate generation from acceptance. This approach outputs continuous confidence scores without requiring external verifiers, supporting ranking, threshold-based filtering, and rejection mechanisms while remaining compatible with ground-truth-anchored supervision. Experimental results demonstrate that the proposed framework achieves significant improvements in Recall@0.5 on the OMTG-Bench benchmark. Furthermore, it enables flexible adjustment of precision-recall trade-offs for diverse downstream applications.
📝 Abstract
Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro Recall@0.5 from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.
Problem

Research questions and friction points this paper is trying to address.

Video Temporal Grounding
Generative Models
Confidence Scoring
Controllable Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Temporal Grounding
Confidence Scoring
Generative Model
Reinforcement Learning
Controllable Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinhao Chen
Alibaba Group; Beihang University
B
Benlei Cui
Alibaba Group
R
Ruijian Jia
Alibaba Group
Z
Ziheng Wang
Alibaba Group; Fudan University
T
Tianyu Wo
Beihang University
P
Pengfei Sun
Alibaba Group
Longtao Huang
Longtao Huang
Alibaba Group
Knowledge GraphService ComputingData Mining
H
Hui Xue
Alibaba Group
Yitong Yang
Yitong Yang
Shanghai University of Finance and Economics
H
Haiwen Hong
Alibaba Group