Conformal Coverage Guarantees for Any Video Temporal Grounder

πŸ“… 2026-08-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing video temporal localization methods, which typically output a single prediction interval and struggle to quantify reliability or handle annotation ambiguity. To overcome these issues, we propose COVERβ€”a model-agnostic, post-hoc calibration framework that provides finite-sample, distribution-free coverage guarantees: for any given confidence level \(1-\alpha\), the calibrated interval contains the true event interval with probability at least \(1-\alpha\). Without requiring model retraining or white-box access, COVER leverages quantile calibration of temporal inconsistency scores, with tailored score designs for both interval-based and correlation-based localizers. We further establish theoretical analysis specific to this task. Experiments across three benchmark datasets and five diverse localizers demonstrate that COVER achieves precise empirical coverage while uncovering performance differences obscured by conventional point-based evaluation metrics.
πŸ“ Abstract
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is indistinguishable from a right one. COVER changes the output object: a post-hoc, model-agnostic wrapper that turns any grounder, a trained localizer or a black-box video--language model, into one that emits a temporal region containing the true moment with probability at least $1-Ξ±$, by calibrating the quantile of a temporal nonconformity score on held-out labels and widening the base prediction by that amount. The guarantee is finite-sample and distribution-free under exchangeability, and requires neither retraining nor white-box access. We give two score families, a two-sided boundary-widening score for grounders that emit an interval and a super-level-set score for grounders that emit a relevance signal, and develop theory specific to grounding that bounds how large the certified region becomes, when coverage survives conditioning on event length, and how it degrades when moments from one video break exchangeability. Across three benchmarks and five grounders, realized coverage tracks the target, and calibration exposes what point metrics hide.
Problem

Research questions and friction points this paper is trying to address.

video temporal grounding
event boundary ambiguity
conformal coverage
reliability calibration
annotation uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

conformal prediction
temporal grounding
coverage guarantee
model-agnostic calibration
video-language models
πŸ”Ž Similar Papers
No similar papers found.
A
Aseel Mohamed
Department of Electrical and Computer Engineering, Texas A&M University at Qatar, Doha, Qatar
R
Rasul Khanbayov
College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar
E
Erchin Serpedin
Department of Computer and Electrical Engineering, Texas A&M University, College Station, TX, USA
Hasan Kurban
Hasan Kurban
Hamad Bin Khalifa University
Artificial IntelligenceSoftware EngineeringAI for Science