🤖 AI Summary
This study addresses the challenges of missing audio cues and event count calibration bias in multi-segment temporal localization within untrimmed videos by proposing TiTok, an audio-visual large language model. Methodologically, it introduces a novel temporal token interleaving mechanism to align perception with prediction, alongside a Group-reward Decoupled Policy Optimization (GDPO) algorithm designed to enhance localization precision. Furthermore, this work proposes the CountF1 metric and a new evaluation protocol to quantitatively assess count calibration errors. Evaluated on the UnAV-100 benchmark, TiTok achieves state-of-the-art performance with 65.7 mIoU and a CountF1 score of 0.58, significantly improving the accuracy of multi-segment temporal localization.
📝 Abstract
Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration. We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query. For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction. We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO). To assess the performance on AV-MSG, we establish a new UnAV-100-based evaluation protocol, and propose the CountF1 metric for quantifying count miscalibration that overlap metrics fail to capture. TiTok reaches 65.7 mIoU and 0.58 CountF1, achieving state-of-the-art performance. Our code is available at this link.