🤖 AI Summary
This study addresses the instability in local event evidence allocation caused by global Top-K competition in sparse video understanding under fixed budgets. To this end, we propose dKFD, a method that partitions budget capacity into early, middle, and late phases based on full-sequence temporal encoding. By integrating a differentiable selector with a phase supervision mechanism, dKFD breaks conventional global competition patterns and establishes an event-centric, controlled structured evidence allocation paradigm. Experimental results demonstrate that dKFD improves Frame AUC by 30.97% on the DoTA dataset, significantly enhancing the alignment between the selector and target events. Furthermore, its effectiveness is validated on the downstream task of Vulnerable Road User (VRU) accident detection.
📝 Abstract
Sparse video understanding often requires selecting a small set of visual evidence under a fixed frame budget. Most sparse selectors allocate this budget globally, allowing all frames to compete with one another. For temporally localized events, this can be a poor inductive bias: useful evidence is often distributed across pre-event context, the event itself, and post-event consequences. We study fixed-budget evidence allocation for localized event videos and show that globally competitive Top-$K$ selectors can preserve recognition and grounding while producing unstable event evidence. On DoTA Video Anomaly Recognition, Global Top-$K$ obtains competitive recognition and temporal grounding, but low selector-event alignment at $K=12$ (Frame AUC $51.5 \pm 10.8$). We propose dKFD, a phase-structured differentiable selector that reserves evidence capacity across pre-event, event, and post-event phases after full-sequence temporal encoding. Under matched-budget multi-seed evaluation, dKFD improves Frame AUC by $+30.97$ over a matched Global Top-$K$ selector at $K=12$ ($p<0.01$), while yielding modest but statistically significant recognition gains and comparable temporal grounding. Mechanism ablations show that phase supervision is load-bearing: removing it reduces Frame AUC to $41.1 \pm 12.0$ even when phase-partitioned budgets are retained. Downstream diagnostics on VRU-Accident show consistent gains over learned Global Top-$K$ across VLM families, while dense captioning reveals a boundary condition where uniform sampling remains competitive. These results support phase-structured allocation as a controlled fixed-budget approach for event-centric sparse evidence selection, not as a universal video summarization strategy.