Counterfactual Attention Policy Distillation for Temporal Video Grounding

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of multimodal large language models to repetitive actions and visually similar scenes, which compromises temporal localization accuracy in long videos. To this end, it proposes a counterfactual attention policy distillation framework that pioneers the integration of counterfactual intervention into online policy distillation. By masking temporal groups to quantify the causal effect of individual tokens on teacher outputs, the method calibrates attention weights accordingly. This enables token-level distillation weighting grounded in temporal evidence, thereby overcoming the supervision limitations inherent in conventional next-token prediction. Built upon Qwen3-VL and trained with only 2,500 samples, the proposed approach achieves a 12.0% improvement in average recall over GRPO on the TimeLens benchmark while preserving general video understanding capabilities.
📝 Abstract
Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.
Problem

Research questions and friction points this paper is trying to address.

Temporal Video Grounding
Multimodal Large Language Models
On-policy Distillation
Long Videos
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Attention Policy Distillation
Temporal Video Grounding
On-policy Distillation
Multimodal Large Language Models
Counterfactual Influence