🤖 AI Summary
This study addresses the limitations of existing video task adaptation, which relies heavily on extensive data and fine-tuning, while visual in-context learning remains confined to the image domain. This work proposes ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion, enabling training-free unified processing of diverse video tasks. Furthermore, this research identifies and mitigates shortcut effects induced by task internalization, significantly enhancing model robustness. Experimental results demonstrate that ViGeo exhibits strong generalization across varied tasks, transferring effectively to unseen video manipulations and zero-shot modalities such as event cameras. Consequently, this work establishes a novel paradigm for training-free video understanding.
📝 Abstract
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.