🤖 AI Summary
This study addresses the high computational cost induced by dense frame features in video retrieval by proposing an efficient and scalable video-text retrieval framework. Methodologically, it reduces representational redundancy among sampled frames through joint encoding and feature compression, while establishing a reusable caching mechanism to eliminate redundant computations. Furthermore, feature variation prediction is introduced as an auxiliary supervision signal to enhance retrieval accuracy without increasing query overhead. Experimental results demonstrate that the proposed approach significantly improves the R@1 metric across multiple benchmark datasets while substantially reducing both storage and computational costs, thereby achieving an effective balance between efficiency and performance.
📝 Abstract
Finding the right video often requires distinguishing similar scenes in which different events occur. Joint matching improves retrieval, but processing rich video representations for each query is costly. VEDJE compresses features within sampled frames while keeping their representations separate in a reusable cache. Feature-change prediction supplies an auxiliary training signal that improves retrieval from the compressed cache without adding work at query time. On MSR-VTT, MSVD, DiDeMo, and ActivityNet, VEDJE improves R@1 over matched first-stage retrievers in both retrieval directions. On MSR-VTT, it reaches 59.8 text-to-video R@1 with a fine-tuned VideoCLIP-XL first stage. In the VideoPrism configuration, shrinking the per-video cache fourfold to 12 KiB preserves text-to-video recall within 0.2 points. These results show that accurate video search can operate on compact evidence, encoded once and reused as new queries arrive.