TF-PRVR: Training-Free Partially Relevant Video Retrieval

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of semantic dilution and training overfitting caused by fixed decomposition in partial relevant video retrieval. To this end, we propose the first training-free retrieval framework for this task. The method constructs hierarchical representations using frozen vision-language features and adaptively determines temporal boundaries via frequency-domain analysis. Furthermore, a unified multi-scale graph neural network is designed to facilitate cross-scale propagation, coupled with a moment-aware aggregation strategy to achieve precise scoring. Experimental results demonstrate that the proposed framework performs robustly across multiple datasets and effectively suppresses isolated false positives. These findings validate the significant advantages of the training-free paradigm in terms of both generalization capability and practical utility.
📝 Abstract
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and source-domain overfitting induced by task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments demonstrate consistent performance across datasets with diverse visual and temporal characteristics, suggesting a practical direction for training-free PRVR.
Problem

Research questions and friction points this paper is trying to address.

Partially Relevant Video Retrieval
semantic dilution
source-domain overfitting
untrimmed videos
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Partially Relevant Video Retrieval
Hierarchical Representations
Multi-Scale Graph
Frequency-Based Analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.