CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of signal dilution and curvature rigidity caused by fixed geometric spaces in partially relevant video retrieval. To this end, we propose CurvSpec, a novel framework that introduces the first content-adaptive multi-curvature learning mechanism. By discarding fixed geometric priors, CurvSpec employs parallel Euclidean and hyperbolic attention layers to independently learn curvatures and dynamically route representations to optimal geometric manifolds. Furthermore, it integrates semantic centroid matching with geodesic distance metrics to effectively suppress noise. Experimental results demonstrate that the proposed method achieves state-of-the-art retrieval performance on benchmark datasets such as ActivityNet Captions.
📝 Abstract
Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging than standard full-video retrieval. This task presents two intertwined challenges: (1) signal dilution, where coarse global representations blur the brief relevant signal into the dominant irrelevant surroundings;(2) curvature rigidity, where embedding all videos in the same fixed-geometry space distorts representations for videos that range from flat atomic events to deep compositional hierarchies. Existing PRVR methods have improved moment selection and cross-modal matching, but they still typically encode all videos in a single fixed-curvature retrieval space, limiting their ability to model diverse video structures. To address both challenges, we propose CurvSpec, a framework that learns content-adaptive curvature for video retrieval representations rather than imposing a fixed geometric prior. CurvSpec processes features through parallel Euclidean and hyperbolic attention layers, with independently learned curvatures assigned to the hyperbolic layers, and a content-aware fusion mechanism routes each input to its most suitable geometric regime. To further suppress signal dilution, CurvSpec represents each video with semantic centroids whose number is determined by the video's content complexity, projects them onto the learned manifold, and matches each query against its nearest centroid by geodesic distance. Experiments on ActivityNet Captions, TVR, and Charades-STA demonstrate state-of-the-art retrieval performance.
Problem

Research questions and friction points this paper is trying to address.

Partially Relevant Video Retrieval
Signal Dilution
Curvature Rigidity
Video Representation
Cross-modal Matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Partially Relevant Video Retrieval
Adaptive Multi-Curvature Learning
Hyperbolic Attention
Semantic Centroids
Geodesic Distance
🔎 Similar Papers
No similar papers found.