Allocate Before You Embed: Adaptive Visual Input Allocation for Video Embeddings

📅 2026-09-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决大规模视频检索中视觉输入和推理预算紧张问题,提出AllocEmbed框架,通过轻量级分配器动态调整帧分辨率,提高检索性能。
📝 Abstract
Large-scale video retrieval requires embedding models to encode long and diverse videos under tight visual-input and inference budgets. Existing methods typically sample a small, fixed set of frames at their original resolution, limiting temporal coverage and ignoring frame importance. Our empirical analysis shows that expanding temporal coverage improves retrieval even under a fixed visual-input budget. Gains are larger when the original per-frame resolution is preserved, highlighting the complementary roles of temporal coverage and spatial fidelity. Motivated by this finding, we propose AllocEmbed, an allocate-then-embed framework that reallocates a fixed visual-input budget across more frames. A lightweight allocator uses low-cost previews to assign frame-wise resolutions before the embedding backbone, preserving more detail where it most benefits retrieval while reducing visual cost elsewhere. We further introduce Retrieval-Driven Policy Optimization (RDPO), which learns the allocator directly from retrieval feedback using a rank-validated similarity gap and a confidence-guided efficiency incentive. Operating entirely before the backbone, AllocEmbed integrates with existing retrieval systems without modifying the embedding model or downstream pipeline. Experiments on the MMEB-V2 V-QA and V-RET tasks and our LongRet benchmark show that AllocEmbed achieves the best overall retrieval performance among the evaluated budget-matched methods and transfers across embedding backbones. Our code is publicly available at https://github.com/jinsong8/AllocEmbed.
Problem

Research questions and friction points this paper is trying to address.

video retrieval
visual-input budget
temporal coverage
frame importance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Allocate-then-embed
Temporal Coverage
Spatial Fidelity
Retrieval-Driven Policy Optimization
Visual Input Budget
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Song Jin
Song Jin
University of Wisconsin-Madison
nanomaterialsrenewable energysolid state chemistryspintronicsnanobiotechnology
Z
Zhongtao Jiang
Gaoling School of Artificial Intelligence, Renmin University of China
Chenglei Shen
Chenglei Shen
Gaoling School of Artificial Intelligence, Renmin University of China
Recommender systemsLarge language model
Huanxuan Liao
Huanxuan Liao
Institute of Automation, Chinese Academy of Sciences
Natural Language ProcessingLarge Language ModelLong Context Modeling
H
Haozhe Chi
Peking University
Z
Zhiwei Wang
Gaoling School of Artificial Intelligence, Renmin University of China
K
Kun Xu
Gaoling School of Artificial Intelligence, Renmin University of China
Y
Yong Liu
Gaoling School of Artificial Intelligence, Renmin University of China