🤖 AI Summary
This study addresses the prohibitive computational cost of processing all frames in long video understanding and the reliance on expert-crafted agents. We propose an automated framework that generates efficient video question-answering agents through iterative evolution. Methodologically, Monte Carlo Tree Search is introduced to circumvent local optima, combined with uncertainty-aware multi-fidelity verification to reduce evaluation overhead, alongside a mixture-of-experts routing mechanism to enhance inference efficiency. Experimental results demonstrate that the proposed approach establishes new state-of-the-art performance across multiple benchmarks, surpassing the strongest handcrafted agent by 11.2 points while substantially reducing frame consumption. Ultimately, this work achieves low-cost, high-performance long video understanding without manual intervention.
📝 Abstract
Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped around a frozen VLM, address this by observing the video selectively, yet existing harnesses are hand-crafted by experts through slow build-and-test cycles. We propose VidHarness, a framework that automates harness design for cost-efficient long video understanding, in which a harness proposer iteratively evolves harnesses based on execution feedback from an evolution environment. To escape the local optima of greedy refinement, we organize the evolution as Monte Carlo tree search (MCTS), and to reduce the evaluation cost, we integrate uncertainty-aware multi-fidelity validation, which screens new harnesses on a few questions and promotes only the promising ones. Since the best harness varies with the frame budget, we further introduce a mixture-of-harness that routes each question to a harness specialized for its budget. VidHarness sets new state-of-the-art results on LongVideoBench, Video-MME, and Video-Holmes, outperforms the strongest hand-crafted video agent by up to $11.2$ points, and generalizes to the knowledge-intensive benchmarks Video-MMMU and MMVU with fewer than half of the frames of uniform sampling.