CASE: Cost-Aware Stopping for Efficient Long-Video Agents

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe computational inefficiency of existing long-video agents caused by the absence of explicit termination mechanisms. To this end, this work proposes CASE, a plug-and-play termination framework that enables dynamic stopping decisions via policy-conditioned sequences. Specifically, CASE constructs a cost-aware objective over complete trajectories, integrating execution states with auxiliary multiple-choice evaluations, and employs Ridge regression to learn optimal decision intervals. Furthermore, it supports zero-shot transfer to effectively balance evidence accumulation costs against reasoning accuracy. Experimental results demonstrate that CASE achieves the optimal efficiency frontier by reducing average token consumption by 53.63% and runtime by 54.1%, while maintaining or even improving overall accuracy.
📝 Abstract
Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.
Problem

Research questions and friction points this paper is trying to address.

long-video agents
sequential stopping
cost-aware reasoning
efficiency
termination decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

cost-aware stopping
long-video agents
plug-in termination framework
sequential decision making
Pareto efficiency
🔎 Similar Papers
No similar papers found.
Y
Yiming Du
Peking University
C
Chenghao Liu
Peking University
Z
Zhiyuan Liu
Peking University
F
Fangxing Zheng
Peking University
Z
Zhao Wang
Peking University
J
Junnan Nie
Peking University
Songfang Huang
Songfang Huang
Peking University, Alibaba DAMO, IBM Research, The University of Edinburgh
LLMEmbodied AI