Foresight: planning future perception in streaming VLMs without retraining

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing streaming vision-language models, whose fixed computational paths struggle to adapt to dynamic scenarios and incur substantial overhead during online reconfiguration. We propose a foresighted dynamic computation paradigm that requires no retraining. By leveraging the model's intrinsic predictive capabilities, we construct a dual-stream architecture featuring Siamese LLMs for parallel planning and reasoning. Efficient and reliable training-free dynamic configuration is achieved through weight sharing, KV cache reuse, schema-guided decoding, and a lightweight differential update protocol. Experimental results demonstrate that our method surpasses the strongest baseline by 9.5% on the OmniPro metric and 6.7% on StreamingBench, with gains reaching 18.7% in late-evidence scenarios.
📝 Abstract
Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

streaming VLMs
dynamic computation
future anticipation
adaptive perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming VLMs
training-free planning
dual-stream architecture
dynamic computation
anticipatory inference
🔎 Similar Papers
No similar papers found.