OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent conflict between real-time perception and persistent memory in streaming video large language models by proposing an active hierarchical captioning memory and state transition learning framework. The method jointly optimizes query-agnostic evidence recording and task-specific responses through a shared active generation process. Furthermore, it constructs a million-scale streaming interaction dataset and employs a sparse state supervision strategy to unify perception, memory, and response interfaces, thereby significantly enhancing both training and inference efficiency. Experimental results demonstrate that the proposed 4B-parameter model achieves state-of-the-art performance across eight benchmarks. These findings validate that the active generation mechanism effectively strengthens historical question-answering capabilities without compromising real-time perceptual acuity.
📝 Abstract
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
Problem

Research questions and friction points this paper is trying to address.

streaming video interaction
factual memory
real-time perception
proactive response
video LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Video LLM
Proactive Hierarchical Caption Memory
Proactive State Transition Learning
Streaming Data Synthesis
OneStreamer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiangyu Zeng
NJU
Y
Yuandong Yang
NJU
Z
Zhiqiu Zhang
PJLAB
Yuhan Zhu
Yuhan Zhu
Nanjing University, Shanghai AI Lab
Computer VisionVision-Language ModelsVideo Understanding
Xinhao Li
Xinhao Li
Nanjing University
Video UnderstandingMultimodal LLMVision-Language Learning
Q
Qingyi Si
JD
D
Dingyu Yao
CAS
C
Changlian Ma
NJU
H
Haoran Chen
NJU
X
Xinyu Chen
NJU
Y
Yansong Shi
USTC
J
Junhao Zhou
CAS
Y
Yifei Li
THU
J
Jun Zhang
NJU
C
Chuanyu Qin
CAS
Chenxu Yang
Chenxu Yang
Institute of Information Engineering, Chinese Academy of Sciences
NLPDialogue Generation
Xinlei Yu
Xinlei Yu
Beijing University of Posts and Telecommunications
Stochastic Geometry
Kun Ouyang
Kun Ouyang
National University of Singapore
human mobilitymachine learning
Y
Yuchen Shao
CAS
Q
Qianshan Wei
CAS
C
Changhai Zhou
FDU
J
Jun Gao
ZJU
J
Jiaqi Wang
JD
L
Limin Wang
NJU