InteractionBench: A Real-Time Interaction Benchmark for Streaming Video Systems

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely determining when to respond and when to remain silent in streaming video systems, noting that existing offline evaluations fail to predict online behavior. To this end, it constructs a real-time interactive benchmark encompassing queries, event triggers, and continuous updates, and proposes the first silence-compliance metric designed for complete systems comprising models, memory, and controllers. This metric is quantitatively evaluated using negative test sets paired with near-miss event techniques. Experiments reveal that open-source systems perform poorly in silence compliance, with none resolving more than one-third of near-miss scenarios, thereby demonstrating the substantial cost of response restraint. Ultimately, this work fills a critical evaluation gap regarding temporal accuracy and silence decision-making in the real-time stream processing of multimodal large language models.
📝 Abstract
A video assistant must speak when its instruction warrants a response and stay silent otherwise. We introduce a benchmark that evaluates this decision for the complete system of model, memory, and response controller. InteractionBench covers query responses, event triggers, and ongoing updates in 1,060 interactions over 812 videos, with 69 negative streams and 53 suites that pair counted events with look-alike near misses. It scores content accuracy, timing accuracy, and silence compliance on the video clock. Timely speech costs silence across systems. Polled Qwen3-VL-8B reaches 77.8 timing accuracy but 10.9 silence compliance. A native real-time interaction system reaches 29.2 silence compliance at 66.8 timing accuracy, yet emits on 89.9% of negative streams. No open-weight system clears a third of the near-miss suites. Fewer replies help only when chosen, as random deletion merely trades timing for silence. Offline scores miss these failures and mispredict online behavior. Adding restraint is costly, as the native system's controller adds little by itself and agentic systems add it only at about 30 s per poll.Project page: https://www.enxinsong.com/projects/interactionbench/ Code: https://github.com/Espere-1119-Song/InteractionBench Data: https://huggingface.co/datasets/InteractionBench/InteractionBench
Problem

Research questions and friction points this paper is trying to address.

Streaming Video Systems
Real-Time Interaction
Silence Compliance
Timing Accuracy
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Real-time Interaction Benchmark
Streaming Video Systems
Silence Compliance
Timing Accuracy
Response Controller