🤖 AI Summary
This study addresses the limitations of existing models for embodied agents performing real-time 3D spatial reasoning over video streams, which typically lack explicit 3D representations or rely on offline processing. To this end, we propose SpaTime, a streaming vision-language model that pioneers the integration of causal geometric tokens into a streaming VLM architecture to enable real-time 3D reasoning. Furthermore, a differentiable expected response time loss function is designed to supervise the model's answering timing. We also construct a dedicated benchmark, StreamVSTI-Bench, for evaluation. Experimental results demonstrate that SpaTime achieves 49.2% accuracy on this benchmark and reduces the average response time error by 66% compared to the strongest baseline, thereby validating its effectiveness for real-time 3D spatial reasoning in embodied AI applications.
📝 Abstract
Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.