DynStream: Online Streaming 4D Gaussian Reconstruction of Dynamic Worlds from Unposed Video

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of simultaneously achieving online processing and high-fidelity rendering for pose-free dynamic scenes in long video streams. We propose a streaming reconstruction framework based on 4D Gaussian Splatting, which employs local temporal window reconstruction combined with an incremental alignment and fusion strategy. By jointly enforcing cross-window geometric consistency and modeling time-varying content, our method constructs a globally consistent 4D scene representation without requiring per-scene offline optimization. Experiments demonstrate that the proposed approach achieves high-fidelity online reconstruction and novel view synthesis across diverse indoor and outdoor dynamic scenes. It effectively overcomes the limitations of point cloud-based methods regarding dense geometry and rendering quality, attaining state-of-the-art performance in the field.
📝 Abstract
Online reconstruction of dynamic 4D scenes from long, unposed streaming videos requires both continuous processing and photorealistic rendering, which existing methods struggle to achieve simultaneously. Existing feed-forward Gaussian methods are restricted to offline processing, whereas online point-cloud approaches struggle to maintain dense geometry and high-fidelity rendering. We present DynStream, a framework for streaming 4D Gaussian reconstruction from long, unposed videos. Given a continuous video stream, DynStream reconstructs the scene within local temporal windows and incrementally aligns and fuses these local reconstructions into a globally consistent scene, enabling online 4D reconstruction without per-scene optimization. By jointly enforcing cross-window geometric consistency and modeling time-varying scene content, DynStream supports efficient reconstruction and photorealistic rendering over extended video streams. Experiments demonstrate that DynStream enables high-fidelity online dynamic reconstruction and rendering from long video streams, achieving state-of-the-art performance across diverse dynamic indoor and outdoor scenes.
Problem

Research questions and friction points this paper is trying to address.

4D reconstruction
streaming video
unposed video
online processing
photorealistic rendering
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Gaussian Reconstruction
Streaming Video
Unposed Video
Online Reconstruction
Incremental Fusion
D
Dingwei Xian
Wangxuan Institute of Computer Technology, Peking University
Xiaoyu Zhou
Xiaoyu Zhou
Peking University
Computer VisionAutonomous DrivingAI Security
Y
Yajiao Xiong
Wangxuan Institute of Computer Technology, Peking University
Y
Yongtao Wang
Wangxuan Institute of Computer Technology, Peking University, VGI Labs Co., Ltd.
Ming-Hsuan Yang
Ming-Hsuan Yang
University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence