RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of cross-view consistency and unified querying in multi-camera robotic video prediction by proposing a two-stage architecture that lifts synchronized multi-camera videos into a unified metric 4D Gaussian field. The method integrates robot kinematics with geometric priors through feed-forward 4D Gaussian Splatting, an action-conditioned video world model, and cross-view evidence fusion. By jointly optimizing scene appearance and structure under temporal displacement constraints, it achieves spatiotemporally queryable 4D reconstruction. Experimental results demonstrate significant improvements, yielding a 2.15 dB increase in novel view PSNR, a 47% reduction in depth error, and a 61% decrease in robot displacement error.
📝 Abstract
Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.
Problem

Research questions and friction points this paper is trying to address.

video world models
4D reconstruction
cross-view consistency
multi-camera rollouts
robot world model
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Gaussian Splatting
Robot World Model
Feed-Forward Reconstruction
Cross-View Consistency
Articulated Kinematics
💼 Related Jobs
No related jobs found.
J
Jin Hyun Kim
School of Mechanical Engineering, Korea University
Min Young Kim
Min Young Kim
College of AI Convergence, Dongguk University
S
Soohwan Song
College of AI Convergence, Dongguk University
Daekyum Kim
Daekyum Kim
Assistant Professor, Korea University
RoboticsArtificial IntelligenceComputer VisionWearablesSoft Robotics