GeoWM: Efficient Direct World Modeling in Explicit Geometry

๐Ÿ“… 2026-10-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitations of conventional world models, specifically their lack of explicit geometric modeling and the error accumulation and high computational costs associated with recursive prediction. To overcome these issues, this work proposes GeoWM, which leverages a geometry foundation model to transform RGB sequences into geometric history representations. By integrating lightweight camera motion estimation with projective geometric priors, GeoWM employs a flow-matching Transformer to directly predict future scene geometry, establishing the first non-recursive paradigm for direct geometric prediction. Experimental results demonstrate that GeoWM surpasses existing models in depth, pose, and 3D geometry prediction across four datasets. Furthermore, it substantially reduces inference latency over long temporal horizons, achieving efficient and accurate generation of future scene geometry.
๐Ÿ“ Abstract
Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to predict future images or latent representations of the environment and subsequently recover geometry from these predictions. However, this paradigm does not explicitly model geometric structure and typically relies on recursive rollouts to reach longer prediction horizons, leading to error accumulation and increasing computational cost. To address these limitations, we present GeoWM, a geometry world model that directly forecasts future scene geometry at specified future horizons without recursive rollout. The key idea is to leverage a geometry foundation model to transform observed RGB frames into a geometric history, which conditions a flow-matching transformer to predict the scene geometry at a specified future horizon. We further show that a lightweight camera-motion predictor can accurately estimate the future viewpoint, and that projecting the observed geometry into the predicted viewpoint provides an effective geometric prior for future geometry forecasting. Extensive experiments on four datasets spanning urban driving, aerial flight, and dynamic manipulation demonstrate that GeoWM outperforms the evaluated world models in forecasting depth, camera pose, and 3D scene geometry, while substantially reducing inference time at longer horizons.
Problem

Research questions and friction points this paper is trying to address.

World Models
3D Scene Geometry
Autonomous Driving
Error Accumulation
Future Forecasting
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Model
Explicit Geometry Modeling
Flow Matching
Camera Motion Prediction
Autonomous Driving
๐Ÿ”Ž Similar Papers
No similar papers found.