RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited scalability of existing road surface reconstruction methods, which rely on per-scene optimization and trajectory-specific coverage. The authors propose a road structure-aware feedforward reconstruction framework that achieves scalable, test-time-optimization-free reconstruction for the first time. By integrating multi-view images, camera poses, and depth cues, the method leverages a geometric foundation model with a Gaussian head to predict pixel-aligned Gaussian attributes, followed by confidence-weighted mesh fusion in the road-aligned coordinate system. Key innovations include road-aligned planar fusion, category-aware grouping, and a curb-preserving mechanism, collectively enhancing structural completeness and semantic fidelity. Experiments demonstrate superior performance over state-of-the-art approaches in RGB and semantic bird’s-eye-view synthesis, elevation estimation, and novel view synthesis, underscoring the potential of geometric foundation models for road surface reconstruction.
📝 Abstract
Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road--sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird's-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.
Problem

Research questions and friction points this paper is trying to address.

road surface reconstruction
scalable reconstruction
per-scene optimization
large-scale mapping
autonomous driving
Innovation

Methods, ideas, or system contributions that make the work stand out.

road-structure-aware
feed-forward reconstruction
Gaussian surface representation
geometric foundation model
confidence-weighted fusion
🔎 Similar Papers
No similar papers found.