M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of generative novel view synthesis methods that overlook LiDAR geometric information and struggle to simultaneously preserve visual appearance and metric 3D structure. To this end, we propose M3GD, a framework that aligns camera and point cloud foundation models by leveraging the shared spatial structure of projected features, eliminating the need for pretrained cross-modal transformers. Furthermore, lightweight residual adapters are designed to inject explicit geometric statistics into a multimodal flow-matching diffusion generator. Experiments on the GrandTour dataset demonstrate that M3GD significantly outperforms image-only baselines, yielding substantial improvements in both RGB and depth synthesis quality. Real-world deployment validates its practical feasibility, while the Euler integration steps enable flexible trade-offs between generation fidelity and computational cost.
📝 Abstract
Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.
Problem

Research questions and friction points this paper is trying to address.

Novel View Synthesis
Camera-LiDAR Fusion
Multi-Modal Representation
Robotic Perception
3D Geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Modal Novel View Synthesis
Camera-LiDAR Fusion
Geometric Diffusion
Flow Matching
Foundation Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yang Zhou
Yang Zhou
New York University
RoboticsComputer Vision
J
Jiuhong Xiao
Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA
S
Shizhao Ye
Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA
L
Long Quang
U.S. Army Combat Capabilities Development Command, Army Research Laboratory, Adelphi, MD 20783, USA
Carlos Nieto-Granda
Carlos Nieto-Granda
U.S. Army Research Laboratory (ARL)
Multi-robot and multi-agent systemsAutonomous Navigation & ExplorationSLAMHuman-Robot Teams
Giuseppe Loianno
Giuseppe Loianno
UC Berkeley
RoboticsMAVsVisionSensor Fusion