Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

📅 2025-06-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the alignment challenge in joint novel-view image and geometry generation under sparse reference images and coarse geometric priors. We propose the first diffusion-based image-geometry co-generation framework, formulated as cross-modal collaborative inpainting: an off-the-shelf pre-trained geometry predictor provides initial depth/normal maps; a novel cross-modal attention distillation mechanism enforces structural alignment between image and geometry diffusion branches; a proximity-aware mesh-conditioning strategy fuses depth and normal cues to suppress geometric noise; and geometry-guided warped inpainting enhances rendering consistency. Our method achieves high-fidelity extrapolative novel-view synthesis and synchronized geometry reconstruction on unseen scenes. It sets new state-of-the-art interpolation quality, outputs color-aligned point clouds with consistent geometry, and supports end-to-end 3D scene completion.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Generative Models & AutoencodersIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
We introduce a diffusion-based framework that performs aligned novel view image and geometry generation via a warping-and-inpainting methodology. Unlike prior methods that require dense posed images or pose-embedded generative models limited to in-domain views, our method leverages off-the-shelf geometry predictors to predict partial geometries viewed from reference images, and formulates novel-view synthesis as an inpainting task for both image and geometry. To ensure accurate alignment between generated images and geometry, we propose cross-modal attention distillation, where attention maps from the image diffusion branch are injected into a parallel geometry diffusion branch during both training and inference. This multi-task approach achieves synergistic effects, facilitating geometrically robust image synthesis as well as well-defined geometry prediction. We further introduce proximity-based mesh conditioning to integrate depth and normal cues, interpolating between point cloud and filtering erroneously predicted geometry from influencing the generation process. Empirically, our method achieves high-fidelity extrapolative view synthesis on both image and geometry across a range of unseen scenes, delivers competitive reconstruction quality under interpolation settings, and produces geometrically aligned colored point clouds for comprehensive 3D completion. Project page is available at https://cvlab-kaist.github.io/MoAI.
Problem

Research questions and friction points this paper is trying to address.

Aligned novel view image and geometry synthesis
Cross-modal attention for accurate alignment
High-fidelity extrapolative view synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion-based warping-and-inpainting for view synthesis
Cross-modal attention distillation for alignment
Proximity-based mesh conditioning with depth cues
🔎 Similar Papers
No similar papers found.