GenIA: Generative Reconstruction with Test-Time Input Alignment

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistency between generated outputs and observed data in geometry, appearance, and pose during monocular or sparse multi-view 3D reconstruction. To overcome this, we propose a training-free test-time input alignment framework that faithfully aligns generative priors with observations through geometrically derived poses, visibility-biased attention, differentiable rendering guidance, and a lightweight decoder adapter. By incorporating external geometric support for shared appearance recovery of dynamic objects, our approach transcends the limitations of conventional optimization-based and frame-by-frame methods. Extensive evaluations on both synthetic and real-world benchmarks demonstrate that the proposed framework significantly improves pose estimation and reconstruction quality, outperforming existing optimization-based, image-to-3D, and video-to-4D approaches.
📝 Abstract
Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at https://facebookresearch.github.io/GenIA.
Problem

Research questions and friction points this paper is trying to address.

3D reconstruction
generative 3D models
monocular observations
sparse multi-view
input alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Input Alignment
Generative 3D Reconstruction
Differentiable Rendering Guidance
Visibility-Biased Attention
Dynamic Object Canonicalization