🤖 AI Summary
This study addresses the inconsistency between generated outputs and observed data in geometry, appearance, and pose during monocular or sparse multi-view 3D reconstruction. To overcome this, we propose a training-free test-time input alignment framework that faithfully aligns generative priors with observations through geometrically derived poses, visibility-biased attention, differentiable rendering guidance, and a lightweight decoder adapter. By incorporating external geometric support for shared appearance recovery of dynamic objects, our approach transcends the limitations of conventional optimization-based and frame-by-frame methods. Extensive evaluations on both synthetic and real-world benchmarks demonstrate that the proposed framework significantly improves pose estimation and reconstruction quality, outperforming existing optimization-based, image-to-3D, and video-to-4D approaches.
📝 Abstract
Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at https://facebookresearch.github.io/GenIA.