GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse observation integration and occlusion reasoning in 3D object reconstruction from unconstrained images by proposing the GATOR framework. This work introduces a novel architecture that fuses generative and agentic paradigms, incorporating a local modality mixer and text-guided semantic conditioning for asset generation. An "observe-edit-render-review" closed-loop mechanism is established to enable iterative refinement via multimodal agents. Furthermore, the framework integrates cross-view reasoning, patch-aligned feature fusion, and stage-specific adapters. Experimental results demonstrate that GATOR achieves high-fidelity geometric and texture reconstruction alongside accurate pose recovery across both synthetic objects and complex scenes. By effectively disentangling foreground targets from backgrounds, the proposed approach significantly enhances reconstruction efficiency and simulation readiness in heavily occluded scenarios.
📝 Abstract
Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/
Problem

Research questions and friction points this paper is trying to address.

3D object reconstruction
casual images
occlusion inference
scene-aligned pose
sparse observations
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Object Reconstruction
Agentic Framework
Local Modality Mixer
Text-guided Semantic Conditioning
Edit-Render-Review Loop
💼 Related Jobs
No related jobs found.