Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing 3D generation methods struggle to achieve controllable and geometrically consistent shape completion under multi-view, occluded, or editing scenarios, and lack a unified framework to support diverse conditional inputs. To overcome these limitations, we propose Axolotl3D, the first unified multimodal 3D generation framework that jointly models images, visibility masks, camera parameters, and partial point clouds within a shared 3D coordinate system, enabling occlusion-aware and cross-modal reasoning. Built upon a diffusion architecture, Axolotl3D integrates geometric anchors and multi-view alignment, and employs a unified training strategy to synthesize diverse conditions from large-scale data. Experiments demonstrate that Axolotl3D achieves state-of-the-art performance on both Toys4K and OmniObject3D benchmarks, excelling in tasks involving occlusion handling, real-world reconstruction, and geometrically consistent editing.
📝 Abstract
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume complete visibility and single-view inputs, limiting applicability in multi-view, occluded, or editing scenarios. Although prior works address these challenges individually, they lack a unified framework for controllable 3D completion under diverse conditioning signals. We present Axolotl3D, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud serves as a geometric anchor promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment in a shared 3D coordinate system. A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning. Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strong results in real-world reconstruction and geometry-consistent editing.
Problem

Research questions and friction points this paper is trying to address.

3D shape completion
occlusion
multi-view
conditional generation
unified framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D shape completion
multi-modal conditioning
occlusion-aware generation
geometric anchor
unified framework
🔎 Similar Papers