SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects

๐Ÿ“… 2024-10-21
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 4
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenging problem of generating articulated 3D furniture models from a single input image. We propose the first end-to-end, single-image-driven method, overcoming key limitations of prior approachesโ€”namely, their reliance on multi-view or multi-state inputs and coarse-grained joint modeling. Methodologically, we introduce a geometry-motion joint diffusion model that unifies part-level fine-grained articulation synthesis with physically plausible motion constraints. We further design a cross-domain coarse-to-fine generation framework, explicit part connectivity graph modeling, and an abstraction-aware representation mechanism. Quantitatively and qualitatively, our method achieves state-of-the-art performance in realism, image fidelity, and structural reconstruction accuracy. It significantly advances practical utility and generalization capability for single-image articulated object modeling, establishing new benchmarks in this emerging domain.

Technology Category

Computer Vision: Diffusion Models for VisionCognitive Modeling & Cognitive Systems: Computational CreativityKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

User Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
๐Ÿ“ Abstract
We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation process. These limitations hinder the scalability and practicality for articulated object modeling. In this work, we propose a method to generate articulated objects from a single image. Observing the object in resting state from an arbitrary view, our method generates an articulated object that is visually consistent with the input image. To capture the ambiguity in part shape and motion posed by a single view of the object, we design a diffusion model that learns the plausible variations of objects in terms of geometry and kinematics. To tackle the complexity of generating structured data with attributes in multiple domains, we design a pipeline that produces articulated objects from high-level structure to geometric details in a coarse-to-fine manner, where we use a part connectivity graph and part abstraction as proxies. Our experiments show that our method outperforms the state-of-the-art in articulated object creation by a large margin in terms of the generated object realism, resemblance to the input image, and reconstruction quality.
Problem

Research questions and friction points this paper is trying to address.

Generating 3D articulated objects from single image
Overcoming limitations of multi-view or coarse control methods
Ensuring visual consistency and realism in generated objects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Single image generates 3D articulated objects.
Diffusion model learns geometry and kinematics variations.
Coarse-to-fine pipeline with part connectivity graph.
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
Simon Fraser University | Amii
J
Jiayi Liu
Simon Fraser University
D
Denys Iliash
Simon Fraser University
A
Angel X. Chang
Simon Fraser University, Canada-CIFAR AI Chair, Amii
M
M. Savva
Simon Fraser University
A
Ali Mahdavi-Amiri
Simon Fraser University