ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image

๐Ÿ“… 2026-10-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the absence of functional articulation and view-inconsistent scene layouts in single-image 3D desktop reconstruction. To overcome these limitations, this work proposes a framework that generates physically interactive 3D tabletop scenes from a single image. Methodologically, it introduces GRAM-based joint modeling coupled with multimodal large language model-guided joint fitting to recover physical parameters, alongside a Progressive Semantic-Geometric Scene Registration (PSGSR) module to ensure layout consistency across viewpoints. Furthermore, a simulation-ready dataset comprising 100 scenes is constructed for evaluation. Experimental results demonstrate that the proposed approach achieves superior performance in visual fidelity, articulation quality, and physical plausibility. The practical effectiveness of the method is further corroborated through comprehensive user studies.
๐Ÿ“ Abstract
Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.
Problem

Research questions and friction points this paper is trying to address.

single-image 3D reconstruction
articulated objects
tabletop scenes
scene layout
embodied agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Single-image 3D reconstruction
Articulated object modeling
Multimodal large language model
Scene registration
Embodied AI
๐Ÿ”Ž Similar Papers
K
Kai Lv
School of Computer Science, Wuhan University
Y
Yibo Yin
ร‰cole Polytechnique Fรฉdรฉrale de Lausanne (EPFL), Switzerland
L
Lijun Guo
School of Computer Science, Wuhan University
Heng Fan
Heng Fan
Assistant Professor, University of North Texas
Computer VisionMachine LearningArtificial Intelligence
Kaihao Zhang
Kaihao Zhang
Australian National University
Deep learningComputer vision
X
Xingping Dong
School of Computer Science, Wuhan University