๐ค AI Summary
This study addresses the absence of functional articulation and view-inconsistent scene layouts in single-image 3D desktop reconstruction. To overcome these limitations, this work proposes a framework that generates physically interactive 3D tabletop scenes from a single image. Methodologically, it introduces GRAM-based joint modeling coupled with multimodal large language model-guided joint fitting to recover physical parameters, alongside a Progressive Semantic-Geometric Scene Registration (PSGSR) module to ensure layout consistency across viewpoints. Furthermore, a simulation-ready dataset comprising 100 scenes is constructed for evaluation. Experimental results demonstrate that the proposed approach achieves superior performance in visual fidelity, articulation quality, and physical plausibility. The practical effectiveness of the method is further corroborated through comprehensive user studies.
๐ Abstract
Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.