AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deficiency of multimodal large language models in precise 3D spatial reasoning and physical interaction for furniture assembly by proposing a zero-shot assembly framework. The framework introduces a novel anchor-guided boundary assembly decomposition strategy coupled with an iterative post-state feedback correction mechanism. By integrating instruction manual parsing, SE(3) pose updates, and simulation-based release testing, it achieves accurate mapping from semantic relations to 6D poses alongside physical feasibility verification. Experimental results demonstrate that the assembly tree F1 score improves to 62.8% with a significant reduction in Chamfer distance, while also revealing inherent model limitations in handling complex spatial relationships such as collisions.
📝 Abstract
Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components. To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly. It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree. Then, it uses iterative after-state feedback refinement to guide successive (SE(3)) updates and corrections, and validates their physical plausibility through simulation-based release tests. Experiments show that compared with the strongest prior baseline, AssemState improves F1 from 38.58\% to 62.80\% and Tree Exact Match from 28.24\% to 53.92\% for assembly-tree recovery. On 243 independently evaluated part-level operations, our proposed iterative refinement improves judge-accepted operations from 0 to 5.3\% and reduces mean Chamfer distance from 5.4111 to 1.7744. However, visually plausible candidate poses may still suffer from collision, floating, mirror-orientation errors, incomplete seating, and wrong-side attachment. These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.
Problem

Research questions and friction points this paper is trying to address.

3D spatial reasoning
furniture assembly
multimodal large language models
6D pose estimation
physical plausibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-shot Furniture Assembly
Multimodal Large Language Models
Assembly-tree Recovery
Iterative After-state Feedback Refinement
SE(3) Pose Estimation
🔎 Similar Papers
No similar papers found.