FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出FunArt框架,通过单静态配置下的RGB-D观测构建含关节功能的3D场景图,解决机器人识别和理解可动部件及功能交互元素的问题。
📝 Abstract
To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.
Problem

Research questions and friction points this paper is trying to address.

articulated objects
functional 3D scene graphs
RGB-D observations
kinematic models
articulation-aware
Innovation

Methods, ideas, or system contributions that make the work stand out.

articulation-aware
functional 3D scene graphs
O-Voxel representation
sparse-compression VAE
query-based decoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dennis Rotondi
University of Stuttgart, Germany; International Max Planck Research School for Intelligent Systems
A
Abdelrhman Werby
University of Stuttgart, Germany; International Max Planck Research School for Intelligent Systems
Kai O. Arras
Kai O. Arras
Professor of Autonomous Systems
RoboticsSocial RoboticsHuman-Robot InteractionArtificial IntelligenceComputer Vision