🤖 AI Summary
Deformable object manipulation—particularly for cloth—faces challenges in state estimation and dynamics modeling due to high dimensionality, strong nonlinearity, and frequent self-occlusion. To address these, this paper introduces the first Transformer-based diffusion model framework specifically designed for deformable objects. Our method jointly achieves full-state reconstruction from sparse RGB-D observations and action-conditioned long-horizon dynamics prediction. It pioneers the integration of diffusion generative modeling into the perception–control closed loop for deformable objects, decoupling and co-optimizing perception and dynamics modeling to overcome the local-receptive-field limitations of graph neural networks (GNNs). This enables global deformation representation and high-fidelity state generation. Experiments demonstrate significantly improved state reconstruction accuracy and a tenfold reduction in long-horizon prediction error. Furthermore, our approach successfully executes multi-step cloth folding tasks on a real robotic platform.
📝 Abstract
Manipulating deformable objects like cloth is challenging due to their complex dynamics, near-infinite degrees of freedom, and frequent self-occlusions, which complicate state estimation and dynamics modeling. Prior work has struggled with robust cloth state estimation, while dynamics models, primarily based on Graph Neural Networks (GNNs), are limited by their locality. Inspired by recent advances in generative models, we hypothesize that these expressive models can effectively capture intricate cloth configurations and deformation patterns from data. Building on this insight, we propose a diffusion-based generative approach for both perception and dynamics modeling. Specifically, we formulate state estimation as reconstructing the full cloth state from sparse RGB-D observations conditioned on a canonical cloth mesh and dynamics modeling as predicting future states given the current state and robot actions. Leveraging a transformer-based diffusion model, our method achieves high-fidelity state reconstruction while reducing long-horizon dynamics prediction errors by an order of magnitude compared to GNN-based approaches. Integrated with model-predictive control (MPC), our framework successfully executes cloth folding on a real robotic system, demonstrating the potential of generative models for manipulation tasks with partial observability and complex dynamics.