PointCast: One World Model for Rigid, Articulated, and Deformable Object Manipulation

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究提出PointCast模型,通过预测3D点集轨迹来解决刚性、关节和可变形物体的操纵问题,使用扩散变换器作为骨干网络。
📝 Abstract
World models are useful for robotic manipulation because robots can predict how actions change the states of objects before executing them. We present PointCast, a point-set world model that spans rigid, articulated, and deformable object manipulation. Its state is a set of 3D points on the object and the end-effector, mesh-free and topology-agnostic. Each point keeps its identity and is supervised on its own trajectory, which teaches the model where every point goes rather than only the shape the points form. Its backbone is a diffusion transformer that denoises a short window of future point positions, conditioned on the points' recent history and the commanded end-effector motion. The backbone's attention alternates between local and global, and cross-attention to the end-effector carries the coupling. This one architecture at 19.8M parameters and one training recipe cover four regimes, rigid objects, cloth, rope, and multi-joint cabinets, with a separate checkpoint trained for each. Trained on randomized simulation and scored against four baselines on the same metric, it is best on three of four regimes and second on rigid. Trained on a real-world robot teleoperation dataset, it has the lowest mean error in four of its six categories, is second in the other two, and improves on the dataset's own model in all six; zero-shot, its simulation checkpoints are best on two of four captures. Frozen inside sampling-based model-predictive control at one network evaluation per window, it plans four simulated tasks over 64 episodes, competitive with or outperforming every baseline on each. Project website at https://pointcast-wm.github.io.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
world model
rigid objects
deformable objects
articulated objects
Innovation

Methods, ideas, or system contributions that make the work stand out.

point-set world model
diffusion transformer
rigid and deformable object manipulation
trajectory supervision
💼 Related Jobs
No related jobs found.
Hantao Ye
Hantao Ye
Ph.D. Student, University of Minnesota Twin Cities
Robotics
R
Ross Worobel
University of Minnesota, Minneapolis, MN 55455, USA
Z
Zhuoli Xie
University of Minnesota, Minneapolis, MN 55455, USA
M
Mingen Li
University of Minnesota, Minneapolis, MN 55455, USA
Houjian Yu
Houjian Yu
Amazon, University of Minnesota
RoboticsComputer Vision
Y
Youngjin Hong
University of Minnesota, Minneapolis, MN 55455, USA
Changhyun Choi
Changhyun Choi
Assistant Professor, University of Minnesota Twin Cities
robot visionmanipulation