Generalizable Robotic Insertion with World Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究解决了机器人在多样化装配任务中适应性差的问题,通过使用结合机器人本体感受信息和视觉观测的世界模型方法,提高了对未知物体的装配成功率。
📝 Abstract
Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.
Problem

Research questions and friction points this paper is trying to address.

Robotic Assembly
Generalizable Insertion
Diverse Parts
Zero-shot Success
Scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

world models
generalizable insertion
robotic assembly
data-driven
scalability