🤖 AI Summary
This study addresses the limitation of behavioral cloning in generalizing to unseen object instances due to overfitting specific geometric appearances. To overcome this, we propose KeyGen, a framework that extracts canonical semantic keypoints from point clouds via unsupervised learning to construct object-centric structured representations. These representations are deeply integrated with visuomotor diffusion policies to ensure cross-instance geometric correspondence consistency and category-level generalization. Additionally, a planning-driven data generation pipeline is designed to establish a simulation benchmark. Experimental results demonstrate that KeyGen significantly outperforms existing methods under pose variations, scale changes, and real-world scenarios. Furthermore, it exhibits favorable scaling behavior with increasing demonstration data, enabling robust robotic manipulation.
📝 Abstract
Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.