KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of behavioral cloning in generalizing to unseen object instances due to overfitting specific geometric appearances. To overcome this, we propose KeyGen, a framework that extracts canonical semantic keypoints from point clouds via unsupervised learning to construct object-centric structured representations. These representations are deeply integrated with visuomotor diffusion policies to ensure cross-instance geometric correspondence consistency and category-level generalization. Additionally, a planning-driven data generation pipeline is designed to establish a simulation benchmark. Experimental results demonstrate that KeyGen significantly outperforms existing methods under pose variations, scale changes, and real-world scenarios. Furthermore, it exhibits favorable scaling behavior with increasing demonstration data, enabling robust robotic manipulation.
📝 Abstract
Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
category-level generalization
behavior cloning
object-centric representations
policy transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised 3D Keypoints
Object-Centric Representations
Category-Level Generalization
Visuomotor Diffusion Policy
Behavior Cloning
💼 Related Jobs
No related jobs found.
S
Shuxin Cao
Georgia Institute of Technology
L
Liquan Wang
Georgia Institute of Technology
Masoud Moghani
Masoud Moghani
University of Toronto
Robot Learning
B
Benjamin Joffe
Georgia Institute of Technology
Animesh Garg
Animesh Garg
Georgia Institute of Technology, University of Toronto
Robotic ManipulationRobot LearningReinforcement LearningMachine LearningComputer Vision