POIL: Point-based One-Shot Imitation Learning with Stable Dynamical Systems

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of cross-object trajectory transfer and insufficient execution robustness in one-shot imitation learning by proposing a point-set-based policy framework. Methodologically, it leverages multimodal large models to localize functional components and achieves trajectory transfer via point-set correspondence. Theoretically, stable dynamics are extended from SE(3) poses to point-set spaces without requiring prior 3D models, with closed-loop control ensuring global convergence at the end-effector. Simulation and real-world experiments demonstrate that the proposed method successfully transfers demonstrations across object categories and poses while exhibiting strong disturbance rejection capabilities.
📝 Abstract
We present POIL, a point-based one-shot imitation learning framework with stable dynamical systems. While one-shot imitation avoids collecting extensive demonstrations, successful one-shot manipulation requires not only transferring a demonstrated trajectory to a novel object but also executing it robustly under changing scene conditions, grasp configurations, and external disturbances. POIL addresses both problems through a shared representation: a set of 3D points on the object's functional part, used jointly for trajectory transfer and closed-loop execution. The one-shot transfer from the demonstrated trajectory is enabled with point correspondences. POIL grounds the shared functional part with a multi-modal large language model, and transfers the trajectory across viewpoint, pose, and object category changes. During execution, multi-view tracking observes the same points online, and Point-set BCSDM drives them in closed loop by projecting per-point velocities onto a single rigid-body twist computed from the tracked points alone. This extends stable dynamical models from an SE(3) pose to a point set without requiring a known 3D model or pose estimator. We show that at the goal the controller becomes a gradient flow on the classical SO(3) potential, so its terminal phase inherits the almost-global convergence of that potential under a rigid-object assumption. Across simulation and real-robot experiments, POIL transfers a single demonstration across object category, grasp pose, and goal geometry, while recovering from external disturbances during execution. Project page: https://sangminkim-99.github.io/poil
Problem

Research questions and friction points this paper is trying to address.

One-shot imitation learning
Trajectory transfer
Robust execution
Stable dynamical systems
Point-based representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

One-Shot Imitation Learning
Stable Dynamical Systems
Point-set Representation
Multi-modal Large Language Model
Closed-loop Control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sang Min Kim
Department of Electrical and Computer Engineering, Seoul National University, Seoul, Republic of Korea
J
Jinwoo Seo
Department of Electrical and Computer Engineering, Seoul National University, Seoul, Republic of Korea
H
Hyeongjun Heo
Department of Electrical and Computer Engineering, Seoul National University, Seoul, Republic of Korea
Junho Lee
Junho Lee
Seoul National University
Generative ModelRepresentation LearningVideo Understanding
Yonghyeon Lee
Yonghyeon Lee
Postdoctoral Associate @ MIT
Geometric Data AnalysisMachine LearningRobotics
Young Min Kim
Young Min Kim
Department of Electrical and Computer Engineering, Seoul National University
Computer GraphicsComputer VisionComputational GeometryMachine LearningArtificial