Agent Priors-guided Policy Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the information loss inherent in robotic skill composition and generalization, which constrains performance in few-shot scenarios. To overcome this limitation, the work proposes leveraging policy structural priors as a language interface to bridge skill acquisition and composition. During training, an agent automatically segments demonstration data to learn multiple prior-conditioned policies. At inference, the agent dynamically selects and composes these policies via the proposed interface, enabling zero-shot recombination of unseen skills. Extensive evaluations on the MetaWorld and ManiSkill benchmarks demonstrate that this approach significantly enhances out-of-distribution generalization performance while effectively supporting novel skill composition. Furthermore, ablation studies validate the critical role of the prior-based interface in achieving these improvements.
📝 Abstract
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.
Problem

Research questions and friction points this paper is trying to address.

compositional generalization
skill generalization
structural prior
policy learning
few-shot demonstration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structural Priors
Compositional Generalization
Skill Generalization
Policy Learning
Agent Priors
💼 Related Jobs
No related jobs found.
P
Puming Jiang
National University of Singapore
T
Tianrun Hu
National University of Singapore
H
Haozhe Du
National University of Singapore
Yibo Li
Yibo Li
National University of Singapore
LLM
Z
Zhiwei Xue
National University of Singapore
X
Xinhu Li
National University of Singapore
Harold Soh
Harold Soh
Associate Professor at National University of Singapore
Human Robot InteractionMachine LearningTactile PerceptionArtificial IntelligenceRobotics