Sample-Efficient Behavior Cloning Using General Domain Knowledge

📅 2025-01-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address low sample efficiency and poor cross-environment generalization in behavioral cloning, this paper proposes the Knowledge-Guided Imitation Model (KIM). KIM automatically transforms informal, natural-language expert knowledge into semantically precise, differentiable structured policies via large language models, then jointly optimizes these policies with a minimal set of demonstration trajectories (only five). Crucially, KIM enables end-to-end encoding of informal domain knowledge into learnable policy structures—the first such approach—thereby realizing knowledge- and data-coordinated imitation learning. Evaluated on lunar landing and racing control tasks, KIM significantly outperforms knowledge-free baselines while demonstrating robustness to action noise. These results validate its high sample efficiency and strong generalization capability across diverse environments.

Technology Category

Machine Learning: Imitation Learning & Inverse Reinforcement LearningHumans and AI: Human-Aware Planning and Behavior PredictionIntelligent Robots: Behavior Learning & Control

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Behavior cloning has shown success in many sequential decision-making tasks by learning from expert demonstrations, yet they can be very sample inefficient and fail to generalize to unseen scenarios. One approach to these problems is to introduce general domain knowledge, such that the policy can focus on the essential features and may generalize to unseen states by applying that knowledge. Although this knowledge is easy to acquire from the experts, it is hard to be combined with learning from individual examples due to the lack of semantic structure in neural networks and the time-consuming nature of feature engineering. To enable learning from both general knowledge and specific demonstration trajectories, we use a large language model's coding capability to instantiate a policy structure based on expert domain knowledge expressed in natural language and tune the parameters in the policy with demonstrations. We name this approach the Knowledge Informed Model (KIM) as the structure reflects the semantics of expert knowledge. In our experiments with lunar lander and car racing tasks, our approach learns to solve the tasks with as few as 5 demonstrations and is robust to action noise, outperforming the baseline model without domain knowledge. This indicates that with the help of large language models, we can incorporate domain knowledge into the structure of the policy, increasing sample efficiency for behavior cloning.
Problem

Research questions and friction points this paper is trying to address.

Behavioral Cloning
Limited Examples
Expert Knowledge Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge-Guided Models
Few-Shot Learning
Expert Knowledge Integration
🔎 Similar Papers
No similar papers found.