KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人演示数据收集成本高的问题,KnowDemo通过从人类视频中提取结构化操作知识生成多样化的机器人演示。
📝 Abstract
Learning robot manipulation policies typically requires substantial demonstration data, which are costly to collect on real robots. Recent methods generate robot demonstrations from human videos by adapting recovered motion and validating the resulting trajectories in simulation. However, methods centered on motion-reference adaptation can limit behavioral diversity by retaining the demonstrated contact strategies and subtask orders, while insufficient understanding of task requirements and scene relations can reduce demonstration generation efficiency by generating invalid candidates. To address these limitations, we propose KnowDemo, a framework that uses structured manipulation knowledge from human videos to generate diverse robot demonstrations for a target workspace. To distinguish task requirements from demonstration-specific choices, we develop a knowledge extraction and reasoning module based on a vision-language model (VLM) that associates object and action descriptions with inferred task conditions, demonstration references, and permissible execution variations. To translate this knowledge into executable demonstrations, we resolve the descriptions against target-scene entities and geometry to guide candidate generation and screening before motion planning and simulation. The resulting demonstrations exhibit multimodal behavior through alternative contact strategies and valid subtask orders, with structured execution labels. Experiments demonstrate additional verified execution modes beyond a reference-only configuration and improved candidate planning success through task-guided grasp sampling. To validate the generated data for policy learning, we fine-tune the pretrained $π_{0.5}$ model on simulation data, achieving sim-to-real transfer across three tasks. Project page: https://zhiyuan-gao.github.io/knowdemo/
Problem

Research questions and friction points this paper is trying to address.

Robot Demonstration
Behavioral Diversity
Task Requirements
Scene Relations
Demonstration Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Extraction
Reasoning Module
Vision-Language Model
Structured Execution Labels
Task-Guided Grasp Sampling
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhiyuan Gao
University of Bremen
Y
Yanxiang Zhan
University of Bremen
M
Mohammad Khoshnazar
University of Bremen
J
Jeroen Schäfer
University of Bremen
Michael Beetz
Michael Beetz
Intitute for Artificial Intelligence, Computer Science Department, University of Bremen
cognitive roboticsAI-based Roboticsplan-based controlsemantic perceptionknowledge processing for Robots