HybridGen: VLM-Guided Hybrid Planning for Scalable Data Generation of Imitation Learning

📅 2025-03-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address poor generalization in robot imitation learning caused by limited scale and low diversity of complex manipulation demonstration data in real-world scenarios, this paper proposes a vision-language model (VLM)-driven two-stage hybrid data generation framework. First, a VLM parses expert demonstrations to decouple controllable action segments from object pose transformation segments. Second, object-centric pose modeling enables controlled pose perturbations, synthesizing large-scale, high-quality, and diverse training trajectories. The method requires no task-specific data formatting and jointly optimizes control fidelity and trajectory diversity. It constitutes the first data generation paradigm that deeply integrates VLM guidance with hybrid path planning and pose augmentation. Evaluated across seven manipulation tasks and their variants, the approach achieves an average success rate improvement of 5% over baselines, reaching 59.7% on the most challenging variants—outperforming the state-of-the-art MimicGen by 10.2 percentage points.

Technology Category

Intelligent Robots: ManipulationMachine Learning: Imitation Learning & Inverse Reinforcement LearningComputer Vision: Language and Vision

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
📝 Abstract
The acquisition of large-scale and diverse demonstration data are essential for improving robotic imitation learning generalization. However, generating such data for complex manipulations is challenging in real-world settings. We introduce HybridGen, an automated framework that integrates Vision-Language Model (VLM) and hybrid planning. HybridGen uses a two-stage pipeline: first, VLM to parse expert demonstrations, decomposing tasks into expert-dependent (object-centric pose transformations for precise control) and plannable segments (synthesizing diverse trajectories via path planning); second, pose transformations substantially expand the first-stage data. Crucially, HybridGen generates a large volume of training data without requiring specific data formats, making it broadly applicable to a wide range of imitation learning algorithms, a characteristic which we also demonstrate empirically across multiple algorithms. Evaluations across seven tasks and their variants demonstrate that agents trained with HybridGen achieve substantial performance and generalization gains, averaging a 5% improvement over state-of-the-art methods. Notably, in the most challenging task variants, HybridGen achieves significant improvement, reaching a 59.7% average success rate, significantly outperforming Mimicgen's 49.5%. These results demonstrating its effectiveness and practicality.
Problem

Research questions and friction points this paper is trying to address.

Generates large-scale diverse data for robotic imitation learning
Integrates Vision-Language Model and hybrid planning for task decomposition
Improves performance and generalization in complex manipulation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates Vision-Language Model for task parsing
Uses hybrid planning for diverse trajectory synthesis
Expands data via pose transformations without specific formats
🔎 Similar Papers
No similar papers found.