🤖 AI Summary
To address poor generalization in robot imitation learning caused by limited scale and low diversity of complex manipulation demonstration data in real-world scenarios, this paper proposes a vision-language model (VLM)-driven two-stage hybrid data generation framework. First, a VLM parses expert demonstrations to decouple controllable action segments from object pose transformation segments. Second, object-centric pose modeling enables controlled pose perturbations, synthesizing large-scale, high-quality, and diverse training trajectories. The method requires no task-specific data formatting and jointly optimizes control fidelity and trajectory diversity. It constitutes the first data generation paradigm that deeply integrates VLM guidance with hybrid path planning and pose augmentation. Evaluated across seven manipulation tasks and their variants, the approach achieves an average success rate improvement of 5% over baselines, reaching 59.7% on the most challenging variants—outperforming the state-of-the-art MimicGen by 10.2 percentage points.
📝 Abstract
The acquisition of large-scale and diverse demonstration data are essential for improving robotic imitation learning generalization. However, generating such data for complex manipulations is challenging in real-world settings. We introduce HybridGen, an automated framework that integrates Vision-Language Model (VLM) and hybrid planning. HybridGen uses a two-stage pipeline: first, VLM to parse expert demonstrations, decomposing tasks into expert-dependent (object-centric pose transformations for precise control) and plannable segments (synthesizing diverse trajectories via path planning); second, pose transformations substantially expand the first-stage data. Crucially, HybridGen generates a large volume of training data without requiring specific data formats, making it broadly applicable to a wide range of imitation learning algorithms, a characteristic which we also demonstrate empirically across multiple algorithms. Evaluations across seven tasks and their variants demonstrate that agents trained with HybridGen achieve substantial performance and generalization gains, averaging a 5% improvement over state-of-the-art methods. Notably, in the most challenging task variants, HybridGen achieves significant improvement, reaching a 59.7% average success rate, significantly outperforming Mimicgen's 49.5%. These results demonstrating its effectiveness and practicality.