Score
Designs and implements generative models and pipelines that synthesize motion trajectories and demonstrations — from one-step action outputs and few-step action sequences to full multi-step task trajectories — conditioned on inputs such as language, scene state, or predicted velocity fields and optimized for executability (efficient, low-latency, robot-ready control sequences). Builds simulation-driven demonstrators and data-augmentation workflows that convert latent or simulated outputs into executable controls, ensure demonstrability within synthesized scenes, align synthetic and real distributions, and provide supervisory signals for training and distillation.
The field of motion generation lacks a systematic survey grounded in generative methodology. To address this gap, we propose the first deep taxonomy centered on generative strategies, synthesizing state-of-the-art works from top-tier conferences (CVPR, ICCV, SIGGRAPH, CoRL) since 2023. Our framework uniformly analyzes four dominant paradigms—GANs, autoencoders, autoregressive models, and diffusion models—across three dimensions: architectural design, conditional modeling mechanisms, and evaluation protocols for motion sequence synthesis. We consolidate widely adopted datasets and metrics, identifying key challenges including motion coherence, physical plausibility, and cross-domain generalization. By establishing a comprehensive, comparable analytical benchmark with well-defined dimensions, our survey significantly enhances methodological comparability and facilitates precise problem diagnosis. This work serves as a foundational reference for researchers advancing generative motion modeling.
To address the dual bottlenecks of high-cost expert demonstrations and the limitation of existing generative control methods to quasi-static scenarios in fast dynamic robotic tasks, this paper proposes Generative Predictive Control (GPC). GPC employs flow matching to construct a generative policy model trained exclusively on simulated data—eliminating reliance on human teleoperation. It establishes, for the first time, a theoretical connection between sampling-based predictive control and generative modeling. Moreover, GPC supports online warm-start optimization, ensuring temporal consistency and millisecond-level feedback. Evaluated on high-speed, non-quasi-static tasks—including agile grasping and bouncing locomotion—GPC demonstrates real-time performance, stability, and strong generalization. This work introduces a scalable, demonstration-free paradigm for general-purpose robotic policy learning.
This work proposes Dream.exe, a framework that, for the first time, validates the executability of outputs from video generation models through real-world physical execution. Addressing the question of whether generated manipulation videos adhere to physical laws and can be realized by robots, the method integrates video generation models, trajectory extraction algorithms, and a physics simulator to establish an end-to-end video-to-execution evaluation pipeline. Evaluation across 101 manipulation tasks on eight model families reveals that certain models achieve notably high execution success rates, indicating their acquisition of effective physical priors from large-scale training data. The study further uncovers a significant disconnect between visual fidelity and executability, thereby advocating for new evaluation dimensions that go beyond purely perceptual metrics.
Robotic policies often exhibit limited generalization to novel behaviors and unseen environments. Method: This paper proposes DreamGen, a four-stage framework that leverages a video world model to generate embodied-consistent synthetic neural trajectories (i.e., robot action videos), followed by latent-action modeling or inverse dynamics modeling to recover high-fidelity pseudo-action labels. It is the first work to adapt image-to-video generation models for embodied agents, establishing a “neural-trajectory-driven generalization” paradigm. Contribution/Results: We introduce DreamGen Bench—the first benchmark explicitly designed for generalization evaluation—and empirically demonstrate a strong positive correlation between video generation quality and downstream policy success rates. Using only teleoperated data from a single task in a single environment, DreamGen achieves zero-shot transfer of 22 novel behaviors across both seen and unseen environments, significantly improving cross-behavior and cross-environment generalization performance.
Bridging the domain gap between real-world RGB-D images and robot simulation environments remains challenging for digital twin task generation. Method: This paper proposes a simulation-task alignment framework leveraging vision-language models (VLMs) and an iterative routing mechanism to generate executable simulation tasks end-to-end from single-frame RGB-D input. The method integrates SAM2 for precise object segmentation, VLM-driven semantic understanding, dynamic matching against a simulation asset library, and automated generation of self-validating test suites—forming a closed-loop “perceive–match–generate–verify” optimization pipeline. Contribution/Results: It achieves the first high-fidelity geometric-semantic alignment between real-scene objects and simulation assets while ensuring physical feasibility and executability within physics engines. Evaluated on multiple real-world benchmarks, the approach significantly improves object correspondence accuracy (+23.6%), task success rate (+31.4%), and cross-scene generalization.
This work proposes a general-purpose physics-based character controller capable of executing tasks with natural, realistic, and diverse motions. The approach discretizes motion data using Finite Scalar Quantization (FSQ) and integrates a GPT-style autoregressive Transformer with end-to-end reinforcement learning to build a transferable generative motion controller amenable to downstream task fine-tuning. A key innovation lies in the joint optimization of the discrete action vocabulary and the control policy, replacing conventional pipeline-based training procedures. Experimental results demonstrate that the method achieves a 99.98% motion reproduction success rate on large-scale motion datasets and exhibits robust behaviors such as perturbation response and fall recovery, proving effective across a range of downstream control tasks.
This work addresses the challenge of data-efficient and generalizable robot motion generation, aiming to learn from few demonstrations and adapt across diverse scenarios. The proposed approach integrates object-centric neural fields with a temporal mixture-of-experts (MoE) architecture, decomposing complex behaviors into object-based motion primitives through spatiotemporal compositionality. It further introduces canonical neural fields and latent-conditioned deformations to model 3D scene structure. Innovatively combining vision-based structural priors with language instructions, the method achieves category-level, cross-scenario systematic generalization. Experiments demonstrate that it accomplishes long-horizon manipulation tasks in simulation using significantly less training data than baselines, while exhibiting robustness to noise and enabling language-driven 3D manipulation in real-world environments.
This work addresses the challenge that pretrained generative robotic policies often fail under distribution shifts in real-world deployment, while existing adaptation methods either require large amounts of data or rely on online reinforcement learning, compromising efficiency or safety. To overcome this, the authors propose FlowDAgger, a method that efficiently fine-tunes a frozen generative policy in latent space through human intervention. Its key innovation is an action inversion mechanism that combines inverse time integration with local noise optimization to map human-provided actions back to corresponding latent variables. These latent targets are then used to train a lightweight latent policy that guides the base generative model. Experiments demonstrate that FlowDAgger achieves superior performance over supervised fine-tuning and latent-space reinforcement learning baselines in both simulated and real-world single- and dual-arm manipulation tasks, requiring only minimal human intervention while preserving the original policy’s capabilities.
Traditional animation production relies heavily on labor-intensive manual keyframing, resulting in high technical barriers and low efficiency. This work proposes the first language-driven animation generation framework, leveraging a large language model (LLM) to interpret natural language semantics and integrating the Segment Anything Model (SAM) for visual grounding and scene geometry understanding. The framework automatically generates high-quality animations that respect perspective constraints, depth structure, and occlusion logic. It supports complex animation types such as contour-following motion, depth-aware camera trajectories, and perspective-aligned transformations. Extensive evaluations across diverse scenes demonstrate the method’s feasibility and practicality, significantly lowering the technical threshold for animation creation.