MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出MimicAgent框架,通过生成参考轨迹而非设计奖励函数来学习四足机器人的动态技能,使用编码代理生成粗略参考轨迹以训练策略。
📝 Abstract
We present MimicAgent, a prompt-to-trajectory generation framework for learning dynamic quadruped skills. Although reward shaping is extensively used when training quadruped policies, navigating the resulting reward landscape is notoriously difficult, requiring hours of "graduate student descent". Eureka attempts to automate reward design with LLMs, but we find that it struggles to generalize across diverse skills and morphologies. Our key observation is that it is far easier for a human - and by association, an LLM - to generate reference motions than to shape reward functions. Our hypothesis is motivated by the success of example-guided RL for humanoids, which exploits large-scale motion capture datasets as references for training locomotion policies. Unlike humanoids, quadrupeds lack such reference motion data. Towards this end, we propose MimicAgent, an agentic harness that, given a skill prompt, generates quadruped reference trajectories with coding agents. These coarse reference trajectories are then used to train example-guided RL policies that are deployable in simulation and in the real-world. Notably, we find that when prompting Claude Fable 5.1 within our agentic harness, 87% of prompts yield semantically aligned reference trajectories.
Problem

Research questions and friction points this paper is trying to address.

reward shaping
quadruped skills
large language models
reference motions
motion capture datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt-to-trajectory
coding agents
example-guided RL
quadruped skills
reference trajectories
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.