๐ค AI Summary
This work addresses the challenge of enabling robots to flexibly adapt their motion execution styles across diverse scenarios while lacking reusable, task-agnostic control mechanisms. The authors propose MoMo, a novel framework that models motion modalities as composable and transferable continuous behavioral factors. By integrating a spatiotemporal action tokenizer with a conditional behavior cloning Transformer, MoMo jointly encodes task instructions and motion modalities to enable controllable skill generation. The approach supports seamless switching among steady, dynamic, and intermediate motion styles and achieves compositional generalization on unseen taskโmodality combinations. Experiments demonstrate that MoMo successfully generates human-distinguishable motion styles across six real-world manipulation tasks, maintaining high task success rates while enabling cross-task style transfer.
๐ Abstract
To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present \textbf{MoMo}, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady, dynamic, and intermediate behaviors that human raters can distinguish and that differ in joint speed, acceleration, and end-effector approach pitch. On tasks demonstrated in only one mode, MoMo transfers the unseen requested mode while largely preserving task success. Together, these results provide evidence of compositional generalization to unseen task--mode combinations and show that motion mode can be reused across tasks to control how a manipulation skill is performed.