DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of learning diverse, long-horizon robotic manipulation policies from limited demonstrations. To this end, it proposes a language-guided diffusion framework that, for the first time, integrates Dynamic Movement Primitives (DMPs) with language-conditioned diffusion models for data augmentation. The approach leverages cross-attention mechanisms to enable multi-skill trajectory generation and employs large language models (LLMs) for task decomposition, thereby supporting both natural language interaction and autonomous execution. The primary contribution lies in achieving multi-skill composition and orientation-sensitive behavior control within a single generative policy. Validated across real-world robotic and simulated environments, the proposed method significantly outperforms baseline approaches, demonstrating robust multi-skill generalization from few-shot demonstrations and reliable execution of long-horizon tasks.
📝 Abstract
Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM leverages Dynamic Movement Primitives (DMPs) to augment a small set of expert demonstrations into expressive multi-skill datasets, substantially reducing data collection while improving spatial generalization beyond the demonstrated workspace. Building upon Motion Planning Diffusion (MPD), we extend the diffusion architecture to support language-conditioned multi-skill trajectory generation through cross-attention, allowing a single model to generate skill-consistent motions for a diverse set of manipulation primitives, including orientation-sensitive behaviors that are difficult to design using conventional motion planning or hard-coded controllers. For long-horizon manipulation, a large language model decomposes high-level operator requests into executable sequences of skills, enabling natural language interaction and autonomous task execution. We validate DROM on a Franka Emika Panda robot, a FANUC CRX25ia robot, and in MuJoCo simulation across a wide range of manipulation tasks. Experimental results demonstrate that DROM outperforms Motion Planning Diffusion and Behavior Cloning baselines, achieves robust multi-skill generalization, and composes learned skills to reliably execute long-horizon manipulation tasks from natural language instructions using only a limited number of human demonstrations. Datasets, simulation environments, and more at https://github.com/automation-robotics-machines/drom.
Problem

Research questions and friction points this paper is trying to address.

Robotic Manipulation
Multi-Skill Learning
Long-Horizon Tasks
Limited Demonstrations
Language-Guided Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Policy
Dynamic Movement Primitives
Language-Conditioned Generation
Long-Horizon Manipulation
Large Language Models
V
Vincenzo Pomponi
Institute of Systems and Technologies for Sustainable Production (ISTePS), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland.
R
Rocco Felici
Institute of Systems and Technologies for Sustainable Production (ISTePS), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland.
P
Paolo Franceschi
Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland.
Stefano Baraldo
Stefano Baraldo
SUPSI - University of Applied Sciences of Southern Switzerland
Industrial RoboticsMetal Additive ManufacturingMachine Learning
O
Oliver Avram
Institute of Systems and Technologies for Sustainable Production (ISTePS), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland.
L
Loris Roveda
Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Department of Innovative Technologies, University of Applied Science and Arts of Southern Switzerland (SUPSI), Via la Santa 1, Lugano, CH-6900, Ticino, Switzerland. Mechanical Department, Politecnico di Milano (PoliMi), Via Giuseppe Candiani, Milan, 20155, Italy.
L
Luca Maria Gambardella
Istituto Dalle Molle di studi sull’intelligenza artificiale (IDSIA), Faculty of Informatics, Università della Svizzera Italiana (USI), Via Buffi 13, Lugano, CH-6900, Ticino, Switzerland.
Anna Valente
Anna Valente
SUPSI