DreamFormer: Dream Imitation with a Transformer World Model for Language-Conditioned Robotic Manipulation

๐Ÿ“… 2026-10-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenges of covariate shift and the high computational cost of long-horizon imagination in robotic manipulation by proposing a latent-space reinforcement learning framework built upon Transformer-based world models. Methodologically, it introduces a single-token high-resolution multi-view encoding scheme that substantially reduces the overhead of long-horizon imagination. By imitating expert policies within the latent space and incorporating an intrinsic reward mechanism, the approach mitigates distributional shift through online training on imagined trajectories, thereby enabling language-conditioned multi-task skill acquisition. Experimental results demonstrate that the proposed method surpasses LUMOS on the CALVIN benchmark and achieves nearly twice the zero-shot transfer performance of HULC, validating the efficient generalization capability of the learned dynamics model.
๐Ÿ“ Abstract
We introduce DreamFormer, a model-based agent that acquires language-conditioned, multi-task skills by imitating expert demonstrations within the latent imagination of a learned world model. DreamFormer first learns a task-agnostic Transformer world model from unstructured play data, then acquires task-specific behaviors by optimizing an intrinsic reward that aligns agent-generated rollouts with expert demonstrations in latent space. Since the policy is trained on-policy inside imagination, it is exposed to its own errors during training, mitigating the covariate shift inherent to offline behavioral cloning. To make long-horizon imagination affordable, DreamFormer encodes a high-resolution multi-view robotic observation into a single input token, avoiding both spatial downsampling and the multi-token representations used by prior Transformer world models. On the long-horizon CALVIN benchmark, DreamFormer outperforms LUMOS, the comparable model-based agent, on single-environment evaluation (2.52 vs 2.34 average tasks completed per chain of five) while keeping imagination rollouts tractable. Against HULC, the behavior cloning baseline, it nearly doubles performance on zero-shot transfer to an unseen environment (1.30 vs 0.67), indicating that dynamics learned by the world model transfer more readily than a directly cloned policy. This is consistent with the role attributed to internal models in biological agents, where a model of environment dynamics supports behavior in situations not previously encountered.
Problem

Research questions and friction points this paper is trying to address.

language-conditioned robotic manipulation
world model
covariate shift
long-horizon imagination
zero-shot transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer World Model
Latent Imagination
Language-Conditioned Manipulation
Single-Token Encoding
Zero-Shot Transfer
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Mostafa Kotb
Knowledge Technology Group, Department of Informatics, University of Hamburg, 20146 Hamburg, Germany
Cornelius Weber
Cornelius Weber
Universitรคt Hamburg
neural networkscomputational neurosciencecognitive robotics
Muhammad Burhan Hafez
Muhammad Burhan Hafez
School of Electronics and Computer Science, University of Southampton, Southampton SO17 1BJ, UK
S
Stefan Wermter
Knowledge Technology Group, Department of Informatics, University of Hamburg, 20146 Hamburg, Germany