๐ค AI Summary
This study addresses the challenges of covariate shift and the high computational cost of long-horizon imagination in robotic manipulation by proposing a latent-space reinforcement learning framework built upon Transformer-based world models. Methodologically, it introduces a single-token high-resolution multi-view encoding scheme that substantially reduces the overhead of long-horizon imagination. By imitating expert policies within the latent space and incorporating an intrinsic reward mechanism, the approach mitigates distributional shift through online training on imagined trajectories, thereby enabling language-conditioned multi-task skill acquisition. Experimental results demonstrate that the proposed method surpasses LUMOS on the CALVIN benchmark and achieves nearly twice the zero-shot transfer performance of HULC, validating the efficient generalization capability of the learned dynamics model.
๐ Abstract
We introduce DreamFormer, a model-based agent that acquires language-conditioned, multi-task skills by imitating expert demonstrations within the latent imagination of a learned world model. DreamFormer first learns a task-agnostic Transformer world model from unstructured play data, then acquires task-specific behaviors by optimizing an intrinsic reward that aligns agent-generated rollouts with expert demonstrations in latent space. Since the policy is trained on-policy inside imagination, it is exposed to its own errors during training, mitigating the covariate shift inherent to offline behavioral cloning. To make long-horizon imagination affordable, DreamFormer encodes a high-resolution multi-view robotic observation into a single input token, avoiding both spatial downsampling and the multi-token representations used by prior Transformer world models. On the long-horizon CALVIN benchmark, DreamFormer outperforms LUMOS, the comparable model-based agent, on single-environment evaluation (2.52 vs 2.34 average tasks completed per chain of five) while keeping imagination rollouts tractable. Against HULC, the behavior cloning baseline, it nearly doubles performance on zero-shot transfer to an unseen environment (1.30 vs 0.67), indicating that dynamics learned by the world model transfer more readily than a directly cloned policy. This is consistent with the role attributed to internal models in biological agents, where a model of environment dynamics supports behavior in situations not previously encountered.