๐ค AI Summary
This work addresses the challenge of achieving unified and scalable multitask decision-making across a vast array of heterogeneous reinforcement learning environments. It introduces LDM-v0, a Transformer-based universal policy model trained offline at scale on multimodal trajectory data from approximately 1,000 diverse domainsโincluding robotics, autonomous driving, inventory management, cybersecurity, trading, and gaming. The model performs supervised next-action prediction conditioned on historical observations, actions, rewards, and termination signals. For the first time, it demonstrates that a single Transformer policy can match the performance of task-specific policies across over a thousand heterogeneous tasks, establishing a unified paradigm for large-scale multitask reinforcement learning.
๐ Abstract
Recent progress in large-scale sequence modeling has shown that a single model can learn useful representations across highly diverse data distributions. Inspired by these advances, we investigate whether a unified transformer policy can be trained across large collections of heterogeneous reinforcement learning environments.
We introduce LDM-v0, a Large Decision Model trained offline on trajectories collected from thousands of environments spanning multiple domains and modalities. LDM-v0 is a multi-task, multi-modal transformer policy conditioned on histories of observations, actions, rewards, and termination signals, and trained through supervised next-action prediction over offline trajectories. We describe the environment infrastructure, automated data generation pipeline, model architecture, and training methodology used to build LDM-v0, and evaluate its performance across diverse environments. We show that a single pretrained model matches the performance of independently trained task-specific reference policies on approximately 1,000 environments including robotics, autonomous driving, inventory management, cybersecurity, trading, and video games. These results demonstrate the feasibility of large-scale offline pretraining across heterogeneous reinforcement learning environments using a single transformer policy.