🤖 AI Summary
Existing offline reinforcement learning methods—such as Decision Transformers—are hindered by redundant tokenization and quadratic attention complexity, compromising both computational efficiency and generalization. To address this, we propose Unified Token Representation (UTR), which encodes return, state, and action into a single compact token, drastically shortening sequence length and reducing computational overhead. Theoretically, UTR yields a tighter Rademacher complexity bound, thereby enhancing generalization. Leveraging UTR, we design two architectures: UDT (a Transformer-based variant) and UDC (a gated CNN-based variant). On standard offline RL benchmarks, both achieve state-of-the-art or competitive performance while accelerating inference by 2–5× and reducing memory footprint by 30%–60%. Moreover, UTR demonstrates strong cross-architecture transfer robustness.
📝 Abstract
Transformers have demonstrated strong potential in offline reinforcement learning (RL) by modeling trajectories as sequences of return-to-go, states, and actions. However, existing approaches such as the Decision Transformer(DT) and its variants suffer from redundant tokenization and quadratic attention complexity, limiting their scalability in real-time or resource-constrained settings. To address this, we propose a Unified Token Representation (UTR) that merges return-to-go, state, and action into a single token, substantially reducing sequence length and model complexity. Theoretical analysis shows that UTR leads to a tighter Rademacher complexity bound, suggesting improved generalization. We further develop two variants: UDT and UDC, built upon transformer and gated CNN backbones, respectively. Both achieve comparable or superior performance to state-of-the-art methods with markedly lower computation. These findings demonstrate that UTR generalizes well across architectures and may provide an efficient foundation for scalable control in future large decision models.