Unified token representations for sequential decision models

📅 2025-10-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing offline reinforcement learning methods—such as Decision Transformers—are hindered by redundant tokenization and quadratic attention complexity, compromising both computational efficiency and generalization. To address this, we propose Unified Token Representation (UTR), which encodes return, state, and action into a single compact token, drastically shortening sequence length and reducing computational overhead. Theoretically, UTR yields a tighter Rademacher complexity bound, thereby enhancing generalization. Leveraging UTR, we design two architectures: UDT (a Transformer-based variant) and UDC (a gated CNN-based variant). On standard offline RL benchmarks, both achieve state-of-the-art or competitive performance while accelerating inference by 2–5× and reducing memory footprint by 30%–60%. Moreover, UTR demonstrates strong cross-architecture transfer robustness.

Technology Category

Machine Learning: Online Learning & BanditsSearch and Optimization: Learning to SearchReasoning under Uncertainty: Sequential Decision Making

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
Transformers have demonstrated strong potential in offline reinforcement learning (RL) by modeling trajectories as sequences of return-to-go, states, and actions. However, existing approaches such as the Decision Transformer(DT) and its variants suffer from redundant tokenization and quadratic attention complexity, limiting their scalability in real-time or resource-constrained settings. To address this, we propose a Unified Token Representation (UTR) that merges return-to-go, state, and action into a single token, substantially reducing sequence length and model complexity. Theoretical analysis shows that UTR leads to a tighter Rademacher complexity bound, suggesting improved generalization. We further develop two variants: UDT and UDC, built upon transformer and gated CNN backbones, respectively. Both achieve comparable or superior performance to state-of-the-art methods with markedly lower computation. These findings demonstrate that UTR generalizes well across architectures and may provide an efficient foundation for scalable control in future large decision models.
Problem

Research questions and friction points this paper is trying to address.

Reduces redundant tokenization in decision transformers
Addresses quadratic attention complexity in offline RL
Improves scalability for resource-constrained control systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified token representation merges return state action
UTR reduces sequence length and model complexity
UDT UDC variants achieve superior performance efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhuojing Tian
Intelligent Game and Decision Lab(IGDL), Beijing, China
Y
Yushu Chen
Tsinghua University, Department of Computer Science and Technology, Beijing, China