🤖 AI Summary
This work proposes Decision MetaMamba to address the performance limitations of Mamba-based models in offline reinforcement learning, where the selective scanning mechanism tends to overlook critical sequential information. By replacing the original token mixer with a dense-layer-driven heterogeneous sequence mixing mechanism and refining the positional structure to better preserve local context, Decision MetaMamba enables full-channel joint sequence modeling at the Mamba frontend. This design effectively mitigates information loss caused by selective scanning and residual gating. The resulting model achieves state-of-the-art performance across multiple offline reinforcement learning benchmarks while maintaining a compact parameter count, demonstrating both high efficiency and strong practical potential.
📝 Abstract
Mamba-based models have drawn much attention in offline RL. However, their selective mechanism often detrimental when key steps in RL sequences are omitted. To address these issues, we propose a simple yet effective structure, called Decision MetaMamba (DMM), which replaces Mamba's token mixer with a dense layer-based sequence mixer and modifies positional structure to preserve local information. By performing sequence mixing that considers all channels simultaneously before Mamba, DMM prevents information loss due to selective scanning and residual gating. Extensive experiments demonstrate that our DMM delivers the state-of-the-art performance across diverse RL tasks. Furthermore, DMM achieves these results with a compact parameter footprint, demonstrating strong potential for real-world applications.