Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing

📅 2026-02-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes Decision MetaMamba to address the performance limitations of Mamba-based models in offline reinforcement learning, where the selective scanning mechanism tends to overlook critical sequential information. By replacing the original token mixer with a dense-layer-driven heterogeneous sequence mixing mechanism and refining the positional structure to better preserve local context, Decision MetaMamba enables full-channel joint sequence modeling at the Mamba frontend. This design effectively mitigates information loss caused by selective scanning and residual gating. The resulting model achieves state-of-the-art performance across multiple offline reinforcement learning benchmarks while maintaining a compact parameter count, demonstrating both high efficiency and strong practical potential.

Technology Category

Search and Optimization: Metareasoning and MetaheuristicsMachine Learning: Mixture of Experts (MoE)Reasoning under Uncertainty: Sequential Decision Making

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Mamba-based models have drawn much attention in offline RL. However, their selective mechanism often detrimental when key steps in RL sequences are omitted. To address these issues, we propose a simple yet effective structure, called Decision MetaMamba (DMM), which replaces Mamba's token mixer with a dense layer-based sequence mixer and modifies positional structure to preserve local information. By performing sequence mixing that considers all channels simultaneously before Mamba, DMM prevents information loss due to selective scanning and residual gating. Extensive experiments demonstrate that our DMM delivers the state-of-the-art performance across diverse RL tasks. Furthermore, DMM achieves these results with a compact parameter footprint, demonstrating strong potential for real-world applications.
Problem

Research questions and friction points this paper is trying to address.

offline RL
Mamba
selective mechanism
sequence modeling
information loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision MetaMamba
Selective SSM
Offline RL
Sequence Mixing
Information Preservation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wall Kim
Samsung Electronics, Hwaseong, South Korea
C
Chaeyoung Song
Seoul National University of Science and Technology, Seoul, South Korea
Hanul Kim
Hanul Kim
Seoul National University of Science and Technology
Computer VisionMachine Learning