${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出M^2Tok方法,通过多头多码本动作标记化技术减少重建误差,提高视觉-语言-动作模型的性能。
📝 Abstract
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This ``discretization bottleneck'' significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose $\mathcal{M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the $\mathcal{M}^2$Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at \href{https://github.com/cpaaax/M2Tok}{https://github.com/cpaaax/M2Tok}.
Problem

Research questions and friction points this paper is trying to address.

discretization bottleneck
reconstruction loss
fine-grained dynamics
Vision-Language-Action models
discrete action tokenizers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-head
Multi-codebook
Action Tokenization
Reconstruction Loss
Vision-Language-Action Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Chunpu Xu
Chunpu Xu
PolyU
Multimodal learningNatural language processing
Zhixuan Liang
Zhixuan Liang
University of Hong Kong
Embodied AIMachine LearningRoboticsComputer Vision
Yuhao Zhang
Yuhao Zhang
Shanghai Jiao Tong University
Theoritical Computer Science
Chi-Min Chan
Chi-Min Chan
HKUST
Large Language ModelsPost-TrainingAlignmentLLM Agents
J
Jessie Wang
The Hong Kong Polytechnic University, HongKong SAR, China
Y
Yang Xiao
The Hong Kong Polytechnic University, HongKong SAR, China
Mengkang Hu
Mengkang Hu
University of Hong Kong
Natural Language ProcessingEmbodied AILLM Agent
X
Xiaokang Yang
Shanghai Jiao Tong University, Shanghai, China
Y
Yao Mu
Shanghai Jiao Tong University, Shanghai, China