Score
Design and implement codecs and tokenizers that convert continuous actions or motion trajectories into compact, ordered discrete token sequences and codebooks (e.g., neural action codecs and VQ‑VAE‑style encoders); this includes scaling and expanding codebook capacity, ensuring low-latency encoding, and using synthetic-augmented token generation to learn richer motion primitives and improve compositionality and coverage for autoregressive models.
Current Vision-Language-Action (VLA) models lack a systematic understanding of action tokens, hindering principled model design and research direction. This work establishes action tokenization as a unifying analytical lens, proposing the first taxonomy of VLA models grounded in eight distinct action tokenization paradigms. We introduce a comprehensive analytical framework evaluating representation capacity, generalizability, and physical realizability. Leveraging multimodal machine learning principles and systematic literature review, we characterize fundamental trade-offs among accuracy, computational efficiency, and deployment feasibility across paradigms—and identify critical domain blind spots. Our analysis yields a structured cognitive framework for VLA models, clarifies the appropriate application contexts and intrinsic limitations of each action representation, and provides a reproducible methodology and clear evolutionary roadmap for generalizable, embodied-action modeling.
This work addresses the challenge that existing discrete action tokenizers in vision–language–action models struggle to simultaneously achieve high compression ratios, low latency, and strong downstream task performance. Inspired by neural audio codecs, the authors propose a multi-scale residual vector-quantized generative adversarial network (RVQGAN) that treats robot action trajectories as multi-channel one-dimensional signals for efficient compression and reconstruction. Key innovations include a shifted codebook architecture that yields a structured and compact action token space, a Vocos-style decoder with an inverse short-time Fourier transform (ISTFT) head, and tailored time-domain and non-Mel spectral reconstruction losses combined with an adversarial discriminator. Evaluated on LIBERO-10, RoboMimic, and real-world tasks, the method achieves significantly lower reconstruction error and higher task success rates than conventional binning, FAST, and existing vector-quantization tokenizers, at comparable or better compression ratios.
Existing action tokenizers overly prioritize reconstruction fidelity while neglecting their direct impact on vision–language–action (VLA) model training, lacking design principles explicitly optimized for VLA tasks. This work proposes, for the first time, an information-theoretic framework for action tokenizer design grounded in VLA training objectives: maximizing temporal token overlap, minimizing vocabulary redundancy, and enhancing both multimodal mutual information and token independence. Based on these principles, we develop ActionCodec, an efficient tokenizer enabling end-to-end autoregressive VLA training. Integrated into SmolVLM2-2.2B, ActionCodec achieves a 95.5% success rate on the LIBERO benchmark without any robot pretraining; with architectural enhancements, performance further improves to 97.4%, establishing a new state of the art in VLA.
Existing motion capture datasets suffer from limited diversity, which constrains the generalization capabilities of generative models on rare, highly dynamic, and compositionally complex actions. To address this limitation, this work proposes a method that leverages large-scale synthetic human motion data combined with physics-based plausibility constraints to jointly expand both the training distribution and the size of the discrete codebook. By reconstructing the VQ-VAE motion tokenizer beyond the confines of real-data distributions, the approach substantially broadens the coverage and compositional capacity of the discrete motion representation space. This leads to consistent performance gains in text-to-motion generation and motion in-betweening tasks, and the enhanced representations can be seamlessly integrated into existing frameworks such as MotionGPT, demonstrating both the effectiveness and generalizability of the proposed representation expansion.
Conventional per-dimension, per-frame binning-based action tokenization fails for vision-language-action (VLA) models in high-frequency, dexterous robotic tasks, leading to poor reconstruction of continuous actions from discrete token predictions. Method: We propose Frequency-domain Adaptive Spatio-Temporal tokenization (FAST), which leverages discrete cosine transform (DCT) and compressed sensing to project action sequences into a low-dimensional frequency domain and discretize them—thereby overcoming temporal-domain binning limitations. We further introduce FAST+, a universal black-box tokenizer enabling unified encoding of multimodal and multi-frequency action signals. Results: Trained on 1 million real-world robot trajectories (10k hours), an autoregressive Transformer VLA model using FAST+ matches diffusion-based methods in performance while achieving up to 5× faster training convergence.
This work addresses the challenge that discrete action tokenization in vision–language–action (VLA) models struggles to accurately reconstruct continuous control signals conditioned on robot proprioceptive states, such as joint configurations and object poses. To overcome this limitation, the authors propose SA-VLA, a state-aware action tokenizer that integrates current state information into the discrete action decoding process. This is achieved through either state–action cross-attention or a lightweight state adapter, enabling each discrete token to represent a family of state-dependent continuous actions. By conditioning action generation on real-time robot states, SA-VLA significantly enhances representational capacity while preserving the efficiency of discrete modeling. Empirical results demonstrate substantial improvements: average success rates on 12 RoboTwin tasks rise from 0.29 to 0.56, and zero-shot sim-to-real transfer performance on three real-world tasks improves from 0.15 to 0.33.