🤖 AI Summary
Existing robotic action tokenization relies solely on reconstruction objectives, neglecting the predictability and robustness required by downstream policies. This work proposes ProAct, which reframes action tokenization as a policy interface that balances fidelity, predictability, and robustness. By employing a policy-agnostic training mechanism that optimizes the tokenizer using only action data, ProAct generalizes across diverse architectures and Vision-Language-Action (VLA) models, effectively enhancing policy generalization and decoding under unseen observations. Experiments demonstrate that ProAct improves average success rates by 11.3% on simulation benchmarks and yields gains of 21.8% and 36.7% in real-world robotic manipulation scenarios.
📝 Abstract
Autoregressive action-token policies such as vision-language-action models require action tokenizers to translate discrete token sequences into precise control actions in continuous space. Many action tokenizers learn the mapping between tokens and actions via a reconstruction objective. However, as we show through extensive analysis, sufficiently accurate action reconstruction is only one part of what makes a downstream robot policy successful. It is also critical that the policy is able to predict the right tokens for new observations, and that unseen policy token predictions still decode into reasonable actions. These properties are downstream of tokenizer training and are not directly incentivized by a reconstruction objective alone. In this work, we introduce Predictable and Robust Action Tokenization (ProAct), a tokenizer training method that strategically augments reconstruction with the goal of improving downstream predictability and robustness. ProAct is policy-agnostic and uses only action datasets for training. Across the Robomimic, LIBERO, and RoboTwin benchmarks and a diverse set of tokenizer architectures, ProAct improves rollout success by an average of 11.3 percentage points. These improvements also translate to vision-language-action policies and real-world robotic manipulation, yielding average gains of 21.8 and 36.7 percentage points, respectively. These results suggest that effective action tokenization should be designed as a policy interface that balances fidelity, predictability, and robustness, rather than as a reconstruction problem alone.