Beyond Reconstruction: What Matters in Action Tokenization for Robot Policies?

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing robotic action tokenization relies solely on reconstruction objectives, neglecting the predictability and robustness required by downstream policies. This work proposes ProAct, which reframes action tokenization as a policy interface that balances fidelity, predictability, and robustness. By employing a policy-agnostic training mechanism that optimizes the tokenizer using only action data, ProAct generalizes across diverse architectures and Vision-Language-Action (VLA) models, effectively enhancing policy generalization and decoding under unseen observations. Experiments demonstrate that ProAct improves average success rates by 11.3% on simulation benchmarks and yields gains of 21.8% and 36.7% in real-world robotic manipulation scenarios.
📝 Abstract
Autoregressive action-token policies such as vision-language-action models require action tokenizers to translate discrete token sequences into precise control actions in continuous space. Many action tokenizers learn the mapping between tokens and actions via a reconstruction objective. However, as we show through extensive analysis, sufficiently accurate action reconstruction is only one part of what makes a downstream robot policy successful. It is also critical that the policy is able to predict the right tokens for new observations, and that unseen policy token predictions still decode into reasonable actions. These properties are downstream of tokenizer training and are not directly incentivized by a reconstruction objective alone. In this work, we introduce Predictable and Robust Action Tokenization (ProAct), a tokenizer training method that strategically augments reconstruction with the goal of improving downstream predictability and robustness. ProAct is policy-agnostic and uses only action datasets for training. Across the Robomimic, LIBERO, and RoboTwin benchmarks and a diverse set of tokenizer architectures, ProAct improves rollout success by an average of 11.3 percentage points. These improvements also translate to vision-language-action policies and real-world robotic manipulation, yielding average gains of 21.8 and 36.7 percentage points, respectively. These results suggest that effective action tokenization should be designed as a policy interface that balances fidelity, predictability, and robustness, rather than as a reconstruction problem alone.
Problem

Research questions and friction points this paper is trying to address.

Action Tokenization
Robot Policies
Reconstruction Objective
Predictability
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action Tokenization
Robot Policies
Reconstruction Objective
Predictability and Robustness
Vision-Language-Action Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoran Chen
Toyota Technological Institute at Chicago
Jingtian Ji
Jingtian Ji
Toyota Technological Institute at Chicago
robot learning
S
Samuel Wheeler
Argonne National Laboratory
Kaylene Caswell Stocking
Kaylene Caswell Stocking
Research Assistant Professor at Toyota Technological Institute at Chicago
roboticscognitive scienceneural engineering
M
Matthew Walter
Toyota Technological Institute at Chicago