Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the semantic misalignment between action representations and autoregressive backbones caused by existing action tokenizers, which limits robot learning. We propose CATok, which reformulates action tokenization as a causal generative process, achieving coarse-to-fine semantic alignment via conditional annealed flow matching. This approach constructs a causal token space that inherently decouples high-level reasoning from low-level execution without requiring explicit attention masks, while integrating a multimodal diffusion Transformer (MMDiT) with a discrete bottleneck architecture to enhance modeling capacity. Experiments demonstrate that CATok significantly improves reconstruction fidelity, inference efficiency, and Vision-Language-Action (VLA) success rates across both simulated and real-world tasks, establishing a foundation for high-performance autoregressive robot learning.
📝 Abstract
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
Problem

Research questions and friction points this paper is trying to address.

Action Tokenization
Autoregressive VLA Models
Semantic Misalignment
Robot Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Action Tokenization
Conditional Annealing
Flow Matching
Vision-Language-Action Models
Multimodal Diffusion Transformer
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chenyu Zhang
IIIS, Tsinghua University; Shanghai Qizhi Institute
Yuhang Cao
Yuhang Cao
MMLab The Chinese University of Hong Kong
Multi-Modal Large Language ModelObject DetectionFew Shot Object Detection
D
Daru Du
IIIS, Tsinghua University; Shanghai Qizhi Institute
Y
Yingxi Lu
IIIS, Tsinghua University; Shanghai Qizhi Institute
Jing Shao
Jing Shao
Research Scientist, Shanghai AI Laboratory/Shanghai Jiao Tong University
Computer VisionMulti-Modal Large Language Model
R
Ruoqu Chen
IIIS, Tsinghua University; Shanghai Qizhi Institute
J
Jiajun Liu
IIIS, Tsinghua University; Shanghai Qizhi Institute
L
Liu Cao
IIIS, Tsinghua University; Shanghai Qizhi Institute
Yicheng Liu
Yicheng Liu
Tsinghua University
Robotics
H
Hang Zhao
IIIS, Tsinghua University; Shanghai Qizhi Institute
Mengdi Xu
Mengdi Xu
Stanford University
RoboticsMachine Learning