Behavior-Aligned Action Tokenization for Robot Policy Learning

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决机器人策略学习中动作表示不一致的问题,提出了一种基于软动态时间规整的行为对齐动作标记化方法(BAAT),通过该方法提高了下游任务的学习效果。
📝 Abstract
Autoregressive robot policies learn continuous control by predicting discrete action tokens from observations. Different tasks often share local motions, yet behavioral correspondence across demonstrations receives limited explicit supervision in existing tokenizers. Motions with different timing can therefore lack a shared representation despite following similar patterns. We propose Behavior-Aligned Action Tokenization (BAAT), which uses soft dynamic time warping (Soft-DTW) to select corresponding action chunks and aligns their quantized coordinates jointly with reconstruction. This objective encourages similar motions across tasks to occupy nearby quantized representations while retaining executable action detail. A history-conditioned diffusion decoder reconstructs continuous action chunks from these tokens, and a downstream autoregressive policy learns to predict them. We evaluate BAAT on selected tasks from three simulation benchmarks and two real robot tasks. BAAT achieves a mean simulation success rate of approximately 45.2%, exceeding OAT by approximately 7.2 percentage points. In the controlled LIBERO-All alignment ablation, policy success rises from 70.2% to 79.0% while trajectory replay success decreases. These results support behavioral correspondence as supervision for organizing shared motion structure in action tokenizers and improving downstream robot policy learning.
Problem

Research questions and friction points this paper is trying to address.

Behavior-Aligned
Action Tokenization
Soft-DTW
Robot Policy Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavior-Aligned Action Tokenization
Soft-DTW
autoregressive policies
quantized representations
diffusion decoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junbo Dong
Southern University of Science and Technology
Ze Chen
Ze Chen
Alibaba Group
Comuter Vision
Z
Zhendong Xie
Southern University of Science and Technology
Junjie Li
Junjie Li
University Of Science And Technology Of China
Few-shot learningDomain adaptationImage inpainting
L
Lixin Xu
National University of Singapore
Xuemin Chi
Xuemin Chi
Zhejiang University
Motion Planning
Y
Yiming Song
Southern University of Science and Technology
Z
Zhaoyuan Ma
Southern University of Science and Technology