LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited reusability of atomic interactions in vision-language-action (VLA) models on unseen tasks. To this end, it proposes a dual-layer codebook mechanism grounded in a retrievable lexicon of atomic action words, where global and detail codebooks capture shared structures and execution discrepancies, respectively. By integrating visual-atomic action alignment, trajectory reconstruction, and a scene-aware adapter, the method achieves cross-task generalization without requiring skill experts or parameter updates. Trained on the AtomAction dataset, the approach improves success rates on unseen RLBench tasks by 17.5% to 71.08%. Furthermore, real-world experiments validate its capability for stepwise execution and fault recovery.
📝 Abstract
Vision-language-action (VLA) models struggle to reuse recurring interactions in unseen tasks. Our diagnostic study reveals that reliable task completion does not imply consistent execution of constituent atomic actions across task contexts. We present LexiconVLA, a retrievable atomic-action lexicon for cross-task reuse. Global and detail codebooks capture shared interaction structure and fine-grained execution variation, respectively, preserving both reusable patterns and execution details. Visual-Atomic Action Alignment couples trajectory reconstruction from visual state changes with visual outcome prediction from action codes, grounding the lexicon in motion and its effects. We learn these codebooks with trajectory reconstruction and visual alignment on our AtomAction Dataset of 57,803 segments from 69 tasks. A planner and scene-aware adapter translate new goals into code-conditioned subtasks for a shared policy, without skill-specific experts or deployment-time parameter updates. Across five policy backbones on 26 RLBench tasks, LexiconVLA largely maintains performance on 18 seen tasks while improving success on 8 tasks held out from policy training. With BridgeVLA, unseen-task success rises from 16.67% to 34.17% (+17.50 percentage points), and overall success reaches 71.08%, the highest among methods with reported results. Real-robot experiments demonstrate stepwise execution and failure recovery.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
atomic action reuse
unseen tasks
cross-task generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Atomic Action Codebooks
Cross-Task Reuse
Visual-Action Alignment
Zero-Shot Generalization
🔎 Similar Papers
No similar papers found.