YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of aligning instructions with physical manipulation in Vision-Language-Action (VLA) models, which arises from the lack of fine-grained interaction semantics in demonstration data. To this end, we propose YUBI-STAG, a spatiotemporal annotation framework that integrates contact-object segmentation, vision-language model (VLM) reasoning, and knowledge distillation to automatically annotate robot manipulation videos with detailed contact, object, and action labels. Furthermore, we distill an efficient model, YUBI-VLM, capable of recovering action structures solely from wrist-view observations. Experimental results demonstrate that our approach maintains high annotation accuracy while significantly reducing inference costs, effectively enhancing the instruction-following capabilities of bimanual robots and their fine-grained control over object identity and spatial relationships.
📝 Abstract
Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions are executed, including which gripper acts, which object is contacted, and how it is grasped and moved. We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language. Combining contact-object segmentation with vision-language models, YUBI-STAG annotates object identities, attributes and states, per-gripper actions, bimanual coordination, and spatially grounded interactions. To address YUBI-STAG's reliance on localized sequences and multi-stage VLM inference, we distill it into YUBI-VLM. YUBI-VLM directly recovers action structure and annotations from raw, unsegmented video in few inference calls and operates from wrist views alone. We evaluate both frameworks on YUBI-STAG-Bench across temporal, semantic, and spatial grounding tasks. YUBI-VLM retains much of YUBI-STAG's annotation accuracy with fewer inference calls and shorter runtime while generalizing to unseen manipulations. Finally, post-training VLA policies on these annotations aligns them with fine-grained language and contact-aware structure. Bimanual experiments demonstrate improved performance and instruction following, including control over object identity, acting gripper, target location, and spatial relations absent from original labels.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
fine-grained alignment
robot manipulation
video-language grounding
contact-aware semantics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Automated Video-Language Grounding
Spatio-Temporal Annotation
Knowledge Distillation
Bimanual Manipulation
🔎 Similar Papers
No similar papers found.