FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of long-horizon bimanual robotic manipulation, where the lack of fine-grained annotations hinders complex multi-step task execution. To this end, we construct a densely annotated bimanual trajectory dataset and propose an action model based on a Vision-Language-Action (VLA) architecture. The core innovation lies in an intermediate training strategy incorporating self-predicting subtask objectives alongside large-scale data augmentation techniques, which jointly endow the model with robust subtask planning and execution capabilities. Experimental results demonstrate that the proposed method achieves 100% success in spatial disambiguation and improves performance on unseen long-horizon tasks to 76%. Furthermore, it enables zero-shot cross-hardware generalization with minimal fine-tuning data, significantly reducing deployment costs.
📝 Abstract
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask, and show that mid-training it this way yields substantial gains. Specifically, success on a spatial disambiguation task increases from 32.0% to 100.0%, and step-by-step human subtask guidance lifts success on an unseen long-horizon task from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires one-tenth the data of baselines without mid-training and generalizes zero-shot to completely unseen tasks on the new hardware. We open-source the full dataset, model weights, and training code.
Problem

Research questions and friction points this paper is trying to address.

bimanual manipulation
long-horizon tasks
manipulation datasets
fine-grained annotation
subtask labeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bimanual Manipulation
Vision-Language-Action Model
Fine-grained Annotation
Subtask Prediction
Zero-shot Generalization
🔎 Similar Papers
No similar papers found.