ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of skill selection ambiguity, state inheritance bias, and evaluation difficulties in long-horizon mobile manipulation. Leveraging the BEHAVIOR-1K simulation environment, we construct a skill dataset and benchmark comprising 50 activities. Methodologically, we introduce explicit subtask instruction pairing, initial-state sensitivity measurement, and a state-recovery-based local success criterion, integrating vision-language-action models with reinforcement learning for policy training. Our analysis reveals that conventional aggregated scores obscure inter-skill disparities, and joint perturbations reduce success rates by 56%. The proposed skill policies achieve a local success rate of 78.7% and improve full-task success to 18%, demonstrating the efficacy of fine-grained skill decomposition and robust evaluation mechanisms for complex sequential manipulation tasks.
📝 Abstract
Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction. Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested. We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities. Its dataset contains 137,899 segments across 21 skill types and 417 subtasks, and its benchmark contains 1,260 test instances. Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages. Evaluations of representative vision-language-action (VLA) policies show that similar aggregate scores can hide substantial per-skill differences. The tested starting-state perturbations also degrade execution: on the full benchmark, joint perturbations reduce success rates by approximately 56% relative to those from demonstrated starting states. On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations. The trained skills further support complete-task execution on these activities, as coordinating the task and skill policies through a planner raises full-task success from 4.0% to 18.0%.
Problem

Research questions and friction points this paper is trying to address.

long-horizon mobile manipulation
manipulation skill evaluation
skill selection ambiguity
starting-state perturbation
skill-level diagnosis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-horizon manipulation
Skill-level benchmark
Vision-language-action policies
State perturbation sensitivity
Subtask instruction
🔎 Similar Papers
No similar papers found.