🤖 AI Summary
This study addresses the absence of multi-task simulation benchmarks for contact manipulation on underwater floating bases, which hinders the evaluation of visuomotor policies. We construct the first underwater dynamics simulation benchmark comprising ten expert-solvable tasks across dual platforms. A standardized evaluation protocol is proposed to decouple expert feasibility, policy completion, and execution effort, enabling systematic assessment of algorithms including ACT, Diffusion Policy, and SmolVLA. Our analysis reveals that the absence of integral action constitutes the primary cause of task failure, demonstrating the significant impact of action interfaces on success rates. Furthermore, this work quantifies the task-dependent effects of current disturbances on both completion performance and energy consumption, providing critical insights for developing robust underwater manipulation strategies.
📝 Abstract
Underwater manipulation couples visual decisions, contact forces and a thruster-controlled floating base. We introduce WasserMan, to our knowledge the first multi-task simulation benchmark for visuomotor learning of floating-base underwater contact manipulation. It provides ten expert-solvable tasks, two vehicle-arm platforms, a bimanual configuration and underwater dynamics. Nine tasks have learned-policy evaluations. We compare ACT, diffusion policies (DP) and behavioral cloning on six tasks, with three training runs and equal sampled-window budgets. Pretrained SmolVLA adds evaluations on three tasks. In the tested settings, removing integral action can prevent completion, while changing action interfaces can reduce learned-policy success despite successful expert replay. Currents produce task-dependent changes in completion and actuation effort. Versioned tasks, demonstrations and per-episode evidence support a protocol separating expert feasibility, learned completion and execution effort.