🤖 AI Summary
This work addresses the challenge of integrating low-level robotic manipulation with high-level symbolic reasoning under physical constraints in assembly tasks. To this end, we introduce WorkBenchMark, the first hierarchical and scalable benchmark for LEGO Duplo assembly, comprising 400 tasks across four levels of complexity. We propose a planning-driven baseline method that uniquely combines open-vocabulary visual perception with a disassembly-based backward reasoning strategy to enable semantically guided, reliable assembly. Experimental results demonstrate that our approach consistently outperforms existing vision-language-action models across all complexity levels. The benchmark, simulation environment, and source code will be publicly released to foster research in intelligent manufacturing.
📝 Abstract
We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufacturing League. Robotic assembly couples low-level manipulation with task-level symbolic reasoning under physical constraints, a combination that current end-to-end learning methods do not yet solve reliably. The benchmark provides 400 tasks across four complexity tiers. We provide an open-vocabulary perception, Assembly-by-Disassembly baseline solution. Our planning-based pipeline outperforms a modern vision-language-action approach across all tiers. The benchmark, simulation environment, and baseline implementation will be released openly to support the broader robotic assembly community.