🤖 AI Summary
This study addresses the challenge of simultaneously achieving long-horizon coordination and fine-grained execution in multi-robot systems by proposing a hierarchical cooperative framework based on semantic communication. The proposed approach decouples the high-level reasoning and planning capabilities of vision-language models (VLMs) from the low-level execution facilitated by vision-language-action models (VLAs), while leveraging semantic communication techniques to enable efficient collaboration among distributed agents. Key contributions include the first VLM-VLA hierarchical architecture integrated with semantic communication and the introduction of the RoboPoly benchmark dataset. Experimental results demonstrate that this method significantly enhances multi-robot task performance, validating the effectiveness of both the hierarchical orchestration strategy and the underlying semantic communication mechanisms.
📝 Abstract
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.