🤖 AI Summary
This study addresses the challenges of poor semantic coherence across multi-turn interactions and difficulty in constructing complex scenes in existing robotic painting systems. We propose a hierarchical planning framework for progressive human-robot co-creation on physical canvases. This framework introduces a novel multi-stage structured reasoning mechanism that effectively decouples high-level semantic intent inference, mid-level spatial localization, and low-level stroke execution control, achieving a paradigm shift from single-pass rendering to continuous content generation. Experimental results demonstrate that our method significantly outperforms baseline models in semantic alignment, spatial stability, and action plausibility, substantially enhancing both the robustness and expressiveness of interactive physical painting.
📝 Abstract
Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coherent multi-round collaboration or to construct complex, content-rich scenes over time. We present CoBrush, a hierarchical framework that formulates multi-round co-painting as a coordinated semantic, spatial, and execution process. By separating high-level intent inference from spatial grounding and stroke-level control, the system supports progressive scene development on real acrylic canvases. We evaluate the framework through real human-robot painting sessions, stress tests, and user studies. Compared to single-turn baselines, our approach achieves stronger semantic alignment, more stable spatial progression, and higher perceived plausibility of robot actions. These results demonstrate that structured multi-stage reasoning improves the coherence and robustness of interactive painting and supports the progressive development of content-rich physical artworks.