🤖 AI Summary
This study addresses the limitation of existing AI design tools that overlook verbal context during sketching, thereby failing to capture dynamically evolving creative intent. To this end, we propose a multimodal interaction system that jointly interprets concurrent sketches and speech by introducing a novel dynamic alignment mechanism. Integrating natural language processing with computer vision, the system enables real-time fusion and co-evolution of both modalities throughout the ideation process. User studies demonstrate that the proposed approach significantly enhances human-AI perceptual alignment (p<.05), effectively improving the efficiency of creative expression and the collaborative design experience. This work establishes a new paradigm for multimodal human-AI co-creation.
📝 Abstract
Designers often speak while sketching when explaining ideas, yet AI design tools often rely on sketches or prompts, overlooking context expressed as ideas develop. We developed a sketch-based AI design interface that jointly interprets sketches and concurrent speech. Through a between-subjects study ($N=24$), we examined how speaking while sketching steers human--AI design ideation compared with sketches alone. For creativity support, concurrent speech supported natural expression of design intent and efficient visualisation. For human--AI collaboration, speech helped establish a shared understanding of design intent, supported significantly higher perceived alignment ($p<.05$), and enabled participants to guide AI contributions as ideas co-evolved. We discuss how future human--AI design tools could support dynamic alignment, broader multimodal expression, and human--AI co-creativity.