🤖 AI Summary
Designers face two critical bottlenecks in text-prompted generative AI–assisted ideation: difficulty in prompt formulation and limited expressiveness of visual concepts—both severely impeding cognitive fluency. To address this, we introduce the first embedded multimodal generative AI framework that enables real-time, context-aware, low-latency AI feedback by jointly processing speech and freehand sketch inputs. The system integrates on-device automatic speech recognition, sketch understanding, and multimodal generative models, embedding AI deeply within the design thinking process—not merely optimizing outputs. A user study demonstrates that our approach significantly reduces prompting effort (p < 0.01), increases ideational fluency by 38%, and enhances conceptual diversity by 42%. These results empirically validate the critical value of multimodal interaction—particularly speech and sketch—to early-stage concept generation.
📝 Abstract
Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output.