TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech

📅 2025-11-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Designers face two critical bottlenecks in text-prompted generative AI–assisted ideation: difficulty in prompt formulation and limited expressiveness of visual concepts—both severely impeding cognitive fluency. To address this, we introduce the first embedded multimodal generative AI framework that enables real-time, context-aware, low-latency AI feedback by jointly processing speech and freehand sketch inputs. The system integrates on-device automatic speech recognition, sketch understanding, and multimodal generative models, embedding AI deeply within the design thinking process—not merely optimizing outputs. A user study demonstrates that our approach significantly reduces prompting effort (p < 0.01), increases ideational fluency by 38%, and enhances conceptual diversity by 42%. These results empirically validate the critical value of multimodal interaction—particularly speech and sketch—to early-stage concept generation.

Technology Category

Humans and AI: Game Design — Procedural Content Generation & StorytellingNatural Language Processing: GenerationMachine Learning: Multimodal Learning

Application Category

Economics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactionsSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSearch and Retrieval-Augmented AI: Assisted, interactive, and conversational search
📝 Abstract
Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output.
Problem

Research questions and friction points this paper is trying to address.

Addresses text-based prompting disrupting creative design flow
Integrates freehand drawing with real-time speech for ideation
Supports fluid concept development through multimodal AI responses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates freehand drawing with real-time speech input
Captures verbal descriptions during sketching process
Generates context-aware AI responses for fluid ideation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Weiyan Shi
Singapore University of Technology and Design, Singapore
S
Sunaya Upadhyay
Carnegie Mellon University, United States
G
Geraldine Quek
Singapore University of Technology and Design, Singapore
K
K. Choo
Singapore University of Technology and Design, Singapore