multimodal prompt design

Designs and evaluates prompt templates, exemplar demonstrations, and input encodings that combine multiple data modalities (e.g., text, images, sketches, audio) and multiple languages to steer a generative model’s behavior. This includes selecting few‑shot examples, structuring multimodal inputs and prompt sequences for ideation and iteration, and specifying constraints or framing to align model outputs with designer intent.

multimodalpromptdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$177K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

PromptMap: Supporting Exploratory Text-to-Image Generation

Oct 03, 2025
YG
Yuhan Guo
🏛️ Peking University | University of Nottingham

In text-to-image generation, users often become disoriented due to the vast design space, struggling to systematically track exploration trajectories, reuse prior creative ideas, or discover novel inspirations. To address this, we propose a *Design Exploration Model* that formalizes the nonlinear creative process as a representable and navigable structure. Based on this model, we design PromptMap—an interactive visualization tool supporting memory of past prompts, identification of promising directions, and iterative decision-making via multi-scale layout, exploration-path tracking, and semantic clustering. Technically, PromptMap integrates diffusion model output analysis, user behavior modeling, and dynamic graph visualization. An empirical user study (N=18) demonstrates that our approach significantly reduces cognitive load (p<0.01), improves exploration completeness (+37%), and increases inspiration reuse rate (+42%). This work provides a scalable methodology and practical framework for AI-augmented creative exploration.

Addressing user disorientation in vast design spacesProviding visual representation for non-linear exploration processesSupporting exploratory text-to-image generation process

Current generative AI image editing tools rely heavily on textual prompts or localized inpainting for fine-grained tasks such as layout and proportion adjustments, often suffering from low efficiency, ambiguity, and cumbersome iteration. This work proposes a multimodal prompting paradigm that integrates freehand sketches with semantic annotations alongside text to form a unified visual-textual input prototype. Through the first systematic user study comparing the efficacy of textual, visual, and hybrid prompts in image refinement, the study demonstrates that visual prompts significantly enhance spatial editing accuracy and efficiency while reducing cognitive load; textual prompts remain better suited for semantic and global adjustments; and hybrid prompting yields the best overall performance. The findings further reveal a strong dependency of optimal modality choice on task type, offering a novel interaction paradigm for generative AI design tools.

design workflowsGenerative AIimage refinement

Exploring Visual Prompts: Refining Images with Scribbles and Annotations in Generative AI Image Tools

Mar 05, 2025
HP
Hyerim Park
🏛️ BMW Group | University of Stuttgart | LMU Munich

Designers struggle to effectively integrate text prompts, annotations, and scribbles—three distinct input modalities—during the image refinement phase in generative AI image tools. Method: We conducted the first systematic comparative study with seven professional designers using a digital-paper prototype, combining contextual interviews and task-based behavioral observation to analyze input strategies, cognitive load, and AI misinterpretation patterns. Results: We identify clear functional boundaries and complementary synergies among modalities: annotations excel at spatial referencing and element identification; scribbles support precise shape and positional control; text prompts best stimulate semantic creativity. Key bottlenecks include frequent AI misinterpretation of visual cues and high cognitive cost in crafting effective text prompts. Grounded in empirical evidence, we propose multimodal prompt design principles that balance expressivity, interpretability, and efficiency—offering both theoretical foundations and practical guidelines for next-generation GenAI design tool interaction paradigms.

Addressing challenges in AI interpretation and effective prompt creationComparing text prompts, annotations, and scribbles for design refinementExploring input methods for refining images in GenAI tools

This paper addresses the challenge of operationalizing generative AI within collaborative software engineering teams. Drawing on a design study with 39 industry experts—including field observations, semi-structured interviews, and multi-role workshops—we systematically investigate how prompt engineering supports cross-functional AI prototyping and iterative co-design. Our study is the first to characterize three core phenomena in collaborative prompt prototyping: (1) the emergent construction of shared coordination norms, (2) dynamic role evolution across developers, domain experts, and AI specialists, and (3) context-sensitive evaluation mechanisms for prompt efficacy. We propose a generative-content-feature-driven rapid iteration paradigm and distill a reusable prompt prototyping strategy framework. Key technical challenges—including model opacity and example overfitting—are empirically identified. The findings provide both methodological grounding and actionable practice guidelines for industrial software teams, advancing the shift from generative AI as a technical capability to a collaborative design enabler.

Addressing challenges like model interpretability in prototypingExploring prompt engineering in generative AI designUnderstanding collaborative team dynamics in AI prototyping

Latest Papers

What's happening recently
View more

In environmental design, generative AI–assisted image creation faces dual challenges: poor local detail control and weak global consistency—long LLM-expanded prompts hinder precise identification of key visual terms, while localized inpainting often disrupts semantic coherence. This paper proposes an interactive image optimization framework centered on a traceable bidirectional mapping mechanism between prompt tokens and image regions. Designers can click any image region to reverse-locate and edit its corresponding prompt token, enabling fine-grained, semantically coherent iterative refinement. The method integrates LLM-based prompt enhancement, text-to-image generation, and controllable localized inpainting. A user study (N=20) and field deployments across two professional studios demonstrate statistically significant improvements in prompt interpretability, editing accuracy, and workflow efficiency (p < .01).

Enhancing controllability over specific visual elements in environment designImproving traceability of AI-generated prompts for image refinementMaintaining global consistency during localized image editing processes

This work addresses the limitations of current generative models, which rely heavily on precise textual prompts and thus struggle to support the open-ended and ambiguous visual exploration typical of early-stage creative ideation. The authors propose a prompt-free, feedforward generative framework that synthesizes semantically meaningful and visually coherent image combinations from only two input images. By eliminating dependence on language, the method constructs training data exclusively from visual triplets and leverages a CLIP-based sparse autoencoder to extract disentangled editing directions from the CLIP latent space, enabling non-literal recombination of visual concepts. This approach empowers designers to conduct intuitive and efficient visual exploration during the initial phases of the creative process, fostering inspiration without the constraints of explicit textual guidance.

generative explorationimage generationnon-literal visual combinations

TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech

Nov 08, 2025
WS
Weiyan Shi
🏛️ Singapore University of Technology and Design | Carnegie Mellon University

Designers face two critical bottlenecks in text-prompted generative AI–assisted ideation: difficulty in prompt formulation and limited expressiveness of visual concepts—both severely impeding cognitive fluency. To address this, we introduce the first embedded multimodal generative AI framework that enables real-time, context-aware, low-latency AI feedback by jointly processing speech and freehand sketch inputs. The system integrates on-device automatic speech recognition, sketch understanding, and multimodal generative models, embedding AI deeply within the design thinking process—not merely optimizing outputs. A user study demonstrates that our approach significantly reduces prompting effort (p < 0.01), increases ideational fluency by 38%, and enhances conceptual diversity by 42%. These results empirically validate the critical value of multimodal interaction—particularly speech and sketch—to early-stage concept generation.

Addresses text-based prompting disrupting creative design flowIntegrates freehand drawing with real-time speech for ideationSupports fluid concept development through multimodal AI responses

This work addresses the creative stagnation often induced by existing generative design tools that directly output complete images. To overcome this limitation, the authors propose a multi-stage, compositional AI-assisted design approach that emulates professional designers’ workflows: it first structurally interprets ambiguous design requests, then generates candidate elements—such as objects, backgrounds, typography, layout, and composition—separately, and finally enables interactive recombination. This method formalizes real-world design processes into a computable system for the first time, decoupling requirement interpretation, element generation, and composition to substantially enhance prompt diversity and alignment with user intent. User studies demonstrate that the system outperforms baseline approaches in both requirement comprehension accuracy and designer-rated quality, revealing a productive trade-off between structured workflow, creative clarity, and efficiency despite slightly longer generation times.

AI-assisted designdesign briefgraphic design

Hot Scholars

MZ

Minfeng Zhu

Zhejiang University
VisualisationMath
XM

Xingjun Ma

Fudan University
Trustworthy AIMultimodal AIGenerative AIEmbodied AI
LL

Linjie Li

Microsoft
Vision and Language
WM

Walid Maalej

University of Hamburg
Generative Software EngineeringAI EngineeringRequirements EngineeringUser Feedback