Score
Designs and evaluates prompt templates, exemplar demonstrations, and input encodings that combine multiple data modalities (e.g., text, images, sketches, audio) and multiple languages to steer a generative model’s behavior. This includes selecting few‑shot examples, structuring multimodal inputs and prompt sequences for ideation and iteration, and specifying constraints or framing to align model outputs with designer intent.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
In text-to-image generation, users often become disoriented due to the vast design space, struggling to systematically track exploration trajectories, reuse prior creative ideas, or discover novel inspirations. To address this, we propose a *Design Exploration Model* that formalizes the nonlinear creative process as a representable and navigable structure. Based on this model, we design PromptMap—an interactive visualization tool supporting memory of past prompts, identification of promising directions, and iterative decision-making via multi-scale layout, exploration-path tracking, and semantic clustering. Technically, PromptMap integrates diffusion model output analysis, user behavior modeling, and dynamic graph visualization. An empirical user study (N=18) demonstrates that our approach significantly reduces cognitive load (p<0.01), improves exploration completeness (+37%), and increases inspiration reuse rate (+42%). This work provides a scalable methodology and practical framework for AI-augmented creative exploration.
Current generative AI image editing tools rely heavily on textual prompts or localized inpainting for fine-grained tasks such as layout and proportion adjustments, often suffering from low efficiency, ambiguity, and cumbersome iteration. This work proposes a multimodal prompting paradigm that integrates freehand sketches with semantic annotations alongside text to form a unified visual-textual input prototype. Through the first systematic user study comparing the efficacy of textual, visual, and hybrid prompts in image refinement, the study demonstrates that visual prompts significantly enhance spatial editing accuracy and efficiency while reducing cognitive load; textual prompts remain better suited for semantic and global adjustments; and hybrid prompting yields the best overall performance. The findings further reveal a strong dependency of optimal modality choice on task type, offering a novel interaction paradigm for generative AI design tools.
本文提出一种生成图像模型的可视化界面设计空间,通过分解用户界面、可控模型对象及映射函数来分析不同交互技术,以促进新界面设计。
Designers struggle to effectively integrate text prompts, annotations, and scribbles—three distinct input modalities—during the image refinement phase in generative AI image tools. Method: We conducted the first systematic comparative study with seven professional designers using a digital-paper prototype, combining contextual interviews and task-based behavioral observation to analyze input strategies, cognitive load, and AI misinterpretation patterns. Results: We identify clear functional boundaries and complementary synergies among modalities: annotations excel at spatial referencing and element identification; scribbles support precise shape and positional control; text prompts best stimulate semantic creativity. Key bottlenecks include frequent AI misinterpretation of visual cues and high cognitive cost in crafting effective text prompts. Grounded in empirical evidence, we propose multimodal prompt design principles that balance expressivity, interpretability, and efficiency—offering both theoretical foundations and practical guidelines for next-generation GenAI design tool interaction paradigms.
This paper addresses the challenge of operationalizing generative AI within collaborative software engineering teams. Drawing on a design study with 39 industry experts—including field observations, semi-structured interviews, and multi-role workshops—we systematically investigate how prompt engineering supports cross-functional AI prototyping and iterative co-design. Our study is the first to characterize three core phenomena in collaborative prompt prototyping: (1) the emergent construction of shared coordination norms, (2) dynamic role evolution across developers, domain experts, and AI specialists, and (3) context-sensitive evaluation mechanisms for prompt efficacy. We propose a generative-content-feature-driven rapid iteration paradigm and distill a reusable prompt prototyping strategy framework. Key technical challenges—including model opacity and example overfitting—are empirically identified. The findings provide both methodological grounding and actionable practice guidelines for industrial software teams, advancing the shift from generative AI as a technical capability to a collaborative design enabler.
In environmental design, generative AI–assisted image creation faces dual challenges: poor local detail control and weak global consistency—long LLM-expanded prompts hinder precise identification of key visual terms, while localized inpainting often disrupts semantic coherence. This paper proposes an interactive image optimization framework centered on a traceable bidirectional mapping mechanism between prompt tokens and image regions. Designers can click any image region to reverse-locate and edit its corresponding prompt token, enabling fine-grained, semantically coherent iterative refinement. The method integrates LLM-based prompt enhancement, text-to-image generation, and controllable localized inpainting. A user study (N=20) and field deployments across two professional studios demonstrate statistically significant improvements in prompt interpretability, editing accuracy, and workflow efficiency (p < .01).
This work addresses the limitations of current generative models, which rely heavily on precise textual prompts and thus struggle to support the open-ended and ambiguous visual exploration typical of early-stage creative ideation. The authors propose a prompt-free, feedforward generative framework that synthesizes semantically meaningful and visually coherent image combinations from only two input images. By eliminating dependence on language, the method constructs training data exclusively from visual triplets and leverages a CLIP-based sparse autoencoder to extract disentangled editing directions from the CLIP latent space, enabling non-literal recombination of visual concepts. This approach empowers designers to conduct intuitive and efficient visual exploration during the initial phases of the creative process, fostering inspiration without the constraints of explicit textual guidance.
本文提出InterIL模型,通过联合生成背景图像和布局来解决图形设计模板创建问题,利用双向交互模块改进组合和谐度。
Designers face two critical bottlenecks in text-prompted generative AI–assisted ideation: difficulty in prompt formulation and limited expressiveness of visual concepts—both severely impeding cognitive fluency. To address this, we introduce the first embedded multimodal generative AI framework that enables real-time, context-aware, low-latency AI feedback by jointly processing speech and freehand sketch inputs. The system integrates on-device automatic speech recognition, sketch understanding, and multimodal generative models, embedding AI deeply within the design thinking process—not merely optimizing outputs. A user study demonstrates that our approach significantly reduces prompting effort (p < 0.01), increases ideational fluency by 38%, and enhances conceptual diversity by 42%. These results empirically validate the critical value of multimodal interaction—particularly speech and sketch—to early-stage concept generation.
This work addresses the creative stagnation often induced by existing generative design tools that directly output complete images. To overcome this limitation, the authors propose a multi-stage, compositional AI-assisted design approach that emulates professional designers’ workflows: it first structurally interprets ambiguous design requests, then generates candidate elements—such as objects, backgrounds, typography, layout, and composition—separately, and finally enables interactive recombination. This method formalizes real-world design processes into a computable system for the first time, decoupling requirement interpretation, element generation, and composition to substantially enhance prompt diversity and alignment with user intent. User studies demonstrate that the system outperforms baseline approaches in both requirement comprehension accuracy and designer-rated quality, revealing a productive trade-off between structured workflow, creative clarity, and efficiency despite slightly longer generation times.