prompt engineering for visualization

Designs and authors textual prompts that drive the generation of visual outputs (images, diagrams, or visual scenes) from models or generative systems. This includes translating conceptual or data-driven representations into scene descriptions, specifying composition, style, and fidelity through wording, and iteratively refining prompts to control visual attributes and rendering outcomes.

promptengineeringforvisualization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.62
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

PromptMap: Supporting Exploratory Text-to-Image Generation

Oct 03, 2025
YG
Yuhan Guo
🏛️ Peking University | University of Nottingham

In text-to-image generation, users often become disoriented due to the vast design space, struggling to systematically track exploration trajectories, reuse prior creative ideas, or discover novel inspirations. To address this, we propose a *Design Exploration Model* that formalizes the nonlinear creative process as a representable and navigable structure. Based on this model, we design PromptMap—an interactive visualization tool supporting memory of past prompts, identification of promising directions, and iterative decision-making via multi-scale layout, exploration-path tracking, and semantic clustering. Technically, PromptMap integrates diffusion model output analysis, user behavior modeling, and dynamic graph visualization. An empirical user study (N=18) demonstrates that our approach significantly reduces cognitive load (p<0.01), improves exploration completeness (+37%), and increases inspiration reuse rate (+42%). This work provides a scalable methodology and practical framework for AI-augmented creative exploration.

Addressing user disorientation in vast design spacesProviding visual representation for non-linear exploration processesSupporting exploratory text-to-image generation process

VisualPrompter: Prompt Optimization with Visual Feedback for Text-to-Image Synthesis

Jun 29, 2025
SW
Shiyu Wu
🏛️ Institute of Automation, Chinese Academy of Sciences | University of Chinese Academy of Sciences | Beijing Academy of Artificial Intelligence

In text-to-image generation, a semantic gap exists between user prompts and model priors, yielding aesthetically pleasing yet semantically inaccurate images. To address this, we propose a training-free, vision-feedback-driven prompt optimization framework. Our method employs an automatic self-reflection module to localize missing concepts in generated images and leverages CLIP-based semantic analysis for fine-grained, goal-directed prompt revision. Designed as a plug-and-play module, it requires no model fine-tuning and is compatible with mainstream diffusion models. Evaluated on multiple semantic alignment benchmarks, our approach achieves state-of-the-art performance: it significantly improves content accuracy while preserving visual fidelity. The core innovations are (1) a training-free online self-reflection mechanism that dynamically identifies semantic deficiencies, and (2) a semantic–visual co-optimization paradigm that jointly refines prompts and image generation through cross-modal feedback.

Bridges gap between user and model-preferred promptsImproves semantic alignment in text-to-image synthesisOptimizes prompts without training for better image quality

What Do You Want? User-centric Prompt Generation for Text-to-image Synthesis via Multi-turn Guidance

Aug 23, 2024
YL
Yilun Liu
🏛️ Huawei | Huawei Canada | Waseda University

In text-to-image synthesis (TIS), novice users struggle to generate desired outputs due to limited prompt engineering skills, while existing automatic prompt generation methods lack interpretability and interactivity. To address this, we propose DialPrompt—the first multi-turn conversational prompt generation framework designed specifically for novices. It guides users iteratively to clarify 15 key visual attribute dimensions (e.g., style, composition, lighting), enabling interpretable mapping between prompt elements and visual attributes, as well as real-time user intervention. DialPrompt is trained on a novel, self-constructed multi-turn prompt optimization dataset and integrates preference-aware conditional generation with feedback-driven iterative refinement. Experiments demonstrate that DialPrompt improves image quality by 5.7% over state-of-the-art prompt engineering baselines, increases user-centeredness scores by 46.5%, and achieves an expert overall rating of 7.9/10.

Enhancing interactivity through multi-turn dialogue guidanceGenerating user-friendly prompts for text-to-image synthesisImproving novice users' control over image generation process

This work addresses the challenges of semantic distortion and structural inconsistency in generating images from complex textual prompts involving multiple objects with specified attributes, quantities, and spatial relationships. The authors propose a scene graph–based, zero-shot soft visual guidance mechanism that leverages a lightweight language model during inference to produce conditional signals that steer a diffusion model. The key innovation lies in the ASQL Conditioner module, which enables the first unified zero-shot conditioning framework jointly modeling Attribute, Size, Quantity, and Location. This approach significantly enhances semantic fidelity and structural coherence in generated images under complex prompts while preserving output diversity and computational efficiency.

complex promptsscene graphsemantic fidelity

Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study

Mar 24, 2025
RZ
R. Ziatdinov
🏛️ Keimyung University | Ufa State Petroleum Technological University

This work addresses the inconsistent modeling quality of current text-to-3D and image-to-3D generative tools. We present the first systematic, critical evaluation of mainstream online generators—including diffusion-based, NeRF-based, and implicit surface reconstruction methods—assessing their real-world output performance. Using a multi-prompt diversity benchmark coupled with a hybrid human-and-automated evaluation framework, we identify pervasive geometric distortions in complex topologies and fine-grained structures, and uncover key bottlenecks at the intersection of prompt engineering, geometric fidelity, and semantic consistency. Our contributions include actionable prompt design principles and a standardized set of quality evaluation metrics. These provide an empirical benchmark for applications such as digital twins and advance the development of next-generation generative 3D modeling technologies.

Automating 3D model generation to save timeComparing AI-based 3D model creation methodsEvaluating quality of text-to-3D and image-to-3D tools

Latest Papers

What's happening recently
View more

This work addresses the challenges in text-to-image generation—namely, modeling difficulty, detail loss, and limited editability—stemming from the absence of intermediate semantic representations. To this end, the authors propose Visual Prompt Engineering (VPE), a unified framework that autoregressively generates semantic visual tokens (e.g., SigLIP features) as “visual prompts” and uses them to condition diffusion-based image synthesis, thereby tightly coupling semantic planning with image generation. VPE is the first approach to seamlessly integrate intermediate semantic guidance within a single-stage model, circumventing the information bottleneck inherent in conventional two-stage pipelines. Experiments demonstrate that, at comparable model scales, VPE substantially improves editing fidelity (PSNR: 26.76 vs. 19.92), accelerates convergence, elevates the upper bound of generation quality, and supports diverse tasks including class-conditional generation, text-to-image synthesis, and image editing.

diffusion modelsimage editingsemantic representation

I Prompt, it Generates, we Negotiate. Exploring Text-Image Intertextuality in Human-AI Co-Creation of Visual Narratives with VLMs

Nov 05, 2025
MG
Mengyao Guo
🏛️ Harbin Institute of Technology | The University of Sydney | Hong Kong Polytechnic University | Aarhus University | Tsinghua University | University of Illinois at Urbana-Champaign | Swinburne University of Technology

This study investigates the intertextuality mechanisms between human-authored textual intent and AI-generated images in collaborative visual storytelling involving novice users and vision-language models (VLMs). Using GPT-4o’s image generation capability, we conducted a three-phase qualitative study integrated with fuzzy-set qualitative comparative analysis (fsQCA) to identify three core collaborative strategies: prompt iteration, semantic expansion, and multimodal complementarity. We propose a theoretical framework of “text–image intertextuality,” characterizing four collaborative patterns and three empirically derived pathways to successful collaboration—namely, the Educational Collaborator, Technical Expert, and Visual Thinker. Findings demonstrate that AI-induced semantic overflow positively enhances creative ideation, while revealing critical challenges: insufficient cultural representation, weak visual consistency, and difficulties in narrative translation. The work provides empirical grounding and interface-design implications for developing human-centered, role-adaptive AI assistants in creative authoring contexts.

Addressing cultural gaps and visual consistency in AI-generated narrativesExploring text-image intertextuality in human-AI visual narrative co-creationInvestigating how novices navigate sequential visual storytelling with VLMs

This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.

documentationevaluationprompt engineering

Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation

Dec 03, 2025
SK
Subin Kim
🏛️ KAIST | POSTECH | Adobe | Meta

In text-to-visual generation, misalignment between user intent and generated outputs remains a persistent challenge, with single-step generation often failing to meet precise requirements. To address this, we propose PRIS (Prompt Revision via Inference-time Self-refinement), the first framework enabling joint scaling of prompt optimization and visual generation. PRIS establishes a closed-loop prompt refinement mechanism comprising adaptive prompt revision, feedback-driven analysis of generated outputs, and an element-level factual correctness validator. Crucially, this validator enables fine-grained, interpretable alignment assessment. Evaluated on both text-to-image and text-to-video generation tasks, PRIS achieves substantial quality improvements—yielding a 15% average score gain on the VBench 2.0 benchmark. These results empirically validate the effectiveness and generalizability of co-optimizing prompts and generation during inference.

Addresses misalignment between text prompts and generated visuals in text-to-image/video modelsImproves fine-grained attribute matching through adaptive prompt redesignOvercomes quality plateaus from fixed prompts during inference-time scaling

This work addresses the creative stagnation often induced by existing generative design tools that directly output complete images. To overcome this limitation, the authors propose a multi-stage, compositional AI-assisted design approach that emulates professional designers’ workflows: it first structurally interprets ambiguous design requests, then generates candidate elements—such as objects, backgrounds, typography, layout, and composition—separately, and finally enables interactive recombination. This method formalizes real-world design processes into a computable system for the first time, decoupling requirement interpretation, element generation, and composition to substantially enhance prompt diversity and alignment with user intent. User studies demonstrate that the system outperforms baseline approaches in both requirement comprehension accuracy and designer-rated quality, revealing a productive trade-off between structured workflow, creative clarity, and efficiency despite slightly longer generation times.

AI-assisted designdesign briefgraphic design

Hot Scholars

HD

Henghui Ding

Fudan University
Computer VisionMachine LearningSegmentationAIGC
CM

Chia-Mu Yu

National Yang Ming Chiao Tung University
AI SecurityData PrivacyData AnonymizationCryptography
FS

Fahad Shahbaz Khan

MBZUAI, Linköping University Sweden
Computer VisionObject RecognitionGenerative AIAI for Science
BZ

Baobao Zhang

Syracuse University
Political SciencePublic PolicyTechnology Policy
TH

Tsun-Hsuan Wang

Massachusetts Institute of Technology
roboticsmachine learningsimulation