spatial prompting

Designs, builds, and evaluates prompting and conditioning techniques that encode spatial layouts, semantic labels, and structured constraints to guide multimodal models or agents. This work combines sketches, language, and visual targets to preserve user-authored spatial scaffolds, enforce layout conditioning, and deliver rubric-like or structured prompts for controlled generation, grounded reference, or stepwise guided interaction.

spatialprompting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing approaches that integrate large language models with spatial layouts struggle to support progressive refinement, often suffering from misalignment between user interaction and model revisions, inconsistent human–AI intent, and insufficient fine-grained customization. This work proposes S-PRISM, a semantic prompting framework that uniquely unifies spatial semantic interaction awareness with user intent inference. By introducing a position-aware revision mechanism, S-PRISM enables interactive, incremental, and customizable narrative generation. Experimental results demonstrate that S-PRISM significantly improves revision accuracy and alignment between human and model intent. An empirical study with 14 users further validates that the system effectively supports progressive narrative formalization, exhibiting high flexibility, efficiency, and trustworthiness.

human-LLM alignmentincremental narrativesemantic interaction

PromptMap: Supporting Exploratory Text-to-Image Generation

Oct 03, 2025
YG
Yuhan Guo
🏛️ Peking University | University of Nottingham

In text-to-image generation, users often become disoriented due to the vast design space, struggling to systematically track exploration trajectories, reuse prior creative ideas, or discover novel inspirations. To address this, we propose a *Design Exploration Model* that formalizes the nonlinear creative process as a representable and navigable structure. Based on this model, we design PromptMap—an interactive visualization tool supporting memory of past prompts, identification of promising directions, and iterative decision-making via multi-scale layout, exploration-path tracking, and semantic clustering. Technically, PromptMap integrates diffusion model output analysis, user behavior modeling, and dynamic graph visualization. An empirical user study (N=18) demonstrates that our approach significantly reduces cognitive load (p<0.01), improves exploration completeness (+37%), and increases inspiration reuse rate (+42%). This work provides a scalable methodology and practical framework for AI-augmented creative exploration.

Addressing user disorientation in vast design spacesProviding visual representation for non-linear exploration processesSupporting exploratory text-to-image generation process

This study addresses the challenge of accurately interpreting users’ multimodal spatial reference expressions—combining speech and gestures—in immersive environments. Drawing on spatial cognition theory, the authors analyze unconstrained multimodal interactions in virtual reality through a Wizard-of-Oz experiment. They formally introduce, for the first time, a triadic structure for spatial referencing composed of Source, Anchor, and Frame, along with associated composition strategies. Building upon this foundation, they develop a multimodal interaction pipeline integrating large language models to enable generalizable and extensible spatial reasoning. The resulting prototype system demonstrates improved accuracy and practicality in understanding user spatial intent, as validated through technical evaluation.

multimodal interactionreference-based manipulationspatial reasoning

PromptMap: An Alternative Interaction Style for AI-Based Image Generation

Mar 12, 2025
KA
Krzysztof Adamkiewicz
🏛️ Lodz University of Technology | TU Wien | German Research Center for Artificial Intelligence

To address the challenge novice users face in crafting effective prompts for text-to-image generation, this paper introduces a semantic map–based interactive paradigm. It constructs a large-scale, high-quality prompt–image pair repository (containing over one million entries, automatically generated by LLMs), integrated with CLIP-based cross-modal embeddings and semantic clustering to enable multi-scale visual navigation and prompt exploration. This work is the first to deeply unify LLM-driven prompt synthesis, semantic embedding–based clustering, and a scalable map-style interface, facilitating human–AI collaborative prompt engineering. A user study (n = 72) demonstrates that our system significantly improves prompt authoring efficiency (+63%) and generation satisfaction (p < 0.01) over baseline tools. Furthermore, it empirically validates the feasibility and practical utility of large-scale, navigable prompt knowledge bases for generative AI applications.

Difficulty in crafting effective prompts for AI-based image generation.Need for tools to help users explore and discover relevant prompts.Novice users struggle with text-to-image AI interaction.

Latest Papers

What's happening recently
View more

Existing text-to-3D generation methods struggle to accurately convey users’ spatial intent regarding part placement, inter-part relationships, and overall structure. This work proposes a novel interactive paradigm that treats coarse 3D sketches drawn in virtual reality as spatial prompts—rather than geometry to be reconstructed—and combines them with natural language instructions to jointly guide state-of-the-art multimodal 3D generative models. By leveraging sketch segmentation, multi-view part guidance, and structured prompt engineering, the approach effectively channels user-defined spatial skeletons into the generation process. Experiments on 20 open-domain examples demonstrate that the method significantly outperforms baselines relying solely on text or sketches, while a user study confirms its capability to support efficient initial creation and flexible subsequent editing.

3D content creationpart relationshipsspatial intent

Existing unified multimodal large language models struggle to achieve controllable image generation under complex spatial instructions and logical constraints. This work proposes ATLAS, a novel framework that introduces, for the first time in a unified model, a human-like “think–plan–draw” three-stage paradigm, using layout as a shared representation to jointly perform spatial reasoning, object arrangement planning, and image rendering. The framework supports layout alignment, instruction editing, and multimodal grounding, and introduces the ATLAS-Reasoning evaluation benchmark. Experimental results demonstrate that ATLAS significantly outperforms current state-of-the-art methods in image generation, achieving an average improvement of 65.31% over the best existing layout-based unified models, and a 23.06% gain on spatial reasoning tasks.

controllable image generationlayout-aware reasoninglogical constraints

This study addresses the challenge of translating users’ intuitive spatial expressions into executable constraints for controllable, collaborative generative 3D design. To this end, we propose an XR system leveraging the Apple Vision Pro that integrates spatial sketches drawn with a Logitech Muse 3D pen and spoken language prompts, jointly encoding them for the first time as actionable constraints for an AI generative model (Meshy). The approach enables multiple users to synchronously collaborate and iteratively refine designs within a shared 3D space. Experimental results demonstrate the effectiveness of this workflow in enhancing design intuitiveness and fostering group consensus, while also highlighting the need for further improvements in generative efficiency and clarity of system feedback.

Collaborative CreationExecutable ConstraintsExtended Reality

Existing unified multimodal models exhibit limitations in fine-grained compositional understanding and controllable generation. This work proposes COMPASS, a novel framework that unifies composition-aware perception and generation within a single system for the first time, leveraging a shared expert token τ_c as an anchor for compositional intent to shift from passive analysis to explicit layout control. Built upon a Mixture-of-Experts (MoE) backbone, the approach introduces lightweight composition experts, expert token distillation, global conditioning via denoising trajectories, and an inference-enhanced compositional annotation strategy. Additionally, the authors curate Comp-11, a large-scale compositional instruction dataset. Experiments demonstrate that COMPASS significantly advances category-level compositional understanding and generates outputs with superior composition consistency and prompt fidelity compared to strong baselines.

compositioncomposition recognitioncontrollable generation

This work addresses the high sensitivity of text-to-image generation to prompt phrasing, where semantically similar descriptions often yield inconsistent outputs due to linguistic variations. To mitigate this, the authors propose APE, a lightweight prompt enhancement framework that leverages deployable small language models for prompt rewriting. APE supports both a single-agent variant (SAPE) and a role-specialized multi-agent collaborative approach (MAPE), improving prompt-generation consistency without modifying downstream visual models. Through a task-aware reward mechanism and a structured pipeline of routing, rewriting, and composition, APE significantly outperforms baseline models across multiple image generation and editing benchmarks. Notably, MAPE excels in complex compositional tasks, effectively narrowing the performance gap with proprietary large-model enhancers.

image editingimage generationlanguage model dependency

Hot Scholars

AM

Ao Ma

JD.com
Generative AIVideo Generation
SZ

Songyan Zhang

Nanyang Technology University
Computer VisionAutonomous Driving
YY

Yuhui Yuan

Canva CORE, ex-Microsoft Research Asia
Generative AI + DesignComputer Vision
GB

Guanqun Bi

Tsinghua University; UCAS
Social AgentsNatural Language Generation
ZC

Zhuang Chen

中南大学计算机学院
Natural Language ProcessingSocial IntelligenceComputational Psychology