Score
Design and build prompt templates, template engines, and tooling that let practitioners customize prompts without changing their intended meaning, including methods for anchoring non-custom tokens, managing injected embeddings safely, and enforcing constraint-aware placeholders across language and multimodal (e.g., vision) models. Create progressive, pedagogical, and tutor/coach prompt patterns and workflows — plus evaluation metrics and revision procedures — that support elicitation, constraint handling, iterative prompt refinement, and IDE-integrated tutor experiences.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.
Prompt template design for LLM applications remains largely empirical and lacks systematic, principled methodologies. Method: This paper introduces the first industrial-grade prompt template analysis framework: (1) constructing a high-quality dataset of templates from open-source LLM applications (e.g., Uber, Microsoft), curated via LLM-assisted parsing augmented with human verification; (2) establishing the first structured taxonomy of template components; and (3) conducting component-level statistical modeling and A/B-style instruction-following evaluations. Contribution/Results: We identify frequent co-occurrence patterns among template components and quantify their substantial impact on instruction-following performance—yielding up to a 23.6% accuracy gain. Furthermore, we distill reusable, robust design principles and optimization guidelines. This work provides both theoretical foundations and practical paradigms for prompt engineering, advancing systematic, data-driven template design in production LLM systems.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
This paper addresses the challenge of operationalizing generative AI within collaborative software engineering teams. Drawing on a design study with 39 industry experts—including field observations, semi-structured interviews, and multi-role workshops—we systematically investigate how prompt engineering supports cross-functional AI prototyping and iterative co-design. Our study is the first to characterize three core phenomena in collaborative prompt prototyping: (1) the emergent construction of shared coordination norms, (2) dynamic role evolution across developers, domain experts, and AI specialists, and (3) context-sensitive evaluation mechanisms for prompt efficacy. We propose a generative-content-feature-driven rapid iteration paradigm and distill a reusable prompt prototyping strategy framework. Key technical challenges—including model opacity and example overfitting—are empirically identified. The findings provide both methodological grounding and actionable practice guidelines for industrial software teams, advancing the shift from generative AI as a technical capability to a collaborative design enabler.
To address privacy leakage risks in LLM applications within sensitive domains such as finance, this paper proposes an iterative hard prompt optimization method that operates without exposing task-specific context. The core innovation is a novel few-shot meta-prompting mechanism: leveraging the LLM’s intrinsic meta-reasoning capability over minimal examples, it autonomously generates and iteratively refines prompt templates—achieving performance gains without disclosing proprietary data. The method integrates self-prompt optimization, templated prompt engineering, and iterative propagation, strictly preserving syntactic structure and linguistic style consistency. Experiments across diverse contextual tasks demonstrate an average improvement of 103.87%, significantly enhancing grammatical stability and stylistic fidelity of prompts. This work establishes a new paradigm for compliant, privacy-preserving prompt engineering in high-regulation environments.
Non-expert users struggle to efficiently optimize LLM prompts due to limited domain knowledge and insufficient feedback mechanisms. To address this, we propose a beginner-oriented visual prompt engineering system featuring a novel tri-strategy collaborative optimization framework—integrating keyword perturbation, semantic paraphrasing, and optimal few-shot example recommendation. We design a multi-view synchronized interface, an interactive prompt editing environment, and a real-time evaluation mechanism grounded in both semantic similarity and task-specific accuracy. Experiments demonstrate that our system reduces user prompt iteration time by 37%, increases prompt diversity by 2.1×, and improves average accuracy by 11.4% across multiple NLP tasks—significantly outperforming existing prompt interfaces. This work lowers the cognitive barrier to prompt engineering and establishes a new paradigm for LLM interaction that is interpretable, iterative, and empirically evaluable for non-experts.
This work addresses the challenge that software developers face in effectively learning prompt engineering, a skill whose dynamic, interactive, and context-dependent nature is poorly served by traditional instructional methods. To bridge this gap, the paper introduces Prompt Coach—the first system that combines agent-based tutoring with Socratic questioning embedded directly within the IDE. Prompt Coach delivers context-aware, personalized guidance by analyzing the developer’s current codebase and large language model (LLM) interactions to evaluate prompt quality across multiple dimensions and support iterative refinement. The system integrates LLM behavior analysis, multidimensional evaluation metrics, IDE integration, and an interactive agent mechanism. In a 60-minute study with 15 professional developers, participants demonstrated significant improvements in prompt quality—particularly in commonly overlooked dimensions—and consistently reported enhanced prompt-writing capabilities.
This study addresses the lack of a unified and reproducible taxonomy for “prompt patterns” in existing research. Focusing on single-turn textual prompts, it proposes the first systematic classification comprising 30 distinct and well-defined prompt patterns, organized along two orthogonal dimensions. Through a comprehensive literature review, pattern identification, and taxonomic methodology, the work establishes a structured and reproducible knowledge framework. By standardizing terminology and definitions, this classification provides a foundational reference for prompt engineering, significantly enhancing comparability and reproducibility across related studies.
This work addresses the limitations of current text-to-image prompting methods when users struggle to articulate their visual intent due to ambiguous instructions or unfamiliarity with model capabilities. To overcome this, the authors propose Adaptive Prompt Elicitation (APE), a novel approach that integrates language model priors with an information-theoretic framework to dynamically generate interpretable visual queries. These queries guide users in iteratively refining their intentions while automatically compiling high-quality prompts. By moving beyond conventional text-only prompting paradigms, APE significantly improves alignment between user intent and generated outputs, as demonstrated on the IDEA-Bench and DesignBench benchmarks. User studies further reveal a 19.8% improvement in intent alignment for complex tasks without imposing additional cognitive load.
This study addresses the lack of systematic, evidence-driven approaches for evaluating the effectiveness of large language model prompts in educational contexts, where balancing personalization and pedagogical alignment remains challenging. The authors propose a generalizable prompt evaluation framework that integrates six pedagogically informed prompt templates designed to generate follow-up questions within structured dialogues. For the first time in educational prompt engineering, they introduce tournament-style evaluation combined with the Glicko-2 rating system, complemented by multidimensional human assessments—covering format, conversational support, and learner adaptability—and validated through real user interaction data. Across 120 authentic interactions, a prompt template incorporating role specification, contextual management, and metacognitive strategies significantly outperformed others, achieving pairwise win rates of 81%–100%, thereby advancing prompt design from an intuition-based practice toward an evidence-driven paradigm.