Score
Designs and iterates structured question and prompt templates that translate tasks and inputs into model-ready prompts or assessment items, including layout of slots, instruction wording, and incorporation of few‑shot examples or encoded inputs. Builds and evaluates template variants for coherence, information sensitivity, and empirical effectiveness, producing reusable prompt artifacts and associated evaluation procedures.
This study addresses the lack of a unified and reproducible taxonomy for “prompt patterns” in existing research. Focusing on single-turn textual prompts, it proposes the first systematic classification comprising 30 distinct and well-defined prompt patterns, organized along two orthogonal dimensions. Through a comprehensive literature review, pattern identification, and taxonomic methodology, the work establishes a structured and reproducible knowledge framework. By standardizing terminology and definitions, this classification provides a foundational reference for prompt engineering, significantly enhancing comparability and reproducibility across related studies.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.
Prompt template design for LLM applications remains largely empirical and lacks systematic, principled methodologies. Method: This paper introduces the first industrial-grade prompt template analysis framework: (1) constructing a high-quality dataset of templates from open-source LLM applications (e.g., Uber, Microsoft), curated via LLM-assisted parsing augmented with human verification; (2) establishing the first structured taxonomy of template components; and (3) conducting component-level statistical modeling and A/B-style instruction-following evaluations. Contribution/Results: We identify frequent co-occurrence patterns among template components and quantify their substantial impact on instruction-following performance—yielding up to a 23.6% accuracy gain. Furthermore, we distill reusable, robust design principles and optimization guidelines. This work provides both theoretical foundations and practical paradigms for prompt engineering, advancing systematic, data-driven template design in production LLM systems.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
In knowledge-dependent question answering, zero-shot chain-of-thought (CoT) prompting overemphasizes reasoning trace generation while neglecting explicit knowledge acquisition. Method: This paper proposes PREP, a two-stage prompting framework: (1) a language model (LM) actively pre-extracts problem-relevant knowledge, and (2) a second LM generates the answer conditioned on this extracted knowledge. Contribution/Results: PREP introduces the first “knowledge extraction–reasoning” decoupled dual-LM collaborative prompting paradigm, requiring no domain-specific prompt engineering and exhibiting strong generalization. Evaluated on a newly constructed component–material dataset and three public commonsense reasoning benchmarks under zero-shot settings, PREP significantly outperforms zero-shot CoT and other baselines, achieving consistent average accuracy gains. Results empirically validate that explicit knowledge pre-extraction effectively enhances both instruction following and knowledge retrieval capabilities.
This study addresses the lack of systematic, evidence-driven approaches for evaluating the effectiveness of large language model prompts in educational contexts, where balancing personalization and pedagogical alignment remains challenging. The authors propose a generalizable prompt evaluation framework that integrates six pedagogically informed prompt templates designed to generate follow-up questions within structured dialogues. For the first time in educational prompt engineering, they introduce tournament-style evaluation combined with the Glicko-2 rating system, complemented by multidimensional human assessments—covering format, conversational support, and learner adaptability—and validated through real user interaction data. Across 120 authentic interactions, a prompt template incorporating role specification, contextual management, and metacognitive strategies significantly outperformed others, achieving pairwise win rates of 81%–100%, thereby advancing prompt design from an intuition-based practice toward an evidence-driven paradigm.
To address privacy leakage risks in LLM applications within sensitive domains such as finance, this paper proposes an iterative hard prompt optimization method that operates without exposing task-specific context. The core innovation is a novel few-shot meta-prompting mechanism: leveraging the LLM’s intrinsic meta-reasoning capability over minimal examples, it autonomously generates and iteratively refines prompt templates—achieving performance gains without disclosing proprietary data. The method integrates self-prompt optimization, templated prompt engineering, and iterative propagation, strictly preserving syntactic structure and linguistic style consistency. Experiments across diverse contextual tasks demonstrate an average improvement of 103.87%, significantly enhancing grammatical stability and stylistic fidelity of prompts. This work establishes a new paradigm for compliant, privacy-preserving prompt engineering in high-regulation environments.
This study addresses the challenge of low response quality and excessive user interaction in open-ended tasks due to ambiguous prompts in large language models. The authors propose a checklist-style prompting method and systematically evaluate its impact on response quality and user effort across four task categories, comparing it against baseline and clarification-based prompting strategies. Using a unified rubric assessing task completion, correctness, adherence to instructions, and clarity, experiments conducted on ChatGPT, Claude, and Grok demonstrate that checklist prompting achieves an average score of 7.50 out of 8—significantly outperforming both baseline prompting (5.67) and clarification-based prompting (6.67). Moreover, this approach reduces both the number of interaction turns and input token consumption, thereby achieving a superior trade-off between output quality and interaction efficiency.
This study addresses the inefficiency of traditional programming exam design and its limited capacity to holistically assess students’ creativity, problem-solving skills, and domain knowledge. It presents the first systematic application of prompt engineering to the automatic generation of programming examination questions, proposing a method that leverages carefully crafted, diverse prompt templates to guide ChatGPT—without requiring fine-tuning of large language models. The approach autonomously produces high-quality questions and reference solutions spanning theoretical and practical aspects, multiple question types, and varying difficulty levels. Experimental results demonstrate that the generated items match or exceed the quality of human-authored questions while substantially improving item development efficiency. User studies further confirm the method’s effectiveness and practical value in educational settings.
This study addresses the misalignment between AI-generated educational content—such as quiz items produced by the OneClickQuiz Moodle plugin—and Bloom’s taxonomy cognitive levels (Remembering, Applying, Analyzing). We propose a lightweight yet high-precision prompt engineering strategy that departs from conventional concise or role-based prompts. Specifically, we design three structured prompt variants explicitly embedding cognitive objectives, action verbs, and answer constraints. Using an automated classification model augmented by expert human evaluation, we systematically assess the cognitive-level fidelity of generated questions. Experimental results demonstrate that such detailed prompts significantly improve alignment accuracy across all targeted levels—particularly in Applying and Analyzing—without requiring model fine-tuning or additional training. This work contributes a reusable, empirically validated prompt design paradigm for low-cost, highly controllable AI-generated educational content.
This study addresses the challenges in evaluating large language model (LLM) applications—namely, high output stochasticity, multidimensionality, and sensitivity to prompt and model variations—which render traditional testing methods inadequate. The authors propose an evaluation-driven engineering workflow (Define-Test-Diagnose-Fix) and introduce the first hierarchical Minimum Viable Evaluation Suite (MVES) tailored for general-purpose LLMs, retrieval-augmented generation (RAG), and agent-based tool-use scenarios. The framework integrates automated checks, human scoring, and LLM-as-judge to establish a reproducible local evaluation system, validated on the Ollama platform using Llama 3 8B and Qwen 2.5 7B Instruct models. Experiments reveal that while generic prompt templates enhance instruction following, they degrade structured extraction accuracy from 100% to 90% and RAG compliance from 93.3% to 80%, underscoring the necessity of evaluation-driven iteration and advocating systematic assessment over heuristic prompt engineering.
This study addresses the limited reasoning capabilities of small language models (SLMs) in multi-hop question answering by systematically evaluating 24 prompt templates on the HotpotQA dataset. The evaluation encompasses standard RAG prompts, nine existing prompting strategies, and fourteen newly designed hybrid prompts, with experiments conducted on Qwen2.5-3B and Gemma3-4B-It. The work introduces the first efficient hybrid prompting template tailored for SLMs, significantly enhancing multi-hop reasoning performance under resource-constrained conditions. On a test set of 18,720 samples, the proposed approach achieves up to 83% and 84.5% relative accuracy improvements over standard RAG prompts for the two models, respectively, corresponding to an absolute accuracy gain of up to 6%. The paper also provides a reproducible guideline for effective prompt design in SLM-based multi-hop QA systems.