Score
Design and iterate prompt templates that include a small number of input–output demonstrations and explicit format constraints to steer large language models toward accurate, structured outputs; build demonstration-selection and formatting strategies, output parsers, and evaluation comparisons against zero-shot and supervised fine-tuning to analyze marginal performance gains.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.
Prompt template design for LLM applications remains largely empirical and lacks systematic, principled methodologies. Method: This paper introduces the first industrial-grade prompt template analysis framework: (1) constructing a high-quality dataset of templates from open-source LLM applications (e.g., Uber, Microsoft), curated via LLM-assisted parsing augmented with human verification; (2) establishing the first structured taxonomy of template components; and (3) conducting component-level statistical modeling and A/B-style instruction-following evaluations. Contribution/Results: We identify frequent co-occurrence patterns among template components and quantify their substantial impact on instruction-following performance—yielding up to a 23.6% accuracy gain. Furthermore, we distill reusable, robust design principles and optimization guidelines. This work provides both theoretical foundations and practical paradigms for prompt engineering, advancing systematic, data-driven template design in production LLM systems.
This work addresses a longstanding limitation in large language model (LLM) prompt engineering—namely, the exclusive focus on optimizing prompt *content* while neglecting *format* design. We propose a novel paradigm of joint content-and-format optimization. Methodologically, we introduce the first framework that treats prompt format as a learnable dimension, enabling content–format co-optimization through iterative refinement. Our approach integrates natural-language-based prompt mutation, dynamic search over a structured format space, multi-task joint evaluation, and model-agnostic black-box optimization. Extensive experiments across multiple open-source LLMs and diverse downstream tasks demonstrate that our method consistently outperforms content-only baselines, yielding average accuracy improvements of 2.1–5.7 percentage points. The implementation is publicly available.
To address privacy leakage risks in LLM applications within sensitive domains such as finance, this paper proposes an iterative hard prompt optimization method that operates without exposing task-specific context. The core innovation is a novel few-shot meta-prompting mechanism: leveraging the LLM’s intrinsic meta-reasoning capability over minimal examples, it autonomously generates and iteratively refines prompt templates—achieving performance gains without disclosing proprietary data. The method integrates self-prompt optimization, templated prompt engineering, and iterative propagation, strictly preserving syntactic structure and linguistic style consistency. Experiments across diverse contextual tasks demonstrate an average improvement of 103.87%, significantly enhancing grammatical stability and stylistic fidelity of prompts. This work establishes a new paradigm for compliant, privacy-preserving prompt engineering in high-regulation environments.
Prompt engineering heavily relies on manual expertise, while automated optimization methods suffer from a lack of labeled data. Method: This paper proposes a human-in-the-loop interactive prompt optimization framework that uniquely integrates real-time human judgment into the optimization loop. It combines active learning sampling, LLM-generated self-explanations, lightweight performance evaluation, and an interactive visualization interface—enabling domain experts (without programming skills) to iteratively refine prompts based on model explanations, sample-level feedback, and metric analysis. Contribution/Results: Experiments demonstrate significant improvements in prompt quality across diverse tasks. The framework enables non-technical users to efficiently construct high-performance, task-specific prompts. Furthermore, it uncovers critical intrinsic factors governing prompt optimization efficacy—namely semantic consistency, sample representativeness, and explanation credibility—thereby advancing both practical prompt engineering and foundational understanding of LLM behavior.
Non-expert users struggle to efficiently optimize LLM prompts due to limited domain knowledge and insufficient feedback mechanisms. To address this, we propose a beginner-oriented visual prompt engineering system featuring a novel tri-strategy collaborative optimization framework—integrating keyword perturbation, semantic paraphrasing, and optimal few-shot example recommendation. We design a multi-view synchronized interface, an interactive prompt editing environment, and a real-time evaluation mechanism grounded in both semantic similarity and task-specific accuracy. Experiments demonstrate that our system reduces user prompt iteration time by 37%, increases prompt diversity by 2.1×, and improves average accuracy by 11.4% across multiple NLP tasks—significantly outperforming existing prompt interfaces. This work lowers the cognitive barrier to prompt engineering and establishes a new paradigm for LLM interaction that is interpretable, iterative, and empirically evaluable for non-experts.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
This work addresses the limitations of large language models in practical deployment, where textual prompts often fail to enable efficient, stable, and inference-only customization. To overcome this, the paper proposes opening vector prompts as a standardized user interface, establishing a novel customization paradigm. Through vector prompt tuning, attention mechanism analysis, and security evaluation under black-box threat models, experiments demonstrate that vector prompts consistently improve performance with enhanced supervision signals, whereas textual prompts saturate early. Moreover, vector prompts induce globally dense attention patterns, revealing superior controllability and greater potential for model customization compared to conventional textual prompting.
Current large language model evaluation frameworks commonly rely on uniform, static prompt templates, neglecting model-specific prompt optimization and thereby introducing performance distortion and ranking bias. This work presents the first systematic investigation into the impact of prompt optimization on model evaluation and introduces a novel “optimize-then-evaluate” paradigm: prompts are individually optimized for each model prior to performance assessment. Through comprehensive experiments employing diverse prompt optimization techniques across established academic and industrial benchmarks, the study demonstrates that prompt optimization substantially alters model rankings. These findings underscore the critical role of customized prompting in achieving accurate evaluations and informed model selection, effectively bridging the gap between academic assessment protocols and real-world industrial practices.
Manual design of high-quality prompts is challenging in few-shot settings, while existing automated methods suffer from low efficiency and reliance on human-crafted demonstrations. Method: We propose ShapleyPrompt—the first framework to incorporate Monte Carlo Shapley values into in-context learning prompt construction. It quantifies the marginal contribution of each candidate example to model performance, enabling efficient, differentiable-budget-aware example selection and dynamic editing (addition, deletion, or retention). Our approach integrates aggressive sub-sampling, replay buffering, and LLM-based zero-/few-shot evaluation—eliminating the need for human-authored examples. Results: ShapleyPrompt significantly outperforms state-of-the-art automated prompting methods on text simplification, GSM8K, and multi-class classification. With increased computational budget, it establishes new SOTAs across all three tasks, demonstrating superior data efficiency and example quality.
This study addresses the challenges in evaluating large language model (LLM) applications—namely, high output stochasticity, multidimensionality, and sensitivity to prompt and model variations—which render traditional testing methods inadequate. The authors propose an evaluation-driven engineering workflow (Define-Test-Diagnose-Fix) and introduce the first hierarchical Minimum Viable Evaluation Suite (MVES) tailored for general-purpose LLMs, retrieval-augmented generation (RAG), and agent-based tool-use scenarios. The framework integrates automated checks, human scoring, and LLM-as-judge to establish a reproducible local evaluation system, validated on the Ollama platform using Llama 3 8B and Qwen 2.5 7B Instruct models. Experiments reveal that while generic prompt templates enhance instruction following, they degrade structured extraction accuracy from 100% to 90% and RAG compliance from 93.3% to 80%, underscoring the necessity of evaluation-driven iteration and advocating systematic assessment over heuristic prompt engineering.