Score
Designs, builds, and analyzes prompting artifacts and methods for language models, including structured templates, role- and task-conditioned prompts, in‑context example selection, zero‑shot instructions, and strategies for controlled or conditioned generation. Develops prompt representations and search/optimization procedures (e.g., prompt embeddings, variation experiments, reasoning‑guided search), constructs prompt‑based classifiers and evaluation/testing protocols, and systematically measures how prompt choices elicit reasoning or other desired behaviors.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.
Large language models (LLMs) often exhibit low alignment with human expert judgments when identifying theory-driven psychological constructs in textual data. Method: This study proposes an empirically grounded prompting framework that integrates codebook-guided instruction with automated prompt generation, systematically evaluating how construct definitions, task formulations, and exemplar selection affect prompting efficacy. We compare five prompting strategies—codebook-guided selection, automated prompt engineering, role prompting, chain-of-thought, and explanatory prompting—under zero-shot and few-shot settings. Contribution/Results: The few-shot “codebook-guided + automated engineering” strategy achieves statistically significant improvements in human–model agreement across multiple psychological constructs and mainstream LLMs, yielding an average increase of 0.23 in Krippendorff’s α. This work establishes a reproducible, theory-embedded prompting paradigm for construct-driven psychological text analysis.
This paper presents a systematic survey of prompt tuning—a parameter-efficient paradigm for adapting pre-trained language models—focusing specifically on the setting where the backbone model is frozen and only continuous, prefix-based prompt embeddings are optimized. Addressing key challenges including computational inefficiency and training instability, the work introduces the first unified taxonomy encompassing encoder-based, low-rank decomposition, and mixture-of-experts prompt tuning methods, rigorously distinguishing direct prompt learning from transferable prompt learning. Through methodological analysis and visualized performance comparisons across diverse benchmarks, it characterizes fundamental trade-offs among parameter count, optimization convergence, and generalization capability. The study provides both theoretical insights and practical guidelines for enhancing training robustness and extending prompt tuning to multi-task and low-resource scenarios.
Existing surveys predominantly focus on prompt engineering but lack systematic classification and evaluation of prompt optimization strategies. To address this gap, we propose the first unified taxonomy covering 11 distinct optimization paradigms, structured along four dimensions: operational paradigm, target task, large language model (LLM) architecture, and benchmark dataset. We further introduce a standardized evaluation framework enabling consistent, cross-task and cross-model comparative experiments. Our work integrates LLM-based prompt optimization methods, NLP task-specific adaptation techniques, and multi-dataset benchmarking to establish the first open-source knowledge base for prompt optimization. This study fills a critical void in systematic survey literature, providing both theoretical foundations and practical infrastructure for advancing research on prompt optimization mechanisms, designing novel optimizers, and ensuring reproducible, rigorous evaluation.
Large language models exhibit insufficient robustness to out-of-distribution inputs, and while machine-generated optimized prompts effectively steer model outputs, their compositional principles and internal mechanistic pathways remain poorly understood. This work systematically investigates the structure of optimized prompts and their in-model interpretation mechanisms via three complementary approaches: neural activation analysis, token frequency statistics, and cross-model representation trajectory tracking. We make two key discoveries: first, optimized prompts consistently rely heavily on punctuation marks and low-frequency nouns, and exhibit a shared, invariant representation evolution path across diverse instruction-tuned models; second, we identify a sparse, generalizable subset of neural activations that robustly discriminates optimized prompts from natural language across models and tasks. These findings establish an interpretable, transferable mechanistic foundation for enhancing controllability and robustness in large language models.
Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.
This study investigates how prompting affects the quality of internal representations in large language models (LLMs) during zero-shot classification, and how this relates to prompt-task relevance. Methodologically, we construct diverse prompt templates and employ representation probing to systematically assess their impact on the separability and semantic structure of hidden-layer embeddings. Our results reveal that prompting substantially reshapes model representations; however, representation quality does not monotonically improve with increasing prompt-task relevance—in fact, highly relevant prompts sometimes degrade performance. This finding challenges the implicit assumption that “more relevant prompts yield better representations,” exposing the non-intuitive nature of prompt mechanisms in in-context learning. The work provides novel theoretical insights and empirical evidence for understanding zero-shot generalization in LLMs, highlighting the complex, non-linear relationship between prompt design, internal representation geometry, and downstream task performance.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
This study investigates large language models’ (LLMs) capacity for cultural understanding and creative adaptation within poetic contexts. To address limitations in existing prompt engineering for literary tasks, we propose *Poetry Prompt Patterns*—a novel prompting framework that structures poetic expression (e.g., metaphor, meter, imagery directives) to elicit stylistic emulation, canonical work evaluation, and audience-tailored rewriting. Through controlled generative experiments and qualitative literary analysis, we systematically assess LLMs’ performance across literary interpretation, cultural localization, and rhetorical strategy. Results reveal systematic biases in poetic cognition—including stylistic flattening and cultural stereotyping—and expose critical boundaries in rhetorical generation, particularly concerning non-literal meaning and historical contextualization. Our key contribution lies in pioneering the use of poetic form itself as a metalinguistic diagnostic tool for evaluating AI literary intelligence, thereby establishing an interdisciplinary paradigm bridging literary criticism and prompt engineering.
This paper challenges the theoretical foundations and explanatory power of “text gradient”-based automated prompt optimization methods, which metaphorically equate discrete text updates with continuous, differentiable gradient descent. Method: Through systematic LLM prompt fine-tuning experiments, multi-task comparative analysis, ablation studies, and behavioral attribution, we rigorously examine whether these methods operate as genuine gradient-based optimizers. Contribution/Results: We demonstrate that performance gains are not attributable to gradient update logic; instead, “text gradients” function merely as empirical heuristics without theoretical grounding in differentiable optimization. First, we formally establish their non-gradient nature. Second, we propose a novel conceptual framework for prompt optimization explicitly tailored to discrete text spaces. Third, we advocate shifting prompt engineering from analogical transfer (e.g., borrowing optimization metaphors from continuous domains) toward intrinsic, ontology-aware modeling. These findings call for a fundamental methodological rethinking of prompt optimization.