prompt engineering

Designs, builds, and analyzes prompting artifacts and methods for language models, including structured templates, role- and task-conditioned prompts, in‑context example selection, zero‑shot instructions, and strategies for controlled or conditioned generation. Develops prompt representations and search/optimization procedures (e.g., prompt embeddings, variation experiments, reasoning‑guided search), constructs prompt‑based classifiers and evaluation/testing protocols, and systematically measures how prompt choices elicit reasoning or other desired behaviors.

promptengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.88
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

The Prompt Report: A Systematic Survey of Prompting Techniques

Jun 06, 2024
SS
Sander Schulhoff
🏛️ University of Maryland | Stanford | Microsoft | University of Massachusetts Amherst | Texas State University | Icahn School of Medicine | Mount Sinai Beth Israel | Princeton | Vanderbilt

The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.

Generative AILarge Language ModelsPrompt Engineering

Must-Read Papers

Most classic and influential ideas
View more

Large language models (LLMs) often exhibit low alignment with human expert judgments when identifying theory-driven psychological constructs in textual data. Method: This study proposes an empirically grounded prompting framework that integrates codebook-guided instruction with automated prompt generation, systematically evaluating how construct definitions, task formulations, and exemplar selection affect prompting efficacy. We compare five prompting strategies—codebook-guided selection, automated prompt engineering, role prompting, chain-of-thought, and explanatory prompting—under zero-shot and few-shot settings. Contribution/Results: The few-shot “codebook-guided + automated engineering” strategy achieves statistically significant improvements in human–model agreement across multiple psychological constructs and mainstream LLMs, yielding an average increase of 0.23 in Krippendorff’s α. This work establishes a reproducible, theory-embedded prompting paradigm for construct-driven psychological text analysis.

Addressing performance loss from poorly worded prompts in classification tasksAssessing prompt engineering strategies to align with expert judgmentsOptimizing LLM prompts for psychological construct identification

A Survey on Prompt Tuning

Jul 08, 2025
ZL
Zongqian Li
🏛️ University of Cambridge

This paper presents a systematic survey of prompt tuning—a parameter-efficient paradigm for adapting pre-trained language models—focusing specifically on the setting where the backbone model is frozen and only continuous, prefix-based prompt embeddings are optimized. Addressing key challenges including computational inefficiency and training instability, the work introduces the first unified taxonomy encompassing encoder-based, low-rank decomposition, and mixture-of-experts prompt tuning methods, rigorously distinguishing direct prompt learning from transferable prompt learning. Through methodological analysis and visualized performance comparisons across diverse benchmarks, it characterizes fundamental trade-offs among parameter count, optimization convergence, and generalization capability. The study provides both theoretical insights and practical guidelines for enhancing training robustness and extending prompt tuning to multi-task and low-resource scenarios.

Classifies approaches into direct prompt learning and transfer learningIdentifies challenges in computational efficiency and training stabilitySurvey reviews prompt tuning for adapting language models efficiently

The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future

Jun 21, 2025
SS
Summra Saleem
🏛️ RPTU Rheinland-Pfälzische Technische Universität | German Research Centre for Artificial Intelligence | University of Management and Technology

Existing surveys predominantly focus on prompt engineering but lack systematic classification and evaluation of prompt optimization strategies. To address this gap, we propose the first unified taxonomy covering 11 distinct optimization paradigms, structured along four dimensions: operational paradigm, target task, large language model (LLM) architecture, and benchmark dataset. We further introduce a standardized evaluation framework enabling consistent, cross-task and cross-model comparative experiments. Our work integrates LLM-based prompt optimization methods, NLP task-specific adaptation techniques, and multi-dataset benchmarking to establish the first open-source knowledge base for prompt optimization. This study fills a critical void in systematic survey literature, providing both theoretical foundations and practical infrastructure for advancing research on prompt optimization mechanisms, designing novel optimizers, and ensuring reproducible, rigorous evaluation.

Analyzes prompt optimization strategies in NLP tasksClassifies 11 distinct prompt optimization methodologiesEvaluates LLMs and datasets for performance benchmarking

Demystifying optimized prompts in language models

May 04, 2025
RM
Rimon Melamed
🏛️ The George Washington University | LMI Consulting

Large language models exhibit insufficient robustness to out-of-distribution inputs, and while machine-generated optimized prompts effectively steer model outputs, their compositional principles and internal mechanistic pathways remain poorly understood. This work systematically investigates the structure of optimized prompts and their in-model interpretation mechanisms via three complementary approaches: neural activation analysis, token frequency statistics, and cross-model representation trajectory tracking. We make two key discoveries: first, optimized prompts consistently rely heavily on punctuation marks and low-frequency nouns, and exhibit a shared, invariant representation evolution path across diverse instruction-tuned models; second, we identify a sparse, generalizable subset of neural activations that robustly discriminates optimized prompts from natural language across models and tasks. These findings establish an interpretable, transferable mechanistic foundation for enhancing controllability and robustness in large language models.

Analyzing activation patterns for optimized vs natural promptsExploring mechanisms of LM parsing for optimized promptsUnderstanding composition of optimized prompts in LMs

Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.

LLM classificationperformance varianceprompt engineering

Latest Papers

What's happening recently
View more

Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings

Oct 22, 2025
CG
Cesar Gonzalez-Gutierrez
🏛️ Polytechnic University of Catalonia | Bocconi University

This study investigates how prompting affects the quality of internal representations in large language models (LLMs) during zero-shot classification, and how this relates to prompt-task relevance. Methodologically, we construct diverse prompt templates and employ representation probing to systematically assess their impact on the separability and semantic structure of hidden-layer embeddings. Our results reveal that prompting substantially reshapes model representations; however, representation quality does not monotonically improve with increasing prompt-task relevance—in fact, highly relevant prompts sometimes degrade performance. This finding challenges the implicit assumption that “more relevant prompts yield better representations,” exposing the non-intuitive nature of prompt mechanisms in in-context learning. The work provides novel theoretical insights and empirical evidence for understanding zero-shot generalization in LLMs, highlighting the complex, non-linear relationship between prompt design, internal representation geometry, and downstream task performance.

Analyzes factors behind unexpected prompt-representation quality relationshipsExamines if relevant prompts consistently improve representation qualityInvestigates how prompting affects language model embedding representations

This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.

documentationevaluationprompt engineering

Decoding the Black Box: Discerning AI Rhetorics About and Through Poetic Prompting

Dec 04, 2025
PD
P. D. Edgar
🏛️ University of Central Florida

This study investigates large language models’ (LLMs) capacity for cultural understanding and creative adaptation within poetic contexts. To address limitations in existing prompt engineering for literary tasks, we propose *Poetry Prompt Patterns*—a novel prompting framework that structures poetic expression (e.g., metaphor, meter, imagery directives) to elicit stylistic emulation, canonical work evaluation, and audience-tailored rewriting. Through controlled generative experiments and qualitative literary analysis, we systematically assess LLMs’ performance across literary interpretation, cultural localization, and rhetorical strategy. Results reveal systematic biases in poetic cognition—including stylistic flattening and cultural stereotyping—and expose critical boundaries in rhetorical generation, particularly concerning non-literal meaning and historical contextualization. Our key contribution lies in pioneering the use of poetic form itself as a metalinguistic diagnostic tool for evaluating AI literary intelligence, thereby establishing an interdisciplinary paradigm bridging literary criticism and prompt engineering.

Assessing AI models' biases through poetic prompt patternsExploring LLMs' creative boundaries via poetry generation tasksTesting AI adaptation of original works for different audiences

Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization

Dec 15, 2025
DM
Daniel Melcer
🏛️ Northeastern University | AWS AI Labs

This paper challenges the theoretical foundations and explanatory power of “text gradient”-based automated prompt optimization methods, which metaphorically equate discrete text updates with continuous, differentiable gradient descent. Method: Through systematic LLM prompt fine-tuning experiments, multi-task comparative analysis, ablation studies, and behavioral attribution, we rigorously examine whether these methods operate as genuine gradient-based optimizers. Contribution/Results: We demonstrate that performance gains are not attributable to gradient update logic; instead, “text gradients” function merely as empirical heuristics without theoretical grounding in differentiable optimization. First, we formally establish their non-gradient nature. Second, we propose a novel conceptual framework for prompt optimization explicitly tailored to discrete text spaces. Third, we advocate shifting prompt engineering from analogical transfer (e.g., borrowing optimization metaphors from continuous domains) toward intrinsic, ontology-aware modeling. These findings call for a fundamental methodological rethinking of prompt optimization.

Examines if gradient analogy accurately explains optimization behaviorInforms selection and development of prompt optimization strategiesInvestigates textual gradient methods for automatic prompt optimization

Hot Scholars

HW

Haohan Wang

School of Information Sciences, University of Illinois Urbana-Champaign
Computational BiologyAgentic AIAI4ScienceAI security
DH

Dirk Hovy

Bocconi University
Natural Language ProcessingMachine LearningComputational SociolinguisticsComputational Social Science
AP

Aske Plaat

Leiden University
Large Reasoning ModelsMusicReinforcement Learning
JD

Jinho D. Choi

Associate Professor, Emory University
Natural Language ProcessingComputational LinguisticsConversational AI
AS

Adish Singla

Tenured Faculty, Max Planck Institute for Software Systems (MPI-SWS), Germany
Machine TeachingAI for EducationProgramming EducationReinforcement Learning