few-shot prompting

Design and iterate prompt templates that include a small number of input–output demonstrations and explicit format constraints to steer large language models toward accurate, structured outputs; build demonstration-selection and formatting strategies, output parsers, and evaluation comparisons against zero-shot and supervised fine-tuning to analyze marginal performance gains.

few-shotprompting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.75
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

The Prompt Report: A Systematic Survey of Prompting Techniques

Jun 06, 2024
SS
Sander Schulhoff
🏛️ University of Maryland | Stanford | Microsoft | University of Massachusetts Amherst | Texas State University | Icahn School of Medicine | Mount Sinai Beth Israel | Princeton | Vanderbilt

The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.

Generative AILarge Language ModelsPrompt Engineering

Must-Read Papers

Most classic and influential ideas
View more

From Prompts to Templates: A Systematic Prompt Template Analysis for Real-world LLMapps

Apr 02, 2025
YM
Yuetian Mao
🏛️ Technical University of Munich

Prompt template design for LLM applications remains largely empirical and lacks systematic, principled methodologies. Method: This paper introduces the first industrial-grade prompt template analysis framework: (1) constructing a high-quality dataset of templates from open-source LLM applications (e.g., Uber, Microsoft), curated via LLM-assisted parsing augmented with human verification; (2) establishing the first structured taxonomy of template components; and (3) conducting component-level statistical modeling and A/B-style instruction-following evaluations. Contribution/Results: We identify frequent co-occurrence patterns among template components and quantify their substantial impact on instruction-following performance—yielding up to a 23.6% accuracy gain. Furthermore, we distill reusable, robust design principles and optimization guidelines. This work provides both theoretical foundations and practical paradigms for prompt engineering, advancing systematic, data-driven template design in production LLM systems.

Optimizing prompt design to improve LLM performanceReducing reliance on trial-and-error for template creationSystematic analysis of prompt templates in LLMapps

Beyond Prompt Content: Enhancing LLM Performance via Content-Format Integrated Prompt Optimization

Feb 06, 2025
YL
Yuanye Liu
🏛️ Fudan University | Microsoft Research Asia

This work addresses a longstanding limitation in large language model (LLM) prompt engineering—namely, the exclusive focus on optimizing prompt *content* while neglecting *format* design. We propose a novel paradigm of joint content-and-format optimization. Methodologically, we introduce the first framework that treats prompt format as a learnable dimension, enabling content–format co-optimization through iterative refinement. Our approach integrates natural-language-based prompt mutation, dynamic search over a structured format space, multi-task joint evaluation, and model-agnostic black-box optimization. Extensive experiments across multiple open-source LLMs and diverse downstream tasks demonstrate that our method consistently outperforms content-only baselines, yielding average accuracy improvements of 2.1–5.7 percentage points. The implementation is publicly available.

Enhances Large Language Models performanceIntroduces integrated prompt optimization methodologyOptimizes both prompt content and formatting

Optimising Hard Prompts with Few-Shot Meta-Prompting

Jul 09, 2024
SR
Sayash Raaj Hiraou
🏛️ Fidelity Investments

To address privacy leakage risks in LLM applications within sensitive domains such as finance, this paper proposes an iterative hard prompt optimization method that operates without exposing task-specific context. The core innovation is a novel few-shot meta-prompting mechanism: leveraging the LLM’s intrinsic meta-reasoning capability over minimal examples, it autonomously generates and iteratively refines prompt templates—achieving performance gains without disclosing proprietary data. The method integrates self-prompt optimization, templated prompt engineering, and iterative propagation, strictly preserving syntactic structure and linguistic style consistency. Experiments across diverse contextual tasks demonstrate an average improvement of 103.87%, significantly enhancing grammatical stability and stylistic fidelity of prompts. This work establishes a new paradigm for compliant, privacy-preserving prompt engineering in high-regulation environments.

Maintaining regulatory compliance while improving financial task performanceOptimizing prompts without exposing proprietary financial data to LLMsPreserving privacy when adapting LLMs for sensitive financial applications

Prompt engineering heavily relies on manual expertise, while automated optimization methods suffer from a lack of labeled data. Method: This paper proposes a human-in-the-loop interactive prompt optimization framework that uniquely integrates real-time human judgment into the optimization loop. It combines active learning sampling, LLM-generated self-explanations, lightweight performance evaluation, and an interactive visualization interface—enabling domain experts (without programming skills) to iteratively refine prompts based on model explanations, sample-level feedback, and metric analysis. Contribution/Results: Experiments demonstrate significant improvements in prompt quality across diverse tasks. The framework enables non-technical users to efficiently construct high-performance, task-specific prompts. Furthermore, it uncovers critical intrinsic factors governing prompt optimization efficacy—namely semantic consistency, sample representativeness, and explanation credibility—thereby advancing both practical prompt engineering and foundational understanding of LLM behavior.

Assisting non-technical users in generating task-specific promptsBridging manual and automatic prompt optimization for LLMsEnhancing human engagement in interactive prompt refinement

Non-expert users struggle to efficiently optimize LLM prompts due to limited domain knowledge and insufficient feedback mechanisms. To address this, we propose a beginner-oriented visual prompt engineering system featuring a novel tri-strategy collaborative optimization framework—integrating keyword perturbation, semantic paraphrasing, and optimal few-shot example recommendation. We design a multi-view synchronized interface, an interactive prompt editing environment, and a real-time evaluation mechanism grounded in both semantic similarity and task-specific accuracy. Experiments demonstrate that our system reduces user prompt iteration time by 37%, increases prompt diversity by 2.1×, and improves average accuracy by 11.4% across multiple NLP tasks—significantly outperforming existing prompt interfaces. This work lowers the cognitive barrier to prompt engineering and establishes a new paradigm for LLM interaction that is interpretable, iterative, and empirically evaluable for non-experts.

Explore and refine prompts for LLMsSimplify prompt iteration for non-expertsTest prompt performance interactively

Latest Papers

What's happening recently
View more

This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.

documentationevaluationprompt engineering

This work addresses the limitations of large language models in practical deployment, where textual prompts often fail to enable efficient, stable, and inference-only customization. To overcome this, the paper proposes opening vector prompts as a standardized user interface, establishing a novel customization paradigm. Through vector prompt tuning, attention mechanism analysis, and security evaluation under black-box threat models, experiments demonstrate that vector prompts consistently improve performance with enhanced supervision signals, whereas textual prompts saturate early. Moreover, vector prompts induce globally dense attention patterns, revealing superior controllability and greater potential for model customization compared to conventional textual prompting.

customizationinference-only customizationlarge language models

Current large language model evaluation frameworks commonly rely on uniform, static prompt templates, neglecting model-specific prompt optimization and thereby introducing performance distortion and ranking bias. This work presents the first systematic investigation into the impact of prompt optimization on model evaluation and introduces a novel “optimize-then-evaluate” paradigm: prompts are individually optimized for each model prior to performance assessment. Through comprehensive experiments employing diverse prompt optimization techniques across established academic and industrial benchmarks, the study demonstrates that prompt optimization substantially alters model rankings. These findings underscore the critical role of customized prompting in achieving accurate evaluations and informed model selection, effectively bridging the gap between academic assessment protocols and real-world industrial practices.

BenchmarkingLarge Language ModelModel Evaluation

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data

Dec 11, 2025
PB
Pawel Batorski
🏛️ Heinrich Heine Universität Düsseldorf

Manual design of high-quality prompts is challenging in few-shot settings, while existing automated methods suffer from low efficiency and reliance on human-crafted demonstrations. Method: We propose ShapleyPrompt—the first framework to incorporate Monte Carlo Shapley values into in-context learning prompt construction. It quantifies the marginal contribution of each candidate example to model performance, enabling efficient, differentiable-budget-aware example selection and dynamic editing (addition, deletion, or retention). Our approach integrates aggressive sub-sampling, replay buffering, and LLM-based zero-/few-shot evaluation—eliminating the need for human-authored examples. Results: ShapleyPrompt significantly outperforms state-of-the-art automated prompting methods on text simplification, GSM8K, and multi-class classification. With increased computational budget, it establishes new SOTAs across all three tasks, demonstrating superior data efficiency and example quality.

Addresses LLMs' sensitivity to prompt design by automating construction.Generates few-shot examples to augment human instructions efficiently.Optimizes example selection using Monte Carlo Shapley estimation.

This study addresses the challenges in evaluating large language model (LLM) applications—namely, high output stochasticity, multidimensionality, and sensitivity to prompt and model variations—which render traditional testing methods inadequate. The authors propose an evaluation-driven engineering workflow (Define-Test-Diagnose-Fix) and introduce the first hierarchical Minimum Viable Evaluation Suite (MVES) tailored for general-purpose LLMs, retrieval-augmented generation (RAG), and agent-based tool-use scenarios. The framework integrates automated checks, human scoring, and LLM-as-judge to establish a reproducible local evaluation system, validated on the Ollama platform using Llama 3 8B and Qwen 2.5 7B Instruct models. Experiments reveal that while generic prompt templates enhance instruction following, they degrade structured extraction accuracy from 100% to 90% and RAG compliance from 93.3% to 80%, underscoring the necessity of evaluation-driven iteration and advocating systematic assessment over heuristic prompt engineering.

Evaluation-Driven DevelopmentLarge Language ModelsLLM Testing

Hot Scholars

SM

Subhankar Maity

Postdoctoral Researcher, ECE Paris
Natural Language ProcessingAI in EducationLarge Language ModelsGenerative AI
AD

Aniket Deroy

PostDoc at IIT Delhi
Artificial IntelligenceDeep LearningLegal Analytics
VG

Vivek Gupta

Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval
DR

Dan Roth

Professor of Computer Science, University of Pennsylvania
Natural Language ProcessingMachine LearningKnowledge Representation and ReasoningArtificial Intelligence
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing