domain-conditioned prompting

Designs and evaluates prompting strategies and prompt templates that condition large language model inputs on dataset- or domain-level context (e.g., high-level descriptions, domain metadata, or contextual few-shot examples) to steer model outputs. These methods are used to build or analyze techniques that mitigate domain mismatch, improve cross-domain similarity judgments, and enable domain-aware comparisons of biases.

domain-conditionedprompting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

The Prompt Report: A Systematic Survey of Prompting Techniques

Jun 06, 2024
SS
Sander Schulhoff
🏛️ University of Maryland | Stanford | Microsoft | University of Massachusetts Amherst | Texas State University | Icahn School of Medicine | Mount Sinai Beth Israel | Princeton | Vanderbilt

The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.

Generative AILarge Language ModelsPrompt Engineering

Must-Read Papers

Most classic and influential ideas
View more

Effects of Prompt Length on Domain-specific Tasks for Large Language Models

Feb 20, 2025
QL
Qibang Liu
🏛️ Georgia Institute of Technology | Nanjing University of Finance and Economics | Boston University

Large language models (LLMs) face challenges in domain-specific tasks—such as financial sentiment analysis and monetary policy interpretation—due to insufficient domain knowledge activation and limited reasoning accuracy; the relationship between prompt length and model performance remains poorly characterized. This work systematically quantifies the marginal effect of prompt length on professional-domain tasks for the first time, conducting controlled experiments across six financial and legal benchmark datasets using open-source models (e.g., LLaMA, Qwen), augmented by attention visualization and token-level gradient attribution analysis. Results reveal a significant nonlinear relationship between prompt length and performance: excessively long prompts degrade domain-specific reasoning accuracy—up to a 11.7% drop in peak accuracy. We propose a length-adaptive prompt truncation strategy that, when applied at optimal lengths, yields an average F1-score improvement of 3.2%. Our findings provide an interpretable, reusable methodology for prompt engineering in professional-domain applications.

Domain-specific task performanceImpact of prompt length on LLMsPrompt engineering for specialized knowledge

Optimising Hard Prompts with Few-Shot Meta-Prompting

Jul 09, 2024
SR
Sayash Raaj Hiraou
🏛️ Fidelity Investments

To address privacy leakage risks in LLM applications within sensitive domains such as finance, this paper proposes an iterative hard prompt optimization method that operates without exposing task-specific context. The core innovation is a novel few-shot meta-prompting mechanism: leveraging the LLM’s intrinsic meta-reasoning capability over minimal examples, it autonomously generates and iteratively refines prompt templates—achieving performance gains without disclosing proprietary data. The method integrates self-prompt optimization, templated prompt engineering, and iterative propagation, strictly preserving syntactic structure and linguistic style consistency. Experiments across diverse contextual tasks demonstrate an average improvement of 103.87%, significantly enhancing grammatical stability and stylistic fidelity of prompts. This work establishes a new paradigm for compliant, privacy-preserving prompt engineering in high-regulation environments.

Maintaining regulatory compliance while improving financial task performanceOptimizing prompts without exposing proprietary financial data to LLMsPreserving privacy when adapting LLMs for sensitive financial applications

The Prompt is Mightier than the Example

May 24, 2025
SX
Shengzhe Xu
🏛️ Virginia Tech | Stevens Institute of Technology

Synthetic tabular data generation using large language models (LLMs) suffers from poor scalability due to heavy reliance on in-context learning (ICL) examples. Method: This paper proposes Knowledge-Guided Prompting (KGP), a paradigm that explicitly injects structured domain knowledge into prompts to replace a subset of ICL examples. KGP integrates reasoning-aware prompt design with optimized knowledge encoding. Contribution/Results: Through systematic ablation studies and a multi-dimensional evaluation framework, we establish, for the first time, a quantifiable substitution scaling law between knowledge volume and ICL example count. KGP significantly reduces ICL dependency—by up to 80%—while preserving statistical fidelity and improving downstream task performance. This work provides both a novel paradigm and theoretical foundation for high-fidelity, low-overhead LLM-driven tabular data synthesis.

Balancing synthetic data quality and example count trade-offsExploring prompt optimization with domain knowledge injectionReducing reliance on in-context learning examples for LLMs

Exploring Small Language Models with Prompt-Learning Paradigm for Efficient Domain-Specific Text Classification

Sep 26, 2023
HL
Hengyu Luo
🏛️ University of Helsinki | Ingka Group | IKEA

Intent recognition in retail customer service is hindered by severe scarcity of labeled data. Method: This paper investigates synergistic optimization of small language models (SLMs, <1B parameters) with prompt learning. We systematically evaluate SLMs for few-shot and zero-shot text classification—first such study—and propose active few-shot sampling and multi-prompt ensemble strategies. We further identify prompt engineering as the decisive factor governing SLM zero-shot performance. Results: T5-base achieves 75% accuracy using only 15% of the labeled data. After prompt optimization, FLAN-T5-large’s zero-shot accuracy improves from <18% to over 31%, substantially narrowing the gap with GPT-3.5-turbo (55.16%). Our approach establishes an efficient, lightweight paradigm for intent recognition in low-resource domains.

Addresses data scarcity in customer intent recognitionEnhances small models' performance with minimal dataReduces dependency on extensive labeled datasets

On Meta-Prompting

Dec 11, 2023
AD
Adrian de Wynter
🏛️ The University of York | Microsoft

Large language models (LLMs) lack gradient-based parameter updates and rely on in-context learning (ICL) and meta-prompting—yet no rigorous theoretical framework exists to formalize their semantics or behavior. Method: This work introduces the first unified formal framework grounded in category theory, rigorously modeling the semantic structure and behavioral properties of ICL and meta-prompting via categorical constructions, formal semantic analysis, and empirical validation. Contribution/Results: We formally establish the task-agnostic nature of meta-prompting and prove equivalence among mainstream meta-prompting methods. Experiments demonstrate that meta-prompting consistently outperforms standard prompting, yielding significant improvements in output controllability and cross-task generalization. This study provides the first provably sound theoretical foundation for prompt-based adaptation in LLMs without parameter updates.

Comparing effectiveness of meta-prompting vs basic promptingFormalizing in-context learning and task agnosticityTheoretical framework for LLM behavior in meta-prompting

Latest Papers

What's happening recently
View more

Existing LM evaluation frameworks (e.g., HELM) rely on fixed prompts, suffering from poor generalizability and consequently underestimating model performance and yielding inconsistent cross-model rankings. To address this, we propose DSPy+HELM—a novel integrated framework that systematically incorporates structured prompting strategies (including chain-of-thought, self-consistency, least-to-most, and program-of-thought) into standardized evaluation. We conduct reproducible, large-scale assessments of state-of-the-art LMs across seven diverse benchmarks—four general-purpose and three medical—using declarative prompt optimization to explicitly elicit and enhance model reasoning capabilities. Our approach significantly improves evaluation robustness: average accuracy increases by 4%, result variance decreases by 2%, and ranking reversals occur in 3/7 leaderboards, yielding more accurate estimates of true model capability ceilings. All prompt optimization pipelines and integration tools are open-sourced to strengthen decision utility and experimental reproducibility.

Fixed prompts in benchmarks underestimate language model performanceScalable prompting methods reduce sensitivity to prompt design variationsStructured prompting enables more accurate performance ceiling estimation

Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.

LLM classificationperformance varianceprompt engineering

This work addresses the limited generalization of prompt-based large language model (LLM) classifiers in data-scarce scenarios, where insufficient fine-tuning often hinders performance. The authors propose a multi-task prompt fine-tuning approach that designs task-specific prompts while integrating general instruction tuning, substantially improving classification accuracy on unseen domains and novel prompts. Notably, they find that supervised classification training without explicit reasoning capabilities can effectively generalize to reasoning-intensive tasks such as summarization. To mitigate performance degradation caused by prompt variations, a hybrid training strategy is introduced. Experimental results demonstrate strong performance on related unseen tasks, highlighting the potential of classification-oriented training for building versatile, general-purpose monitoring systems.

cross-domain generalizationfine-tuninginstruction following

This work addresses the limited performance of large language models (LLMs) in high-dimensional software engineering optimization tasks, where they often fail to surpass Bayesian optimization. For the first time, it systematically compares human- and AI-generated domain knowledge injection strategies and introduces four novel architectures: Human-feedback-informed Domain Knowledge Prompting (H-DKP), Adaptive Multi-stage Prompting (AMP), Dimension-aware Progressive Refinement (DAPR), and a hybrid approach combining statistical scouting with RAG-enhanced knowledge integration (HKMA). By leveraging a multi-stage, dimension-aware, and hybrid knowledge fusion framework, the proposed methods effectively incorporate structured domain knowledge to significantly enhance LLMs’ ability to generate high-quality initial solutions. Evaluated on the MOOT high-dimensional benchmark, the approaches markedly reduce the Chebyshev distance to the optimal solution and, according to Scott-Knott clustering, outperform existing LLM warm-start baselines.

Domain KnowledgeHigh-Dimensional OptimizationLarge Language Models

This study investigates whether large language models amplify biases present in user prompts and remain susceptible to prompt framing even on factual questions. By employing a controlled experimental design, the authors construct 160 prompts spanning ten topics to systematically disentangle the effects of implicit prompt framing from explicit manipulation on model outputs. The evaluation across six prominent large language models reveals a consistent tendency for models to align their responses with the framing of the prompt, often prioritizing user suggestions over factual consistency—even when objective facts are unambiguous. This work provides the first empirical evidence of the vulnerability of large language models to bias induction in factual domains, highlighting a critical limitation in their reliability despite advances in scale and training.

bias reinforcementfactual consistencylarge language models

Hot Scholars

XW

Xixin Wu

The Chinese University of Hong Kong
SF

Stefan Feuerriegel

Professor, LMU Munich
AI in ManagementBusiness AnalyticsComputational Social ScienceAI for Good
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning
DZ

Dongmei Zhang

Microsoft Research
Software EngineeringMachine LearningInformation Visualization
BJ

Bowen Jiang

University of Pennsylvania, Microsoft Corporation
Artificial IntelligencePost-trainingPersonalizationMultimodality