Score
Designs and evaluates prompting strategies and prompt templates that condition large language model inputs on dataset- or domain-level context (e.g., high-level descriptions, domain metadata, or contextual few-shot examples) to steer model outputs. These methods are used to build or analyze techniques that mitigate domain mismatch, improve cross-domain similarity judgments, and enable domain-aware comparisons of biases.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
The field of prompt engineering lacks a unified taxonomic framework and standardized terminology, resulting in fragmented technical understanding and insufficient practical guidance. Method: We conduct a systematic literature review, bibliometric analysis, and ontology modeling to construct the first cross-modal taxonomy encompassing 58 large language models and 40 multimodal prompting techniques; define 33 core terms; and perform the first comprehensive meta-analysis focused on natural language prefix prompting. Contribution/Results: Our work delivers the most extensive prompt technique classification system to date (98 categories), a standardized lexicon, and an actionable engineering guideline tailored for state-of-the-art models. It systematically addresses critical gaps in terminological inconsistency and ontological absence, establishing a foundational benchmark for the field.
Large language models (LLMs) face challenges in domain-specific tasks—such as financial sentiment analysis and monetary policy interpretation—due to insufficient domain knowledge activation and limited reasoning accuracy; the relationship between prompt length and model performance remains poorly characterized. This work systematically quantifies the marginal effect of prompt length on professional-domain tasks for the first time, conducting controlled experiments across six financial and legal benchmark datasets using open-source models (e.g., LLaMA, Qwen), augmented by attention visualization and token-level gradient attribution analysis. Results reveal a significant nonlinear relationship between prompt length and performance: excessively long prompts degrade domain-specific reasoning accuracy—up to a 11.7% drop in peak accuracy. We propose a length-adaptive prompt truncation strategy that, when applied at optimal lengths, yields an average F1-score improvement of 3.2%. Our findings provide an interpretable, reusable methodology for prompt engineering in professional-domain applications.
To address privacy leakage risks in LLM applications within sensitive domains such as finance, this paper proposes an iterative hard prompt optimization method that operates without exposing task-specific context. The core innovation is a novel few-shot meta-prompting mechanism: leveraging the LLM’s intrinsic meta-reasoning capability over minimal examples, it autonomously generates and iteratively refines prompt templates—achieving performance gains without disclosing proprietary data. The method integrates self-prompt optimization, templated prompt engineering, and iterative propagation, strictly preserving syntactic structure and linguistic style consistency. Experiments across diverse contextual tasks demonstrate an average improvement of 103.87%, significantly enhancing grammatical stability and stylistic fidelity of prompts. This work establishes a new paradigm for compliant, privacy-preserving prompt engineering in high-regulation environments.
Synthetic tabular data generation using large language models (LLMs) suffers from poor scalability due to heavy reliance on in-context learning (ICL) examples. Method: This paper proposes Knowledge-Guided Prompting (KGP), a paradigm that explicitly injects structured domain knowledge into prompts to replace a subset of ICL examples. KGP integrates reasoning-aware prompt design with optimized knowledge encoding. Contribution/Results: Through systematic ablation studies and a multi-dimensional evaluation framework, we establish, for the first time, a quantifiable substitution scaling law between knowledge volume and ICL example count. KGP significantly reduces ICL dependency—by up to 80%—while preserving statistical fidelity and improving downstream task performance. This work provides both a novel paradigm and theoretical foundation for high-fidelity, low-overhead LLM-driven tabular data synthesis.
Intent recognition in retail customer service is hindered by severe scarcity of labeled data. Method: This paper investigates synergistic optimization of small language models (SLMs, <1B parameters) with prompt learning. We systematically evaluate SLMs for few-shot and zero-shot text classification—first such study—and propose active few-shot sampling and multi-prompt ensemble strategies. We further identify prompt engineering as the decisive factor governing SLM zero-shot performance. Results: T5-base achieves 75% accuracy using only 15% of the labeled data. After prompt optimization, FLAN-T5-large’s zero-shot accuracy improves from <18% to over 31%, substantially narrowing the gap with GPT-3.5-turbo (55.16%). Our approach establishes an efficient, lightweight paradigm for intent recognition in low-resource domains.
Large language models (LLMs) lack gradient-based parameter updates and rely on in-context learning (ICL) and meta-prompting—yet no rigorous theoretical framework exists to formalize their semantics or behavior. Method: This work introduces the first unified formal framework grounded in category theory, rigorously modeling the semantic structure and behavioral properties of ICL and meta-prompting via categorical constructions, formal semantic analysis, and empirical validation. Contribution/Results: We formally establish the task-agnostic nature of meta-prompting and prove equivalence among mainstream meta-prompting methods. Experiments demonstrate that meta-prompting consistently outperforms standard prompting, yielding significant improvements in output controllability and cross-task generalization. This study provides the first provably sound theoretical foundation for prompt-based adaptation in LLMs without parameter updates.
Existing LM evaluation frameworks (e.g., HELM) rely on fixed prompts, suffering from poor generalizability and consequently underestimating model performance and yielding inconsistent cross-model rankings. To address this, we propose DSPy+HELM—a novel integrated framework that systematically incorporates structured prompting strategies (including chain-of-thought, self-consistency, least-to-most, and program-of-thought) into standardized evaluation. We conduct reproducible, large-scale assessments of state-of-the-art LMs across seven diverse benchmarks—four general-purpose and three medical—using declarative prompt optimization to explicitly elicit and enhance model reasoning capabilities. Our approach significantly improves evaluation robustness: average accuracy increases by 4%, result variance decreases by 2%, and ranking reversals occur in 3/7 leaderboards, yielding more accurate estimates of true model capability ceilings. All prompt optimization pipelines and integration tools are open-sourced to strengthen decision utility and experimental reproducibility.
Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.
This work addresses the limited generalization of prompt-based large language model (LLM) classifiers in data-scarce scenarios, where insufficient fine-tuning often hinders performance. The authors propose a multi-task prompt fine-tuning approach that designs task-specific prompts while integrating general instruction tuning, substantially improving classification accuracy on unseen domains and novel prompts. Notably, they find that supervised classification training without explicit reasoning capabilities can effectively generalize to reasoning-intensive tasks such as summarization. To mitigate performance degradation caused by prompt variations, a hybrid training strategy is introduced. Experimental results demonstrate strong performance on related unseen tasks, highlighting the potential of classification-oriented training for building versatile, general-purpose monitoring systems.
This work addresses the limited performance of large language models (LLMs) in high-dimensional software engineering optimization tasks, where they often fail to surpass Bayesian optimization. For the first time, it systematically compares human- and AI-generated domain knowledge injection strategies and introduces four novel architectures: Human-feedback-informed Domain Knowledge Prompting (H-DKP), Adaptive Multi-stage Prompting (AMP), Dimension-aware Progressive Refinement (DAPR), and a hybrid approach combining statistical scouting with RAG-enhanced knowledge integration (HKMA). By leveraging a multi-stage, dimension-aware, and hybrid knowledge fusion framework, the proposed methods effectively incorporate structured domain knowledge to significantly enhance LLMs’ ability to generate high-quality initial solutions. Evaluated on the MOOT high-dimensional benchmark, the approaches markedly reduce the Chebyshev distance to the optimal solution and, according to Scott-Knott clustering, outperform existing LLM warm-start baselines.
This study investigates whether large language models amplify biases present in user prompts and remain susceptible to prompt framing even on factual questions. By employing a controlled experimental design, the authors construct 160 prompts spanning ten topics to systematically disentangle the effects of implicit prompt framing from explicit manipulation on model outputs. The evaluation across six prominent large language models reveals a consistent tendency for models to align their responses with the framing of the prompt, often prioritizing user suggestions over factual consistency—even when objective facts are unambiguous. This work provides the first empirical evidence of the vulnerability of large language models to bias induction in factual domains, highlighting a critical limitation in their reliability despite advances in scale and training.