Score
Designs and implements preprocessing pipelines that parse, enrich, and structurally format inputs before they are submitted as prompts — e.g., prompt parsing and structured prompt construction, conversion of raw sensor readings into textual or flagged prompt fragments, and insertion of threshold-aware or compact environmental summaries. Builds prompt templates, enrichment heuristics, and lightweight summarizers/formatters that shape prompt content to improve downstream model relevance, robustness, and accuracy–latency trade-offs.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
This study addresses the lack of a unified and reproducible taxonomy for “prompt patterns” in existing research. Focusing on single-turn textual prompts, it proposes the first systematic classification comprising 30 distinct and well-defined prompt patterns, organized along two orthogonal dimensions. Through a comprehensive literature review, pattern identification, and taxonomic methodology, the work establishes a structured and reproducible knowledge framework. By standardizing terminology and definitions, this classification provides a foundational reference for prompt engineering, significantly enhancing comparability and reproducibility across related studies.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
This study addresses critical limitations of large language models (LLMs) in life science applications—including unreliable responses, hallucination, and multi-turn degradation—by proposing a systematic prompt engineering framework. Methodologically, it synthesizes 58 existing prompt techniques into six high-impact strategies: zero-/few-shot prompting, chain-of-thought generation, model ensembling, self-critique, task decomposition, and structured feedback; these are empirically validated across OpenAI and Anthropic platforms using Claude Code agents and Deep Research capabilities. Its key contribution lies in establishing domain-specific prompt design principles for life sciences, significantly enhancing accuracy and robustness in literature summarization, data extraction, and text editing. Experiments demonstrate a 37% reduction in manual intervention frequency and a 42% improvement in output reliability, advancing prompt engineering from ad hoc experimentation toward a reusable, interpretable scientific infrastructure.
Dependency parsing aims to model syntactic dependency relations among words in a sentence. This paper proposes a purely prompt-driven, text-to-text dependency parsing approach that leverages only pretrained sequence-to-sequence encoders (e.g., T5 or BART), eliminating task-specific decoders and parameters. Our method introduces three key innovations: (1) a structured prompt template that explicitly encodes the topological constraints of dependency trees; (2) a linearized textual serialization of dependency structures; and (3) high-accuracy parsing without any parameter updates—i.e., zero-parameter fine-tuning. Experiments across multilingual benchmarks demonstrate that our approach matches or surpasses state-of-the-art parsers in accuracy, while exhibiting strong cross-model and cross-lingual plug-and-play capability and zero-shot transfer performance. To our knowledge, this is the first work to empirically validate the effectiveness and generalizability of the pure prompting paradigm for syntactic parsing.
To address the high computational cost in large language model (LLM) prompt engineering—stemming from repeated LLM executions required to evaluate prompt-induced syntactic structure generation—this paper introduces a predictive prompt analysis paradigm: the first method capable of forecasting, without executing the LLM, how a given prompt influences the frequency of target syntactic structures. Our core method is the Syntactic Prevalence Analyzer (SPA), a sparse autoencoder (SAE)-based model that maps prompts into a syntactic structure space and quantifies their generative propensity toward specific structures. Evaluated on code synthesis tasks, SPA achieves highly accurate predictions of syntactic structure frequencies—attaining a Pearson correlation coefficient of 0.994—while incurring only 0.4% of the LLM’s inference time overhead. This enables efficient, compute-light prompt design and substantially improves resource utilization in syntactic-aware prompting.
Non-expert users struggle to efficiently optimize LLM prompts due to limited domain knowledge and insufficient feedback mechanisms. To address this, we propose a beginner-oriented visual prompt engineering system featuring a novel tri-strategy collaborative optimization framework—integrating keyword perturbation, semantic paraphrasing, and optimal few-shot example recommendation. We design a multi-view synchronized interface, an interactive prompt editing environment, and a real-time evaluation mechanism grounded in both semantic similarity and task-specific accuracy. Experiments demonstrate that our system reduces user prompt iteration time by 37%, increases prompt diversity by 2.1×, and improves average accuracy by 11.4% across multiple NLP tasks—significantly outperforming existing prompt interfaces. This work lowers the cognitive barrier to prompt engineering and establishes a new paradigm for LLM interaction that is interpretable, iterative, and empirically evaluable for non-experts.
This study investigates how prompt engineering can enhance the performance, reliability, and interpretability of large language models (LLMs) in data analysis tasks while addressing standardization and ethical challenges. We systematically evaluate structured prompting, Chain-of-Thought reasoning, and automated prompt optimization techniques across diverse domains—including healthcare, materials science, finance, and business intelligence—to uncover the interplay among prompt complexity, model architecture, and task performance. Experimental results demonstrate that the proposed approaches yield performance improvements of 6% to over 30% on multiple real-world tasks, underscoring the significant potential of advanced prompting frameworks to strengthen LLMs’ contextual adaptation and practical deployment efficacy.
Large language models (LLMs) exhibit unstable summarization quality and limited controllability over abstraction levels. Method: This paper proposes a controllable abstractive summarization framework based on multi-stage prompt engineering, integrating semantic analysis, topic modeling, and noise-aware control to enable adjustable abstraction granularity. We systematically investigate the impact of prompt length, data noise, and text genre on summarization performance using the CNN/Daily Mail benchmark. Contribution/Results: Experiments demonstrate that medium-length prompts yield statistically significant improvements in ROUGE-L scores; increased input noise degrades performance consistently; and LLMs generalize best on news-domain texts. The framework provides an interpretable, configurable pathway to enhance accuracy, consistency, and abstraction-level control in LLM-generated summaries.
This study addresses the heavy reliance of large language models on prompt design for code summarization tasks and the absence of systematic comparisons and unified evaluation standards across diverse prompting strategies. Through a comprehensive literature review, it integrates and categorizes mainstream approaches—including few-shot prompting, chain-of-thought reasoning, retrieval-augmented generation, and zero-shot learning—and analyzes their effectiveness across different models and scenarios. The work highlights the limitations of current evaluations that overly depend on surface-level overlap metrics, delineates the conditions under which each prompting paradigm performs best, and proposes a unified evaluation framework to guide future research and practical applications in this domain.
Manual design of high-quality prompts is challenging in few-shot settings, while existing automated methods suffer from low efficiency and reliance on human-crafted demonstrations. Method: We propose ShapleyPrompt—the first framework to incorporate Monte Carlo Shapley values into in-context learning prompt construction. It quantifies the marginal contribution of each candidate example to model performance, enabling efficient, differentiable-budget-aware example selection and dynamic editing (addition, deletion, or retention). Our approach integrates aggressive sub-sampling, replay buffering, and LLM-based zero-/few-shot evaluation—eliminating the need for human-authored examples. Results: ShapleyPrompt significantly outperforms state-of-the-art automated prompting methods on text simplification, GSM8K, and multi-class classification. With increased computational budget, it establishes new SOTAs across all three tasks, demonstrating superior data efficiency and example quality.
To address the high inference cost of large language models (LLMs) in agent workflows caused by lengthy prompts and multi-source data streams, this paper proposes an end-to-end prompt-and-data co-compression framework. Methodologically, it innovatively integrates hard prompt compression—pruning low-information tokens via self-information scoring and dependency-aware phrase grouping—with lightweight file-level compression—applying n-gram abbreviation for textual data and uniform quantization for numerical data—to jointly handle heterogeneous text and numeric inputs. The framework further supports real-time visualization of compression decisions and cost–performance Pareto analysis. Evaluated on benchmarks including TAT-QA and FinQA, it achieves up to 60% reduction in token usage and inference cost, while maintaining output quality degradation of less than 5% for Claude-3.5-Sonnet and GPT-4.1-Mini—significantly outperforming existing baselines.