Score
Designs, implements, and evaluates prompt-based elicitation methods that cause language models to produce explicit stepwise deliberation; this includes creating think-prefixed prompts and task-specific post-prompts, manipulating token-activation patterns to force ‘think’ tokens, and measuring the resulting stepwise reasoning, answer compatibility, and context effects.
Large language models (LLMs) exhibit limited multi-step reasoning capabilities—e.g., in elementary mathematics, logical deduction, combinatorial games, and robotic planning—when deployed without fine-tuning. Method: We systematically survey prompt-driven reasoning mechanisms, introducing the first structured taxonomy for LLM reasoning; empirically demonstrate that prompts can elicit metacognitive behaviors such as self-reflection and self-correction; formally define “reasoning by LLMs” as a distinct challenge beyond pattern matching; and integrate chain-of-thought prompting, self-consistency decoding, reasoning-path evaluation, and reinforcement-learning-inspired controllable reasoning frameworks. Contribution/Results: Our work clarifies the fundamental boundaries of LLM reasoning, identifies key open challenges, establishes a unified research paradigm, and proposes a verifiable, controllable, and systematic research agenda for advancing reasoning in foundation models.
The widespread assumption of chain-of-thought (CoT) prompting’s universal effectiveness lacks rigorous empirical validation across diverse models and tasks. Method: We conduct a systematic, multi-model, multi-task evaluation using standardized benchmarks, token-level cost analysis, error-pattern statistics, and controlled cross-model experiments. Contribution/Results: We find—contrary to prevailing assumptions—that CoT yields negligible performance gains for native reasoning models and only marginal improvements (≤1.2% accuracy) for non-reasoning models, at the cost of reduced accuracy stability. It incurs an average 47% increase in response latency and a 3.8× rise in token consumption, while introducing additional logical errors. Our work is the first to empirically refute CoT’s general efficacy, identifying three critical practical bottlenecks: diminishing returns on performance gain, increased output volatility, and sharply escalated inference overhead—providing essential evidence for rational prompt engineering.
This study investigates whether large language models amplify biases present in user prompts and remain susceptible to prompt framing even on factual questions. By employing a controlled experimental design, the authors construct 160 prompts spanning ten topics to systematically disentangle the effects of implicit prompt framing from explicit manipulation on model outputs. The evaluation across six prominent large language models reveals a consistent tendency for models to align their responses with the framing of the prompt, often prioritizing user suggestions over factual consistency—even when objective facts are unambiguous. This work provides the first empirical evidence of the vulnerability of large language models to bias induction in factual domains, highlighting a critical limitation in their reliability despite advances in scale and training.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
This work investigates whether large language models (LLMs) exhibit cognitive dissonance—systematic inconsistency between their stated answers to multiple-choice questions (MCQs) and their revealed beliefs, operationalized as next-token probability distributions over answer options. Method: We propose a novel evaluation framework grounded in raw text completion, distinguishing and quantifying the stated answer versus revealed belief through multi-outcome scenario design, token-level probability analysis, and controlled prompt ablation experiments. Contribution/Results: We find that LLMs frequently select correct answers while simultaneously exhibiting biases in their probability distributions—including causal misattribution, miscalibrated uncertainty estimation, and sluggish evidence updating—thereby exposing critical reasoning flaws obscured by conventional MCQ accuracy metrics. Our findings challenge the reliability of unstructured generative outputs as proxies for robust reasoning and provide both theoretical grounding and empirical evidence for developing more trustworthy, belief-aware evaluation paradigms for LLMs.
Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
This work addresses the heavy reliance of current large language models on prompt engineering for generating task planning explanations, a limitation exacerbated by the lack of systematic understanding of how users construct and refine prompts. To bridge this gap, the authors propose COMPASS, a novel method that formalizes users’ cognitive states—such as attention, comprehension, and uncertainty—and integrates implicit cognitive signals with explicit interaction cues through a Partially Observable Markov Decision Process (POMDP). This framework enables adaptive prompt synthesis tailored to task planning explanations. Experimental evaluations on two cyber-physical system case studies demonstrate that COMPASS effectively leverages user cognition and feedback to significantly enhance both the quality and personalization of generated explanations.
This work addresses the instability of reasoning paths in existing chain-of-thought prompting methods, which often stems from insufficient guidance and the limited generalizability of single-strategy approaches across diverse tasks. To overcome these limitations, we propose Diverge-to-Induce Prompting (DIP), a novel framework that first guides large language models to generate multiple high-level reasoning rationales for the same problem, then refines these into detailed step-by-step drafts, and finally integrates them into a unified reasoning plan. DIP introduces, for the first time, a mechanism for generating and fusing multiple reasoning rationales, significantly enhancing the robustness and accuracy of zero-shot reasoning without requiring extensive sampling. Experimental results demonstrate that DIP consistently outperforms current single-strategy prompting methods across multiple benchmark tasks.
This study addresses the instability of large language models (LLMs) across different prompting styles, a phenomenon whose underlying mechanisms remain poorly understood. By systematically comparing instruction-based and exemplar-based prompts through attention head analysis, representational probing, and behavioral experiments, the work identifies for the first time a shared “lexical task attention head” that operates consistently across prompting paradigms. The activation strength of this attention head strongly correlates with model performance and reliably triggers task-relevant answer generation. Furthermore, the research reveals that competition among internal task representations is a key driver of prompt sensitivity. These findings offer an interpretable framework for understanding LLM internal dynamics and suggest novel directions for enhancing prompt robustness.
Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.