activate deliberate reasoning

Designs, implements, and evaluates prompt-based elicitation methods that cause language models to produce explicit stepwise deliberation; this includes creating think-prefixed prompts and task-specific post-prompts, manipulating token-activation patterns to force ‘think’ tokens, and measuring the resulting stepwise reasoning, answer compatibility, and context effects.

activatedeliberatereasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting

Jun 08, 2025
LM
Lennart Meincke
🏛️ WHU | Otto Beisheim School of Management | University of Pennsylvania | Glowforge

The widespread assumption of chain-of-thought (CoT) prompting’s universal effectiveness lacks rigorous empirical validation across diverse models and tasks. Method: We conduct a systematic, multi-model, multi-task evaluation using standardized benchmarks, token-level cost analysis, error-pattern statistics, and controlled cross-model experiments. Contribution/Results: We find—contrary to prevailing assumptions—that CoT yields negligible performance gains for native reasoning models and only marginal improvements (≤1.2% accuracy) for non-reasoning models, at the cost of reduced accuracy stability. It incurs an average 47% increase in response latency and a 3.8× rise in token consumption, while introducing additional logical errors. Our work is the first to empirically refute CoT’s general efficacy, identifying three critical practical bottlenecks: diminishing returns on performance gain, increased output volatility, and sharply escalated inference overhead—providing essential evidence for rational prompt engineering.

Assesses marginal accuracy gains versus higher costs in reasoning-enabled modelsExamines increased answer variability and occasional errors from Chain-of-Thought promptingInvestigates varying effectiveness of Chain-of-Thought prompting across tasks and models

This study investigates whether large language models amplify biases present in user prompts and remain susceptible to prompt framing even on factual questions. By employing a controlled experimental design, the authors construct 160 prompts spanning ten topics to systematically disentangle the effects of implicit prompt framing from explicit manipulation on model outputs. The evaluation across six prominent large language models reveals a consistent tendency for models to align their responses with the framing of the prompt, often prioritizing user suggestions over factual consistency—even when objective facts are unambiguous. This work provides the first empirical evidence of the vulnerability of large language models to bias induction in factual domains, highlighting a critical limitation in their reliability despite advances in scale and training.

bias reinforcementfactual consistencylarge language models

A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Feb 05, 2024
PS
Pranab Sahoo
🏛️ Indian Institute of Technology Patna | Stanford University | Amazon AI

Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.

Analysis of strengths and limitations of prompting approachesOverview of advancements in prompt engineering techniquesSystematic organization of prompt engineering methods

This work investigates whether large language models (LLMs) exhibit cognitive dissonance—systematic inconsistency between their stated answers to multiple-choice questions (MCQs) and their revealed beliefs, operationalized as next-token probability distributions over answer options. Method: We propose a novel evaluation framework grounded in raw text completion, distinguishing and quantifying the stated answer versus revealed belief through multi-outcome scenario design, token-level probability analysis, and controlled prompt ablation experiments. Contribution/Results: We find that LLMs frequently select correct answers while simultaneously exhibiting biases in their probability distributions—including causal misattribution, miscalibrated uncertainty estimation, and sluggish evidence updating—thereby exposing critical reasoning flaws obscured by conventional MCQ accuracy metrics. Our findings challenge the reliability of unstructured generative outputs as proxies for robust reasoning and provide both theoretical grounding and empirical evidence for developing more trustworthy, belief-aware evaluation paradigms for LLMs.

Assessing LLMs' true reasoning via text-completion vs MCQ responsesEvaluating belief updates in LLMs when given new evidenceInvestigating inconsistencies in LLMs' probability distributions and biases

Demystifying Chains, Trees, and Graphs of Thoughts

Jan 25, 2024
MB
Maciej Besta
🏛️ ETH Zurich | Dell | Cledar | BASF SE

Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.

Analyze performance and cost of different prompting designsDevelop taxonomy for structure-enhanced LLM reasoning schemesEnhance LLM reasoning with structured prompt engineering

Latest Papers

What's happening recently
View more

This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.

documentationevaluationprompt engineering

This work addresses the heavy reliance of current large language models on prompt engineering for generating task planning explanations, a limitation exacerbated by the lack of systematic understanding of how users construct and refine prompts. To bridge this gap, the authors propose COMPASS, a novel method that formalizes users’ cognitive states—such as attention, comprehension, and uncertainty—and integrates implicit cognitive signals with explicit interaction cues through a Partially Observable Markov Decision Process (POMDP). This framework enables adaptive prompt synthesis tailored to task planning explanations. Experimental evaluations on two cyber-physical system case studies demonstrate that COMPASS effectively leverages user cognition and feedback to significantly enhance both the quality and personalization of generated explanations.

explanation generationhuman-AI interactionLarge Language Models

This work addresses the instability of reasoning paths in existing chain-of-thought prompting methods, which often stems from insufficient guidance and the limited generalizability of single-strategy approaches across diverse tasks. To overcome these limitations, we propose Diverge-to-Induce Prompting (DIP), a novel framework that first guides large language models to generate multiple high-level reasoning rationales for the same problem, then refines these into detailed step-by-step drafts, and finally integrates them into a unified reasoning plan. DIP introduces, for the first time, a mechanism for generating and fusing multiple reasoning rationales, significantly enhancing the robustness and accuracy of zero-shot reasoning without requiring extensive sampling. Experimental results demonstrate that DIP consistently outperforms current single-strategy prompting methods across multiple benchmark tasks.

Chain-of-Thought promptinglarge language modelsreasoning instability

This study addresses the instability of large language models (LLMs) across different prompting styles, a phenomenon whose underlying mechanisms remain poorly understood. By systematically comparing instruction-based and exemplar-based prompts through attention head analysis, representational probing, and behavioral experiments, the work identifies for the first time a shared “lexical task attention head” that operates consistently across prompting paradigms. The activation strength of this attention head strongly correlates with model performance and reliably triggers task-relevant answer generation. Furthermore, the research reveals that competition among internal task representations is a key driver of prompt sensitivity. These findings offer an interpretable framework for understanding LLM internal dynamics and suggest novel directions for enhancing prompt robustness.

behavioral variabilitylarge language modelslexical task heads

Large language models exhibit inconsistent performance in social science text classification. This study systematically investigates the impact of three key prompt engineering components—label descriptions, instructional guidance, and few-shot examples—on classification accuracy. Through controlled experiments across multiple mainstream large language models, the authors find that moderately enriching prompt context significantly improves performance, whereas excessive augmentation can degrade it. Moreover, the optimal prompt configuration is highly dependent on the specific model, task, and data batch. The findings underscore the necessity of task- and model-specific validation and reveal a “less-is-more” principle in prompt design, offering practical guidance for achieving efficient and stable text classification in social science applications.

LLM classificationperformance varianceprompt engineering

Hot Scholars

EB

Engin Bumbacher

University of Teacher Education Vaud
STEM EducationComputational ThinkingInnovative Assessment
FM

Francesco Mondada

Ecole Polytechnique Fédérale de Lausanne
RoboticsEducationMechatronics
GA

Giorgia Adorni

Institute of Information Systems and Networking (ISIN), SUPSI
Artificial IntelligenceComputer Science EducationLearning TechnologiesGenerative AI
FM

Francesca Mangili

IDSIA, usi-supsi
statisticsimprecise probabilityprognostics and health management