chain-of-thought prompting

Designs and implements prompt strategies and prompting pipelines that elicit, structure, and generate explicit step-by-step reasoning traces (chain-of-thought) from language models, including retroactive/retrocot and forensic-reconstruction prompts and methods to integrate tool outputs into those traces. Builds evaluation and supervision processes to model, compare, and refine these traces—prompting for diverse or structured reasoning, decomposing multi-step problems, and assessing trace fidelity, consistency, and usefulness for downstream decisions.

chain-of-thoughtprompting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Chain of Methodologies: Scaling Test Time Computation without Training

Jun 08, 2025
CL
Cong Liu
🏛️ Sun Yat-sen University | Temple University

Large language models (LLMs) exhibit limited performance on complex reasoning tasks, primarily due to the absence of structured methodological knowledge—such as divide-and-conquer, abductive reasoning, and analogical reasoning—in their training data. Method: We propose Chain-of-Methodology (CoM), a training-free prompting framework that explicitly encodes general human methodologies as reusable templates embeddable within reasoning chains, augmented by a metacognitive guidance mechanism to elicit systematic thinking and self-unfolding inference. CoM requires no fine-tuning or external tools, relying solely on prompt engineering. Contribution/Results: Evaluated across mathematical reasoning, multi-hop question answering, and scientific reasoning benchmarks, CoM consistently outperforms state-of-the-art prompting methods—including Chain-of-Thought (CoT) and Auto-CoT—demonstrating that injecting structured methodology effectively bridges the gap between LLM inference and human-like reasoning paradigms. This work establishes a novel, training-agnostic paradigm for high-order reasoning.

Activating systematic reasoning without explicit fine-tuningBridging the gap toward human-level reasoning with methodologiesEnhancing LLMs' structured thinking for complex reasoning tasks

Demystifying Chains, Trees, and Graphs of Thoughts

Jan 25, 2024
MB
Maciej Besta
🏛️ ETH Zurich | Dell | Cledar | BASF SE

Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.

Analyze performance and cost of different prompting designsDevelop taxonomy for structure-enhanced LLM reasoning schemesEnhance LLM reasoning with structured prompt engineering

Large language models (LLMs) frequently fail in real-world tool invocation due to intent misinterpretation, incorrect parsing of tool documentation, and parameterization errors. To address this, we propose a curriculum-inspired structured reasoning framework that replaces free-form chain-of-thought prompting with guided, template-based reasoning—explicitly decoupling the process into three sequential stages: *intent parsing*, *tool matching*, and *parameter generation*. Our framework employs stepwise structured prompts to jointly model user goals and tool functionalities, thereby enhancing invocation robustness and decision interpretability. Evaluated across multiple state-of-the-art models (e.g., LLaMA-3, Qwen2) and benchmarks (ToolBench, API-Bank), it reduces relative error rates by 3–12% over strong baselines. The core contribution lies in transforming implicit, unstructured reasoning into an explicit, traceable, and modular pipeline—balancing accuracy with transparency and auditability.

Addresses incorrect parameterization and poor tool selection in LLMsImproves incomplete understanding of user goals and tool documentationSolves misinterpretation of user intent in function-calling tasks

Enhancing Chain of Thought Prompting in Large Language Models via Reasoning Patterns

Apr 23, 2024
YZ
Yufeng Zhang
🏛️ University of Chinese Academy of Sciences | Chinese Academy of Sciences

Existing unsupervised chain-of-thought (CoT) prompting methods rely on semantic similarity for in-context example selection, which often introduces noise and suffers from poor interpretability, thereby limiting multi-step reasoning performance. To address this, we propose a reasoning-pattern-based demonstration selection framework. Our approach explicitly models implicit reasoning processes as structured “reasoning patterns”—a novel conceptualization—leveraging large language model priors, prompt engineering, and pattern clustering to construct task-specific, diverse, and interpretable pattern sets that guide reasoning along semantically coherent paths. By decoupling example selection from surface-level semantics and grounding it in latent reasoning structures, our method significantly reduces selection noise. Empirical evaluation across mathematical reasoning, commonsense reasoning, and other multi-step reasoning benchmarks demonstrates consistent performance gains, enhanced robustness, improved transparency, and greater controllability of CoT generation.

Enhance CoT prompting via reasoning patterns.Improve interpretability and reduce noise in CoT.Select diverse demonstrations using task-specific reasoning patterns.

Chain of Thoughtlessness? An Analysis of CoT in Planning

May 08, 2024
KS
Kaya Stechly
🏛️ Arizona State University

This work investigates the generalization capability of chain-of-thought (CoT) prompting in large language models (LLMs) for reasoning, focusing on the canonical planning domain Blocksworld. Method: We conduct a systematic empirical analysis using two state-of-the-art LLMs on controlled-complexity Blocksworld tasks and scalable CoT benchmark variants. Contribution/Results: We find that CoT performance critically depends on strict structural alignment—e.g., stack height—between exemplars and queries, exhibiting negligible generalization across problem complexity or syntactic form. Its gains stem from problem-specific pattern matching rather than acquisition of general algorithms. This study provides the first evidence of a fundamental generalization bottleneck for CoT in classical planning and quantifies a significant trade-off between CoT efficacy and the human effort required to engineer high-quality reasoning traces. These findings challenge the prevailing hypothesis that CoT enables implicit algorithm learning.

Chain of thought prompts require highly specific examples for improvement.LLM performance on reasoning problems lacks generalization.Performance gains from CoT depend on problem-specific prompt engineering.

Latest Papers

What's happening recently
View more

Generating Verifiable CoT from Execution-Traces

Nov 28, 2025
ST
Shailja Thakur
🏛️ IBM Research

Existing synthetic chain-of-thought (CoT) data often relies on teacher models to generate “plausible-sounding” yet unverifiable reasoning steps, leading language models to internalize logical hallucinations. To address this, we propose Execution-Traced CoT: a method that instruments code execution to capture ground-truth program traces and structurally maps them to natural-language reasoning steps—each strictly verifiable via observable program behavior. This enables bidirectional verifiability: forward (execution → reasoning) and backward (reasoning → execution). Using this approach, we construct high-fidelity training data and perform supervised fine-tuning of language models. On code reasoning benchmarks, our method improves prediction accuracy by up to 30 percentage points (output) and 28 percentage points (input), while substantially enhancing logical consistency and trustworthiness in both code generation and explanation.

Addresses logical flaws in synthetic training dataGenerates verifiable reasoning from execution tracesImproves code reasoning and generation tasks

This work addresses the susceptibility of large language models to hallucination, reasoning drift, and insufficient interpretability in safety-critical tasks by proposing the first structured prompt engineering framework tailored for locally deployed scenarios. The framework explicitly guides models to generate reliable and auditable chains of thought through four complementary dimensions: contextual and scope control, evidence anchoring with traceability, structured reasoning with cognitive control, and safety-specific analytical constraints. Empirical evaluations demonstrate that the approach yields up to a 40% improvement in reasoning performance across multiple model families, maintains consistent gains across model scales, and achieves high inter-annotator agreement in human assessments (Cohen’s κ > 0.80), substantially enhancing reasoning completeness, robustness to interference, and practical utility.

Chain-of-Thoughthallucinationprompt engineering

This work addresses the inefficiency and logical fragmentation often observed in traditional Chain-of-Thought (CoT) prompting during complex multi-step reasoning, which frequently arises from redundant intermediate steps. To overcome these limitations, the authors propose Hierarchical Chain-of-Thought (Hi-CoT), a novel approach that introduces a structured, hierarchical reasoning paradigm. Hi-CoT alternates between high-level directive planning and low-level step-by-step execution, thereby decomposing intricate tasks into logically coherent sub-steps. Empirical evaluations demonstrate that this method substantially enhances both accuracy and efficiency in long-horizon reasoning for large language models. Across multiple mainstream models and mathematical reasoning benchmarks, Hi-CoT achieves an average accuracy improvement of 6.2%—reaching up to 61.4% in certain cases—while simultaneously reducing reasoning trajectory length by 13.9%, underscoring the critical role of a strict hierarchical structure in boosting performance.

Chain-of-Thoughthierarchical reasoninglarge language models

Large language models (LLMs) often produce incorrect final answers in complex reasoning tasks due to undetected errors in intermediate reasoning steps; existing chain-of-thought (CoT) prompting methods lack explicit mechanisms for error identification and correction. To address this, we propose Error-Reflective Prompting (ERP), the first prompting framework that integrates automated error detection, attribution, and correction directly into the CoT process: after generating an initial answer, the model autonomously backtracks through its reasoning trace, pinpoints erroneous steps, constructs an “error profile,” and regenerates a corrected solution. ERP requires no additional training or fine-tuning—only structured prompting enables self-reflection. Experiments across mathematical reasoning and commonsense question answering demonstrate that ERP significantly improves accuracy and stability while enhancing interpretability and robustness of reasoning traces. ERP thus provides a general, lightweight, prompt-level solution for trustworthy LLM reasoning.

Addressing Chain-of-Thought's inability to reflect on mistakesEnhancing error recognition and correction in language modelsImproving reasoning reliability through error analysis steps

Current evaluations of large language model reasoning predominantly rely on final answer accuracy or superficial statistical features, which inadequately capture the quality of reasoning processes in open-ended outputs. This work proposes TRACE, a novel metric that, for the first time, integrates Toulmin’s argumentation model with Flavell’s metacognitive framework to perform fine-grained structural analysis of chain-of-thought reasoning, thereby enabling quantitative assessment of the intrinsic quality of reasoning construction. TRACE can serve as a reward signal in reinforcement learning. Experiments across seven models and 26.3K question-answer pairs demonstrate that TRACE exhibits strong correlation with benchmark accuracy (r = 0.74) and significantly outperforms reinforcement learning baselines that rely solely on answer accuracy.

argumentation structureChain-of-Thoughtlarge language models

Hot Scholars

DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
MY

Mo Yu

WeChat AI, Tencent
NLPQuestion AnsweringInformation ExtractionMachine Reading Comprehension
RZ

Renrui Zhang

Seed ByteDance & MMLab & PKU
Large Multimodal ModelGenerative ModelEmbodied AI
ZG

Ziyu Guo

The Chinese University of Hong Kong
Multi-modality LearningLLM/VLMs3D Vision