Score
Designs and implements prompt strategies and prompting pipelines that elicit, structure, and generate explicit step-by-step reasoning traces (chain-of-thought) from language models, including retroactive/retrocot and forensic-reconstruction prompts and methods to integrate tool outputs into those traces. Builds evaluation and supervision processes to model, compare, and refine these traces—prompting for diverse or structured reasoning, decomposing multi-step problems, and assessing trace fidelity, consistency, and usefulness for downstream decisions.
This work addresses the fragmented development of Chain-of-X (CoX) methods for enhancing large language model (LLM) reasoning. We propose the first systematic, unified CoX taxonomy, organizing over 30 variants—including Chain-of-Thought (CoT), Chain-of-Symbol (CoS), and Chain-of-Uncertainty (CoU)—along two orthogonal dimensions: *node type* (X) and *task scenario*. The framework encompasses 12 X-node categories (e.g., evidence, action, uncertainty) and five major application domains. Through integrated analysis of prompt engineering, cognitive modeling, and empirical evaluation, we uncover the underlying modeling principles, applicability boundaries, and cross-task transfer patterns of distinct X-nodes, distilling reusable prompt design guidelines. We further establish a structured evaluation methodology to assess CoX efficacy rigorously. Our contributions provide both theoretical foundations and practical guidance for interpretable LLM reasoning and domain-adaptive inference.
Large language models (LLMs) exhibit limited performance on complex reasoning tasks, primarily due to the absence of structured methodological knowledge—such as divide-and-conquer, abductive reasoning, and analogical reasoning—in their training data. Method: We propose Chain-of-Methodology (CoM), a training-free prompting framework that explicitly encodes general human methodologies as reusable templates embeddable within reasoning chains, augmented by a metacognitive guidance mechanism to elicit systematic thinking and self-unfolding inference. CoM requires no fine-tuning or external tools, relying solely on prompt engineering. Contribution/Results: Evaluated across mathematical reasoning, multi-hop question answering, and scientific reasoning benchmarks, CoM consistently outperforms state-of-the-art prompting methods—including Chain-of-Thought (CoT) and Auto-CoT—demonstrating that injecting structured methodology effectively bridges the gap between LLM inference and human-like reasoning paradigms. This work establishes a novel, training-agnostic paradigm for high-order reasoning.
Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.
Large language models (LLMs) frequently fail in real-world tool invocation due to intent misinterpretation, incorrect parsing of tool documentation, and parameterization errors. To address this, we propose a curriculum-inspired structured reasoning framework that replaces free-form chain-of-thought prompting with guided, template-based reasoning—explicitly decoupling the process into three sequential stages: *intent parsing*, *tool matching*, and *parameter generation*. Our framework employs stepwise structured prompts to jointly model user goals and tool functionalities, thereby enhancing invocation robustness and decision interpretability. Evaluated across multiple state-of-the-art models (e.g., LLaMA-3, Qwen2) and benchmarks (ToolBench, API-Bank), it reduces relative error rates by 3–12% over strong baselines. The core contribution lies in transforming implicit, unstructured reasoning into an explicit, traceable, and modular pipeline—balancing accuracy with transparency and auditability.
Existing unsupervised chain-of-thought (CoT) prompting methods rely on semantic similarity for in-context example selection, which often introduces noise and suffers from poor interpretability, thereby limiting multi-step reasoning performance. To address this, we propose a reasoning-pattern-based demonstration selection framework. Our approach explicitly models implicit reasoning processes as structured “reasoning patterns”—a novel conceptualization—leveraging large language model priors, prompt engineering, and pattern clustering to construct task-specific, diverse, and interpretable pattern sets that guide reasoning along semantically coherent paths. By decoupling example selection from surface-level semantics and grounding it in latent reasoning structures, our method significantly reduces selection noise. Empirical evaluation across mathematical reasoning, commonsense reasoning, and other multi-step reasoning benchmarks demonstrates consistent performance gains, enhanced robustness, improved transparency, and greater controllability of CoT generation.
This work investigates the generalization capability of chain-of-thought (CoT) prompting in large language models (LLMs) for reasoning, focusing on the canonical planning domain Blocksworld. Method: We conduct a systematic empirical analysis using two state-of-the-art LLMs on controlled-complexity Blocksworld tasks and scalable CoT benchmark variants. Contribution/Results: We find that CoT performance critically depends on strict structural alignment—e.g., stack height—between exemplars and queries, exhibiting negligible generalization across problem complexity or syntactic form. Its gains stem from problem-specific pattern matching rather than acquisition of general algorithms. This study provides the first evidence of a fundamental generalization bottleneck for CoT in classical planning and quantifies a significant trade-off between CoT efficacy and the human effort required to engineer high-quality reasoning traces. These findings challenge the prevailing hypothesis that CoT enables implicit algorithm learning.
Existing synthetic chain-of-thought (CoT) data often relies on teacher models to generate “plausible-sounding” yet unverifiable reasoning steps, leading language models to internalize logical hallucinations. To address this, we propose Execution-Traced CoT: a method that instruments code execution to capture ground-truth program traces and structurally maps them to natural-language reasoning steps—each strictly verifiable via observable program behavior. This enables bidirectional verifiability: forward (execution → reasoning) and backward (reasoning → execution). Using this approach, we construct high-fidelity training data and perform supervised fine-tuning of language models. On code reasoning benchmarks, our method improves prediction accuracy by up to 30 percentage points (output) and 28 percentage points (input), while substantially enhancing logical consistency and trustworthiness in both code generation and explanation.
This work addresses the susceptibility of large language models to hallucination, reasoning drift, and insufficient interpretability in safety-critical tasks by proposing the first structured prompt engineering framework tailored for locally deployed scenarios. The framework explicitly guides models to generate reliable and auditable chains of thought through four complementary dimensions: contextual and scope control, evidence anchoring with traceability, structured reasoning with cognitive control, and safety-specific analytical constraints. Empirical evaluations demonstrate that the approach yields up to a 40% improvement in reasoning performance across multiple model families, maintains consistent gains across model scales, and achieves high inter-annotator agreement in human assessments (Cohen’s κ > 0.80), substantially enhancing reasoning completeness, robustness to interference, and practical utility.
This work addresses the inefficiency and logical fragmentation often observed in traditional Chain-of-Thought (CoT) prompting during complex multi-step reasoning, which frequently arises from redundant intermediate steps. To overcome these limitations, the authors propose Hierarchical Chain-of-Thought (Hi-CoT), a novel approach that introduces a structured, hierarchical reasoning paradigm. Hi-CoT alternates between high-level directive planning and low-level step-by-step execution, thereby decomposing intricate tasks into logically coherent sub-steps. Empirical evaluations demonstrate that this method substantially enhances both accuracy and efficiency in long-horizon reasoning for large language models. Across multiple mainstream models and mathematical reasoning benchmarks, Hi-CoT achieves an average accuracy improvement of 6.2%—reaching up to 61.4% in certain cases—while simultaneously reducing reasoning trajectory length by 13.9%, underscoring the critical role of a strict hierarchical structure in boosting performance.
Large language models (LLMs) often produce incorrect final answers in complex reasoning tasks due to undetected errors in intermediate reasoning steps; existing chain-of-thought (CoT) prompting methods lack explicit mechanisms for error identification and correction. To address this, we propose Error-Reflective Prompting (ERP), the first prompting framework that integrates automated error detection, attribution, and correction directly into the CoT process: after generating an initial answer, the model autonomously backtracks through its reasoning trace, pinpoints erroneous steps, constructs an “error profile,” and regenerates a corrected solution. ERP requires no additional training or fine-tuning—only structured prompting enables self-reflection. Experiments across mathematical reasoning and commonsense question answering demonstrate that ERP significantly improves accuracy and stability while enhancing interpretability and robustness of reasoning traces. ERP thus provides a general, lightweight, prompt-level solution for trustworthy LLM reasoning.
Current evaluations of large language model reasoning predominantly rely on final answer accuracy or superficial statistical features, which inadequately capture the quality of reasoning processes in open-ended outputs. This work proposes TRACE, a novel metric that, for the first time, integrates Toulmin’s argumentation model with Flavell’s metacognitive framework to perform fine-grained structural analysis of chain-of-thought reasoning, thereby enabling quantitative assessment of the intrinsic quality of reasoning construction. TRACE can serve as a reward signal in reinforcement learning. Experiments across seven models and 26.3K question-answer pairs demonstrate that TRACE exhibits strong correlation with benchmark accuracy (r = 0.74) and significantly outperforms reinforcement learning baselines that rely solely on answer accuracy.