Score
Designs, builds, or evaluates structured chain-of-thought artifacts and annotation schemes where each intermediate reasoning step is explicitly linked to specific supporting evidence or observations. This competence includes methods to generate, annotate, validate, and trace stepwise (stepwise CoT) inferences so that each inference is evidence‑anchored and usable for provenance, viewpoint, or proximity-based reasoning.
This work challenges the prevailing assumption that longer chain-of-thought (CoT) reasoning inherently yields better performance, systematically investigating how CoT length, backtracking behavior, and structural properties affect reasoning efficacy in large reasoning models (LRMs) on mathematical and scientific tasks. Method: We model CoT as a directed graph and introduce the Failure Step Fraction (FSF)—the ratio of erroneous or unproductive reasoning steps—as a core structural quality metric. Combining graph-theoretic analysis, token-level measurements, and test-time interventions—including candidate CoT ranking and failure-branch pruning—we conduct causal validation. Contribution/Results: Experiments demonstrate that concise, structurally coherent CoTs significantly outperform lengthy, disorganized ones; FSF predicts answer correctness more reliably than CoT length or backtracking frequency; and targeted editing of failure branches improves model accuracy. This study pioneers a structural perspective on CoT effectiveness, establishing a novel paradigm for interpretable reasoning evaluation and optimization.
This work addresses the challenge of diagnosing errors in chain-of-thought (CoT) reasoning generated by large language models, which are often verbose and prone to logical or factual inaccuracies. To this end, the authors propose the first step-level error detection method that integrates external fact-checking with symbolic logical verification. They further develop ReasonDiag, an interactive visualization system that combines arc diagrams and hierarchical node-link graphs to reveal the reasoning flow and trace error propagation paths. Through technical evaluation, two case studies, and user interviews with 16 participants, the study demonstrates that ReasonDiag effectively supports users in comprehending complex reasoning processes, accurately identifying erroneous steps, and tracing underlying root causes.
This work addresses the limitations of existing chain-of-thought (CoT) reasoning in medical visual question answering (VQA), which often adopts unstructured, free-form formats that misalign with clinicians’ systematic diagnostic workflows, thereby compromising both accuracy and interpretability. To bridge this gap, the authors introduce Step-CoT, a large-scale medical reasoning dataset comprising over 10K clinical cases and 70K question-answer pairs, featuring the first structured multi-step reasoning annotations aligned with real-world clinical workflows. They further propose a teacher–student learning framework coupled with a dynamic graph-based focusing mechanism to guide the model along clinically plausible reasoning paths. Experimental results demonstrate that the proposed approach significantly enhances both the reasoning performance and interpretability of medical VQA systems.
Large language models (LLMs) frequently exhibit “early answering”—producing final answers before generating Chain-of-Thought (CoT) reasoning—raising fundamental questions: Is CoT necessary? Does answer correctness imply correct reasoning? Method: The authors introduce Chain-of-Probe, the first framework to quantitatively decouple CoT necessity from reasoning accuracy. It employs dynamic neuron probing, inter-layer state difference analysis, and step-wise reasoning modeling to assess when and why CoT is required. Contributions/Results: Chain-of-Probe reveals that over 50% of correct answers stem from flawed reasoning; establishes a strong correlation between task complexity and CoT necessity; and enables answer re-ranking based on reasoning trustworthiness. Evaluated on GSM8K and other benchmarks, it improves reasoning trustworthiness by 23%, significantly enhancing both interpretability and reliability of LLM outputs.
Existing LLM reasoning methods—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and ReACT—are hindered by insufficient context utilization, hallucinated intermediate steps, and inefficient iteration, compromising accuracy and robustness. To address these limitations, we propose E2G, a novel “evidence-first” single-agent two-step prompting framework. In the first step, E2G precisely extracts explicit, structured contextual reasoning sequences directly from the input as verifiable evidence; in the second step, it generates answers strictly grounded in this evidence, eliminating unvalidated intermediate reasoning. By integrating retrieval augmentation and reconstructing the CoT paradigm around evidence fidelity, E2G significantly enhances reasoning reliability and context awareness. Experiments demonstrate that E2G achieves 53.8% accuracy on LogiQA—outperforming standard CoT by 18 percentage points—and attains an F1 score of 83.3 on the DROP subset when paired with PaLM2, surpassing Gemini Ultra by 0.9 points.
This study systematically evaluates the faithfulness of open-source reasoning models in chain-of-thought (CoT) generation—specifically, whether their outputs genuinely reflect their underlying reasoning processes. Conducting 41,832 inference trials across 498 MMLU and GPQA questions, the authors inject six categories of reasoning prompts into twelve models and employ keyword-based analysis to distinguish acknowledgment behaviors at the thinking-token versus answer-text levels. The work reveals, for the first time, a significant disconnect between models’ internal cognition and external expression: faithfulness varies widely from 39.7% to 89.9%, with consistency- and flattery-oriented prompts exhibiting the lowest acknowledgment rates. Crucially, faithfulness is primarily influenced by model architecture and training methodology rather than parameter count, offering new empirical insights and a foundation for improving CoT reliability.
This study investigates whether chain-of-thought (CoT) reasoning traces faithfully reflect a model’s actual internal decision-making process, thereby questioning their reliability as a supervisory and auditing mechanism. To this end, the authors propose a step-level Detect-Classify-Compare framework, integrating multidimensional validation techniques—including answer-commitment agents, Patchscopes, tuned-lens probes, causal ablation, truncation experiments, and donor contamination tests. Experiments across nine models and seven reasoning benchmarks reveal that, on average, only 61.9% of CoT steps align with the model’s internal computations; in 58% of misaligned cases, models generate redundant “reasoning” after the answer has already been determined—a phenomenon termed “hallucinated continuation.” Notably, stronger CoT performance correlates with lower temporal fidelity. This work provides the first systematic evidence of a fundamental disconnect between CoT traces and genuine reasoning dynamics, challenging the core assumption that CoT serves as a faithful reasoning log.
This work addresses the unresolved question of which components within chain-of-thought (CoT) reasoning genuinely drive large language models toward correct answers. The authors propose a novel metric termed “potential” to quantify the contribution of individual CoT segments to the final answer and analyze the dynamics of reasoning trajectories on competition-level mathematical problems. Through experiments involving CoT generation, potential-based quantification, and cross-model transfer, the study reveals—for the first time—the non-monotonicity, reasoning leaps, and instances of lucky guessing inherent in CoT processes. Crucially, the research demonstrates that high-potential CoT fragments exhibit strong transferability: merely 20% of such segments extracted from a stronger model can substantially enhance a weaker model’s performance on problems it originally could not solve.
Current evaluations of large language model reasoning predominantly rely on final answer accuracy or superficial statistical features, which inadequately capture the quality of reasoning processes in open-ended outputs. This work proposes TRACE, a novel metric that, for the first time, integrates Toulmin’s argumentation model with Flavell’s metacognitive framework to perform fine-grained structural analysis of chain-of-thought reasoning, thereby enabling quantitative assessment of the intrinsic quality of reasoning construction. TRACE can serve as a reward signal in reinforcement learning. Experiments across seven models and 26.3K question-answer pairs demonstrate that TRACE exhibits strong correlation with benchmark accuracy (r = 0.74) and significantly outperforms reinforcement learning baselines that rely solely on answer accuracy.
This work addresses the significant yet poorly understood performance variations of Chain-of-Thought (CoT) reasoning across different tasks by providing the first theoretical framework for its step-by-step inference process. The authors model CoT as a Markov chain and propose that its effectiveness hinges on the consistency of the transition kernels between reasoning steps, while also quantifying how noise in intermediate steps degrades performance. Through rigorous theoretical analysis, they prove that consistent transition kernels substantially reduce sample complexity. To validate these predictions, they construct synthetic benchmark experiments that align with their theoretical findings and offer a principled explanation for the observed disparities in CoT’s empirical success across real-world tasks. This study thus establishes a novel theoretical lens and analytical framework for understanding and improving CoT reasoning.