chain-of-thought training

Designs and implements supervised fine‑tuning regimes and curricula that train models to generate explicit step‑by‑step reasoning traces (chain‑of‑thought), including long or visual CoT forms and code‑paired traces; builds or curates expert stepwise explanation datasets, selects progressive learning schedules, and fine‑tunes models to improve multi‑step problem solving, correctness, and the clarity/auditability of generated reasoning.

chain-of-thoughttraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs

Jul 03, 2024
HP
Haritz Puerto
🏛️ TU Darmstadt | Queen's University | University of Bath

Large language models (LLMs) lack intrinsic self-refinement capability within a single forward pass, limiting their ability to correct reasoning errors without external feedback or parallel generation. Method: We propose Diversified Chain-of-Thought (DCoT) fine-tuning—a supervised fine-tuning paradigm that constructs structured DCoT datasets by integrating diversity-aware sampling, inter-chain quality comparison, and prompt engineering. This enables the model to generate multiple complementary reasoning chains in one forward pass and perform intra-chain self-correction without external signals or post-hoc aggregation. Contribution/Results: DCoT is the first method to achieve CoT self-refinement *within* a single forward pass, departing from conventional multi-chain parallel generation followed by post-processing. Evaluated across 1.3B–70B models, DCoT consistently outperforms standard CoT baselines—especially on numerically intensive tasks with large state spaces. Human evaluation confirms an intra-chain improvement rate of 68.3%.

Enabling self-improvement in reasoning chains without external feedbackEnhancing LLM reasoning via within-inference CoT refinementImproving performance on tasks with diverse reasoning types

Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training

Nov 10, 2025
DB
Dake Bu
🏛️ City University of Hong Kong | Center for Advanced Intelligence Project, RIKEN | University of Sydney | CFAR and IHPC, Agency for Science, Technology and Research (A⋆STAR) | The University of Tokyo

The empirical benefits of curriculum learning in post-training inference for large language models (LLMs) lack principled theoretical justification. Method: We propose curriculum strategies based on incrementally increasing reasoning-chain depth or decrementally shortening prompt length, and introduce a state-conditioned autoregressive reasoning tree model. This framework enables curriculum-aware fine-tuning via reinforcement learning under outcome-only reward signals. Contribution: We provide the first theoretical proof that curriculum learning can overcome the exponential sample complexity barrier inherent in tree-structured reasoning—reducing it to polynomial order—and establish polynomial-cost scaling guarantees for test-time inference. Experiments demonstrate substantial improvements in reasoning accuracy, alongside significant reductions in sampling overhead and API query costs.

Establishing theoretical guarantees for curriculum learning's polynomial complexity benefitsModeling Chain-of-Thought generation as autoregressive reasoning trees for mathematical problemsUnderstanding why curriculum learning outperforms direct training in reasoning tasks

Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement

Apr 12, 2025
CL
Chengyuan Liu
🏛️ Zhejiang University | Alibaba Group

Existing chain-of-thought (CoT) methods exhibit limited generalization for logic-intensive domain-specific reasoning tasks—such as legal reasoning—that require deep, structured domain knowledge; meanwhile, Monte Carlo Tree Search (MCTS) lacks adaptability to professional reasoning contexts. Method: This paper pioneers the integration of MCTS into domain-specific reasoning via a step-level supervision framework: (i) a knowledge-guided MCTS search space constraint mechanism that explicitly aligns domain rules (e.g., statutory provisions, legal elements) with reasoning steps; and (ii) a learnable reflection-path preference model that enhances self-monitoring and correction of erroneous reasoning trajectories. Contribution/Results: Our approach achieves significant improvements over state-of-the-art baselines across multiple legal reasoning benchmarks. Further analysis reveals a strong positive correlation between fine-grained domain knowledge representation—such as statutory hierarchy and precise legal-element decomposition—and reasoning accuracy, establishing a novel paradigm for expert-level AI reasoning.

Enhancing legal-domain problem-solving using MCTSImproving reflection paths via preference optimizationOptimizing reasoning tasks with domain-specific knowledge

Generating Verifiable CoT from Execution-Traces

Nov 28, 2025
ST
Shailja Thakur
🏛️ IBM Research

Existing synthetic chain-of-thought (CoT) data often relies on teacher models to generate “plausible-sounding” yet unverifiable reasoning steps, leading language models to internalize logical hallucinations. To address this, we propose Execution-Traced CoT: a method that instruments code execution to capture ground-truth program traces and structurally maps them to natural-language reasoning steps—each strictly verifiable via observable program behavior. This enables bidirectional verifiability: forward (execution → reasoning) and backward (reasoning → execution). Using this approach, we construct high-fidelity training data and perform supervised fine-tuning of language models. On code reasoning benchmarks, our method improves prediction accuracy by up to 30 percentage points (output) and 28 percentage points (input), while substantially enhancing logical consistency and trustworthiness in both code generation and explanation.

Addresses logical flaws in synthetic training dataGenerates verifiable reasoning from execution tracesImproves code reasoning and generation tasks

On the Impact of Fine-Tuning on Chain-of-Thought Reasoning

Nov 22, 2024
EL
Elita Lobo
🏛️ University of Massachusetts | University of Virginia | Harvard University

This work investigates how fine-tuning affects the chain-of-thought (CoT) reasoning capabilities of large language models (LLMs), particularly its detrimental impact on reasoning faithfulness and generalization. Through controlled, comparative experiments across four standard CoT benchmarks, we systematically evaluate prevalent fine-tuning paradigms—including supervised fine-tuning (SFT), RLHF, and Q-LoRA—and quantitatively demonstrate, for the first time, that while fine-tuning improves task accuracy, it consistently degrades CoT faithfulness by a significant margin. To diagnose this phenomenon, we introduce a fine-grained faithfulness evaluation framework and empirically establish that fine-tuning induces systematic shifts in internal reasoning mechanisms. Our core contribution is the first causal identification of fine-tuning as a driver of CoT faithfulness degradation, accompanied by a reproducible diagnostic toolkit—laying foundational groundwork for developing trustworthy, interpretable LLM reasoning systems.

Effect of fine-tuning on Chain-of-Thought performanceFaithfulness changes in CoT reasoning post-fine-tuningImpact of fine-tuning on LLM reasoning capabilities

Latest Papers

What's happening recently
View more

This study investigates whether chain-of-thought (CoT) reasoning traces faithfully reflect a model’s actual internal decision-making process, thereby questioning their reliability as a supervisory and auditing mechanism. To this end, the authors propose a step-level Detect-Classify-Compare framework, integrating multidimensional validation techniques—including answer-commitment agents, Patchscopes, tuned-lens probes, causal ablation, truncation experiments, and donor contamination tests. Experiments across nine models and seven reasoning benchmarks reveal that, on average, only 61.9% of CoT steps align with the model’s internal computations; in 58% of misaligned cases, models generate redundant “reasoning” after the answer has already been determined—a phenomenon termed “hallucinated continuation.” Notably, stronger CoT performance correlates with lower temporal fidelity. This work provides the first systematic evidence of a fundamental disconnect between CoT traces and genuine reasoning dynamics, challenging the core assumption that CoT serves as a faithful reasoning log.

answer commitmentChain-of-Thoughtmodel interpretability

This study investigates how the source of chain-of-thought (CoT) data used in supervised fine-tuning affects model generalization, uncovering a paradox wherein low training loss coincides with poor generalization. By comparing CoT trajectories generated by DeepSeek-R1-0528 and gpt-oss-120b, the work identifies two distinct reasoning patterns—convergent deductive and divergent multi-branch reasoning—as key drivers of generalization differences. Building on this insight, the authors propose a trajectory filtering strategy based on branch frequency, which yields consistent performance gains across five reasoning benchmarks, including AIME25 and BeyondAIME, improving average accuracy by 3.6% (up to 5.5%) and substantially enhancing the model’s reasoning generalization capability.

Chain-of-ThoughtGeneralizationLarge Language Models

This work investigates the trade-off between performance and inference cost when large language models are post-trained on compressed reasoning data, a mechanism that remains poorly understood. The study introduces the first taxonomy of compressed chain-of-thought (CoT) reasoning, categorizing it into Explicit, Composed, and Implicit types. Through synthetic compositional reasoning tasks, the authors systematically analyze the effects of compression granularity and data scale using supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and multi-model ablation studies. Key findings reveal that coarse-grained CoT requires larger datasets for compensation, Composed CoT benefits from repeated training, and Implicit CoT is prone to memorization. Moreover, RLVR effectively decouples compression steps learned during SFT, and unidirectional CoT ordering demonstrates superior generalization on long-sequence tasks.

Chain-of-ThoughtCompressed ReasoningPost-Training

This work addresses the performance degradation of student models in supervised fine-tuning caused by style distribution mismatch when using synthetic data generated by strong teacher models. To mitigate this issue, the authors propose TESSY, a novel framework that introduces a teacher–student collaborative mechanism into the data synthesis process. TESSY alternately generates style and content tokens, effectively decoupling and then fusing the teacher’s advanced reasoning capabilities with the student’s linguistic style to produce training data that leverages the strengths of both. Experimental results demonstrate that, when applied to Qwen3-8B, TESSY improves performance by 11.25% on LiveCodeBench-Pro and 6.68% on OJBench compared to conventional teacher-generated data, successfully balancing reasoning ability and stylistic consistency.

reasoning modelstylistic divergencesupervised fine-tuning

This work addresses the frequent lack of faithfulness in chain-of-thought (CoT) reasoning generated by large language models, where answers are often derived by bypassing the intended reasoning process. Framing CoT causally as Z→X→Y (instruction → reasoning chain → answer), the authors propose the CASE framework: during training, counterfactual data and a selective loss function strengthen the causal influence of the reasoning chain on the final answer; during inference, attention masking blocks the direct path from instruction to answer. Integrating causal alignment with structural constraints, CASE achieves an average 37% relative improvement in CoT faithfulness across three models and four benchmarks, significantly enhancing interpretability, reliability, and cross-dataset generalization while maintaining competitive accuracy.

Causal ReasoningChain-of-ThoughtFaithfulness

Hot Scholars

ZL

Zhejian Lai

Master student of Nanjing University
自然语言处理
CW

Chengyu Wang

Alibaba Group
Natural Language ProcessingLarge Language ModelMulti-modal Learning
JW

Jiapeng Wang

South China University of Technology
document understandingvisual information extractionmulti-modal learningCLIP
SH

Shujian Huang

School of Computer Science, Nanjing University
Natural Language ProcessingMachine TranslationMultilingualismLarge Language Models
JT

Jie Tang

UW Madison
Computed Tomography