chain-of-thought cold-start tuning

Design, implement, and apply instruction-tuning procedures that initialize models to produce chain-of-thought (CoT) step-by-step explanations, including techniques to "cold-start" CoT behavior when few or no CoT labels are available (for example via synthetic CoT generation, prompt formats, or specialized loss terms). Build the accompanying datasets, training protocols, and evaluation methods to elicit and measure CoT instruction following and to assess how that initialization affects downstream reasoning performance.

chain-of-thoughtcold-starttuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

May 16, 2025
XL
Xiaomin Li
🏛️ Harvard University | Amazon | NYU

This work identifies a counterintuitive degradation in instruction-following accuracy when large language models (LLMs) are explicitly prompted to perform chain-of-thought (CoT) reasoning. Through systematic evaluation of 15 state-of-the-art models on IFEval and ComplexBench, we find that CoT often diverges from critical instruction constraints, leading to failure. To quantify this phenomenon, we propose *constraint attention*—a novel metric measuring the attenuation of constraint focus during reasoning. We further introduce four deployable mitigation strategies; among them, classifier-guided selective reasoning achieves the best performance: it triggers CoT only when necessary, yielding an average +18.7% accuracy gain on IFEval. Our results demonstrate that robust instruction following hinges less on universal, end-to-end CoT and more on *constraint-aware selective reasoning*—a paradigm prioritizing constraint fidelity over exhaustive inference.

Explicit CoT reasoning reduces instruction-following accuracy in LLMsReasoning diverts attention from instruction-relevant tokens in modelsSelective reasoning strategies mitigate performance drops in LLMs

Unveiling the Mechanisms of Explicit CoT Training: How Chain-of-Thought Enhances Reasoning Generalization

Feb 07, 2025
XY
Xinhao Yao
🏛️ Renmin University of China | Tianjin University of Science and Technology

This work investigates the mechanistic underpinnings of explicit chain-of-thought (CoT) training in enhancing the reasoning generalization capabilities of large language models—specifically addressing (i) the advantages of CoT training and (ii) its intrinsic mechanisms. Method: We design controlled data distributions and a two-hop factual reasoning task, integrating circuit analysis with quantitative generalization evaluation. Contribution/Results: We first demonstrate that CoT training simultaneously improves both in-distribution (ID) and out-of-distribution (OOD) generalization while accelerating convergence. Second, models acquire systematic reasoning abilities even from noisy CoT demonstrations. Third, generalization unfolds via a multi-stage circuit evolution strictly aligned with training steps. Finally, we identify the data ratio λ and pattern structure as critical levers governing generalization behavior. These findings provide interpretable, quantifiable empirical evidence for the mechanistic foundations and generalization boundaries of CoT training.

Enhancing reasoning generalization in LLMsImproving systematic generalization with CoTUnderstanding CoT training mechanisms

This work addresses two key challenges in synthetic instruction data construction: the difficulty of generating high-quality data and the trade-off between reasoning-intensive and non-reasoning tasks. To this end, we propose CoT-Self-Instruct—a novel framework that unifies chain-of-thought (CoT) reasoning with the Self-Instruct paradigm. It guides large language models to perform structured reasoning planning on seed tasks, enabling controllable complexity and semantic richness in generated instructions, and incorporates automated quality filtering metrics for dynamic data curation. Empirically, CoT-Self-Instruct consistently improves performance across both reasoning benchmarks (e.g., MATH500, AMC23) and instruction-following benchmarks (e.g., AlpacaEval 2.0, Arena-Hard), significantly outperforming prior synthetic datasets (e.g., s1k, OpenMathReasoning) as well as human-authored instruction data. These results validate the effectiveness and generalizability of reasoning-guided synthetic data generation for instruction tuning.

Enhancing performance on verifiable and non-verifiable tasksGenerating high-quality synthetic prompts for reasoning tasksImproving LLM training with automatic data filtering

Chain of Thoughtlessness? An Analysis of CoT in Planning

May 08, 2024
KS
Kaya Stechly
🏛️ Arizona State University

This work investigates the generalization capability of chain-of-thought (CoT) prompting in large language models (LLMs) for reasoning, focusing on the canonical planning domain Blocksworld. Method: We conduct a systematic empirical analysis using two state-of-the-art LLMs on controlled-complexity Blocksworld tasks and scalable CoT benchmark variants. Contribution/Results: We find that CoT performance critically depends on strict structural alignment—e.g., stack height—between exemplars and queries, exhibiting negligible generalization across problem complexity or syntactic form. Its gains stem from problem-specific pattern matching rather than acquisition of general algorithms. This study provides the first evidence of a fundamental generalization bottleneck for CoT in classical planning and quantifies a significant trade-off between CoT efficacy and the human effort required to engineer high-quality reasoning traces. These findings challenge the prevailing hypothesis that CoT enables implicit algorithm learning.

Chain of thought prompts require highly specific examples for improvement.LLM performance on reasoning problems lacks generalization.Performance gains from CoT depend on problem-specific prompt engineering.

To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Sep 18, 2024
ZS
Zayne Sprague
🏛️ The University of Texas at Austin | Johns Hopkins University | Princeton University

This study systematically investigates the task boundaries of chain-of-thought (CoT) prompting for enhancing large language model (LLM) performance. Through a quantitative meta-analysis and controlled experiments across 14 models and 20 datasets, we find CoT gains are highly concentrated in mathematical and symbolic reasoning tasks (+12.3% average improvement), yet negligible in commonsense reasoning and language understanding (+0.8%). We propose that CoT’s core mechanism is augmenting symbolic execution—not general-purpose reasoning—and demonstrate that its efficacy is strongly predicted by symbol-triggered behaviors (e.g., equality signs). A planning-execution decoupling analysis further reveals inherent computational paradigm limitations. Consequently, we advocate selective CoT activation to balance performance gains against inference cost, and call for novel intermediate computation architectures integrating explicit symbolic solvers—empirically shown to substantially outperform CoT.

Evaluates when Chain-of-Thought (CoT) improves LLM task performanceIdentifies math and logic tasks as primary beneficiaries of CoTSuggests moving beyond prompt-based CoT for broader LLM applications

Latest Papers

What's happening recently
View more

This study investigates whether chain-of-thought (CoT) reasoning faithfully reflects the true decision-making process of large language models. By training linear probes on residual stream activations preceding CoT generation to predict final answers—and validating causal influence through activation interventions—the work provides the first mechanistic evidence that models typically commit to an answer before generating the CoT. Experimental results show that probes achieve AUC scores up to 0.9 across most tasks, and targeted activation steering flips the model’s answer in over 50% of samples, substantially outperforming baselines. Furthermore, the analysis reveals that when models hold incorrect beliefs, post-hoc CoT reasoning often leads to characteristic failure modes such as non-entailment or hallucination.

chain-of-thoughtfaithfulnessinterpretability

This work addresses the significant yet poorly understood performance variations of Chain-of-Thought (CoT) reasoning across different tasks by providing the first theoretical framework for its step-by-step inference process. The authors model CoT as a Markov chain and propose that its effectiveness hinges on the consistency of the transition kernels between reasoning steps, while also quantifying how noise in intermediate steps degrades performance. Through rigorous theoretical analysis, they prove that consistent transition kernels substantially reduce sample complexity. To validate these predictions, they construct synthetic benchmark experiments that align with their theoretical findings and offer a principled explanation for the observed disparities in CoT’s empirical success across real-world tasks. This study thus establishes a novel theoretical lens and analytical framework for understanding and improving CoT reasoning.

Chain-of-ThoughtMarkov chainreasoning

This study investigates whether chain-of-thought (CoT) training in large language model agents genuinely enhances reasoning capabilities or merely improves the ability to directly predict actions from prompts. Through comparative action prediction analyses, evaluation of training checkpoints, and selective masking of action tokens, the authors find that the primary benefit of CoT training stems from improved prompt-action alignment rather than strengthened internal reasoning. Moreover, as training progresses, models increasingly rely less on CoT for action correction. Building on these insights, the work proposes an intervention strategy that selectively masks action supervision signals during training, which effectively boosts out-of-distribution generalization performance.

action predictionChain-of-Thoughtlanguage-model agents

This work addresses the frequent lack of faithfulness in chain-of-thought (CoT) reasoning generated by large language models, where answers are often derived by bypassing the intended reasoning process. Framing CoT causally as Z→X→Y (instruction → reasoning chain → answer), the authors propose the CASE framework: during training, counterfactual data and a selective loss function strengthen the causal influence of the reasoning chain on the final answer; during inference, attention masking blocks the direct path from instruction to answer. Integrating causal alignment with structural constraints, CASE achieves an average 37% relative improvement in CoT faithfulness across three models and four benchmarks, significantly enhancing interpretability, reliability, and cross-dataset generalization while maintaining competitive accuracy.

Causal ReasoningChain-of-ThoughtFaithfulness

The underlying mechanisms by which Chain-of-Thought (CoT) prompting enhances large language models’ (LLMs’) code generation performance remain poorly understood. Method: This study systematically investigates CoT’s efficacy through an information-theoretic lens—specifically, conditional mutual information (I(Y;C|X))—across a multi-scale model spectrum (7B–480B), six Python and twelve multilingual code-generation benchmarks, and complexity-stratified evaluation. Contribution/Results: We quantitatively establish that CoT effectiveness critically depends on programming language, model scale, and reasoning quality—not merely template structure. Structured CoT yields average Pass@1 improvements of 5–12% over zero-shot CoT, outperforming unstructured variants. Crucially, reasoning fidelity proves more decisive than syntactic formatting; low-quality CoT degrades performance. These findings provide empirically grounded, actionable guidelines for selecting optimal CoT strategies across model sizes and programming languages.

Analyzes how Chain-of-Thought prompting improves code generation in LLMs.Evaluates CoT effectiveness across models, languages, and benchmarks empirically.Investigates the role of reasoning quality and structured guidance in CoT.

Hot Scholars

JX

Jun Xu

Professor, Gaoling School of Artificial Intelligence, Renmin University of China
Information RetrievalLearning to RankSemantic Matching
JZ

Jiahao Zhao

Institute of automation, Chinese Academy of Sciences
LLM Alignment
CL

Chongxuan Li

Associate Professor, Renmin University of China
Machine LearningGenerative ModelsDeep Learning
ZS

Zhongxiang Sun

Renmin University of China
SearchRecommendationLLMLegal