Score
Design, implement, and apply instruction-tuning procedures that initialize models to produce chain-of-thought (CoT) step-by-step explanations, including techniques to "cold-start" CoT behavior when few or no CoT labels are available (for example via synthetic CoT generation, prompt formats, or specialized loss terms). Build the accompanying datasets, training protocols, and evaluation methods to elicit and measure CoT instruction following and to assess how that initialization affects downstream reasoning performance.
Existing chain-of-thought (CoT) fine-tuning research predominantly focuses on technical implementation, lacking systematic analysis grounded in human cognitive mechanisms. This work bridges that gap by introducing, for the first time, a cognitive-dimension classification framework guided by de Bono’s “Six Thinking Hats” theory—systematically categorizing and reorganizing CoT fine-tuning methods along core human reasoning processes: planning, divergent thinking, intuitive judgment, and reflection. Methodologically, we integrate supervised and reinforcement fine-tuning, explicitly modeling CoT data according to empirically grounded human reasoning patterns. Empirically, we conduct comprehensive evaluations across mainstream benchmarks and model architectures, and release a continuously updated GitHub repository with curated resources. Our study fills a critical void at the intersection of CoT fine-tuning and cognitive science, establishing a scalable theoretical framework and practical paradigm for endowing large language models with human-like reasoning capabilities.
This work identifies a counterintuitive degradation in instruction-following accuracy when large language models (LLMs) are explicitly prompted to perform chain-of-thought (CoT) reasoning. Through systematic evaluation of 15 state-of-the-art models on IFEval and ComplexBench, we find that CoT often diverges from critical instruction constraints, leading to failure. To quantify this phenomenon, we propose *constraint attention*—a novel metric measuring the attenuation of constraint focus during reasoning. We further introduce four deployable mitigation strategies; among them, classifier-guided selective reasoning achieves the best performance: it triggers CoT only when necessary, yielding an average +18.7% accuracy gain on IFEval. Our results demonstrate that robust instruction following hinges less on universal, end-to-end CoT and more on *constraint-aware selective reasoning*—a paradigm prioritizing constraint fidelity over exhaustive inference.
This work investigates the mechanistic underpinnings of explicit chain-of-thought (CoT) training in enhancing the reasoning generalization capabilities of large language models—specifically addressing (i) the advantages of CoT training and (ii) its intrinsic mechanisms. Method: We design controlled data distributions and a two-hop factual reasoning task, integrating circuit analysis with quantitative generalization evaluation. Contribution/Results: We first demonstrate that CoT training simultaneously improves both in-distribution (ID) and out-of-distribution (OOD) generalization while accelerating convergence. Second, models acquire systematic reasoning abilities even from noisy CoT demonstrations. Third, generalization unfolds via a multi-stage circuit evolution strictly aligned with training steps. Finally, we identify the data ratio λ and pattern structure as critical levers governing generalization behavior. These findings provide interpretable, quantifiable empirical evidence for the mechanistic foundations and generalization boundaries of CoT training.
This work addresses two key challenges in synthetic instruction data construction: the difficulty of generating high-quality data and the trade-off between reasoning-intensive and non-reasoning tasks. To this end, we propose CoT-Self-Instruct—a novel framework that unifies chain-of-thought (CoT) reasoning with the Self-Instruct paradigm. It guides large language models to perform structured reasoning planning on seed tasks, enabling controllable complexity and semantic richness in generated instructions, and incorporates automated quality filtering metrics for dynamic data curation. Empirically, CoT-Self-Instruct consistently improves performance across both reasoning benchmarks (e.g., MATH500, AMC23) and instruction-following benchmarks (e.g., AlpacaEval 2.0, Arena-Hard), significantly outperforming prior synthetic datasets (e.g., s1k, OpenMathReasoning) as well as human-authored instruction data. These results validate the effectiveness and generalizability of reasoning-guided synthetic data generation for instruction tuning.
This work investigates the generalization capability of chain-of-thought (CoT) prompting in large language models (LLMs) for reasoning, focusing on the canonical planning domain Blocksworld. Method: We conduct a systematic empirical analysis using two state-of-the-art LLMs on controlled-complexity Blocksworld tasks and scalable CoT benchmark variants. Contribution/Results: We find that CoT performance critically depends on strict structural alignment—e.g., stack height—between exemplars and queries, exhibiting negligible generalization across problem complexity or syntactic form. Its gains stem from problem-specific pattern matching rather than acquisition of general algorithms. This study provides the first evidence of a fundamental generalization bottleneck for CoT in classical planning and quantifies a significant trade-off between CoT efficacy and the human effort required to engineer high-quality reasoning traces. These findings challenge the prevailing hypothesis that CoT enables implicit algorithm learning.
This study systematically investigates the task boundaries of chain-of-thought (CoT) prompting for enhancing large language model (LLM) performance. Through a quantitative meta-analysis and controlled experiments across 14 models and 20 datasets, we find CoT gains are highly concentrated in mathematical and symbolic reasoning tasks (+12.3% average improvement), yet negligible in commonsense reasoning and language understanding (+0.8%). We propose that CoT’s core mechanism is augmenting symbolic execution—not general-purpose reasoning—and demonstrate that its efficacy is strongly predicted by symbol-triggered behaviors (e.g., equality signs). A planning-execution decoupling analysis further reveals inherent computational paradigm limitations. Consequently, we advocate selective CoT activation to balance performance gains against inference cost, and call for novel intermediate computation architectures integrating explicit symbolic solvers—empirically shown to substantially outperform CoT.
This study investigates whether chain-of-thought (CoT) reasoning faithfully reflects the true decision-making process of large language models. By training linear probes on residual stream activations preceding CoT generation to predict final answers—and validating causal influence through activation interventions—the work provides the first mechanistic evidence that models typically commit to an answer before generating the CoT. Experimental results show that probes achieve AUC scores up to 0.9 across most tasks, and targeted activation steering flips the model’s answer in over 50% of samples, substantially outperforming baselines. Furthermore, the analysis reveals that when models hold incorrect beliefs, post-hoc CoT reasoning often leads to characteristic failure modes such as non-entailment or hallucination.
This work addresses the significant yet poorly understood performance variations of Chain-of-Thought (CoT) reasoning across different tasks by providing the first theoretical framework for its step-by-step inference process. The authors model CoT as a Markov chain and propose that its effectiveness hinges on the consistency of the transition kernels between reasoning steps, while also quantifying how noise in intermediate steps degrades performance. Through rigorous theoretical analysis, they prove that consistent transition kernels substantially reduce sample complexity. To validate these predictions, they construct synthetic benchmark experiments that align with their theoretical findings and offer a principled explanation for the observed disparities in CoT’s empirical success across real-world tasks. This study thus establishes a novel theoretical lens and analytical framework for understanding and improving CoT reasoning.
This study investigates whether chain-of-thought (CoT) training in large language model agents genuinely enhances reasoning capabilities or merely improves the ability to directly predict actions from prompts. Through comparative action prediction analyses, evaluation of training checkpoints, and selective masking of action tokens, the authors find that the primary benefit of CoT training stems from improved prompt-action alignment rather than strengthened internal reasoning. Moreover, as training progresses, models increasingly rely less on CoT for action correction. Building on these insights, the work proposes an intervention strategy that selectively masks action supervision signals during training, which effectively boosts out-of-distribution generalization performance.
This work addresses the frequent lack of faithfulness in chain-of-thought (CoT) reasoning generated by large language models, where answers are often derived by bypassing the intended reasoning process. Framing CoT causally as Z→X→Y (instruction → reasoning chain → answer), the authors propose the CASE framework: during training, counterfactual data and a selective loss function strengthen the causal influence of the reasoning chain on the final answer; during inference, attention masking blocks the direct path from instruction to answer. Integrating causal alignment with structural constraints, CASE achieves an average 37% relative improvement in CoT faithfulness across three models and four benchmarks, significantly enhancing interpretability, reliability, and cross-dataset generalization while maintaining competitive accuracy.
The underlying mechanisms by which Chain-of-Thought (CoT) prompting enhances large language models’ (LLMs’) code generation performance remain poorly understood. Method: This study systematically investigates CoT’s efficacy through an information-theoretic lens—specifically, conditional mutual information (I(Y;C|X))—across a multi-scale model spectrum (7B–480B), six Python and twelve multilingual code-generation benchmarks, and complexity-stratified evaluation. Contribution/Results: We quantitatively establish that CoT effectiveness critically depends on programming language, model scale, and reasoning quality—not merely template structure. Structured CoT yields average Pass@1 improvements of 5–12% over zero-shot CoT, outperforming unstructured variants. Crucially, reasoning fidelity proves more decisive than syntactic formatting; low-quality CoT degrades performance. These findings provide empirically grounded, actionable guidelines for selecting optimal CoT strategies across model sizes and programming languages.