chain-of-thought distillation

Designs and builds training pipelines, datasets, and compact student models that use chain-of-thought (CoT) or rationale traces from one or more teacher models as supervision signals — including generating weak CoT rationales, producing multi-teacher fine-grained labels, augmenting data with rationales, blending CoT supervision losses, and transferring CoT knowledge across architectures. Analyzes and evaluates the fidelity, stability and generalization of distilled reasoning traces (e.g., divergence over traces, recovery after pruning), and develops methods for weak supervision, rationale distillation, and rationale-based transfer to improve the reasoning behavior of smaller or constrained models.

chain-of-thoughtdistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning

Feb 25, 2025
XC
Xinghao Chen
🏛️ The Hong Kong Polytechnic University | Eastern Institute of Technology | Saarland University | Meituan Inc.

This work investigates efficient distillation of large language models’ (LLMs) chain-of-thought (CoT) reasoning capabilities into small language models (SLMs), balancing computational efficiency and reasoning performance. We conduct large-scale controlled experiments across seven mathematical and commonsense reasoning benchmarks, systematically varying four teacher LLMs, seven student architectures, CoT granularity levels, CoT formatting strategies, and teacher selection criteria. Key findings reveal: (i) a non-monotonic granularity effect in CoT distillation—neither finest nor coarsest granularities yield optimal performance; (ii) minimal impact of CoT formatting on student outcomes; and (iii) no positive correlation between teacher model strength and student performance, highlighting the need to balance teacher diversity and reasoning complexity. Based on these insights, we propose a student-adaptive CoT distillation strategy that significantly improves SLM generalization and stability across multi-task reasoning. All code and datasets are publicly released.

Examine granularity, format, and teacher models.Identify factors for CoT distillation.Optimize CoT strategies for small language models.

This work addresses the limitations of existing chain-of-thought (CoT) distillation methods, which rely on a single teacher model and are thus prone to its capability biases and catastrophic forgetting, hindering the full realization of student models’ reasoning potential. To overcome this, we propose COMPACT, a multi-teacher collaborative distillation framework that dynamically integrates supervision signals compatible with the student’s evolving capabilities. COMPACT introduces a multidimensional compatibility assessment mechanism: it filters erroneous reasoning paths via graph-based consensus, identifies informative teaching moments through mutual information, and mitigates negative transfer by modulating loss difficulty. By integrating graph-structured analysis, mutual information estimation, and dynamic gradient weighting, COMPACT constructs a compatibility-aware distillation system. Experiments demonstrate that COMPACT achieves state-of-the-art performance across multiple reasoning benchmarks, effectively enhancing small language models’ reasoning abilities while alleviating catastrophic forgetting and preserving their original knowledge.

Catastrophic ForgettingChain-of-Thought DistillationMulti-Teacher Learning

Is Human-Written Data Enough? The Challenge of Teaching Reasoning to LLMs Without RL or Distillation

Jul 13, 2025
WD
Wei Du
🏛️ Nvidia Corporation | Institute for Artificial Intelligence Research and Development of Serbia | University of Novi Sad | AwesomeMath | Massachusetts Institute of Technology | University of Pennsylvania | University of Illinois Urbana-Champaign | Georgia Institute of Technology | University of Chicago

This work investigates whether long-chain reasoning capabilities can be effectively elicited in base language models using only a small number of high-quality, human-authored chain-of-thought (CoT) examples or lightweight fine-tuning—without resorting to reinforcement learning or large-model distillation. Method: We propose a synergistic approach integrating prompt engineering, multi-round structured editing, and parameter-efficient fine-tuning, leveraging merely 20 high-precision CoT samples—generated by advanced reasoning models and rigorously validated by human experts—to optimize Qwen2.5-32B. Contribution/Results: Our method yields substantial improvements in mathematical and logical reasoning performance; the fine-tuned model surpasses the larger Qwen2.5-Math-72B-Instruct across multiple benchmarks. Crucially, we empirically demonstrate that carefully curated, human-annotated CoT data exhibits exceptional transfer efficacy for reasoning capability, establishing a cost-effective paradigm for unlocking latent reasoning potential in foundational language models.

Can human-written or non-expert CoT data match reasoning model performance?Can long Chain-of-Thought reasoning be induced in base models without extensive tuning?Does minimal fine-tuning with few high-quality examples unlock strong reasoning capabilities?

Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales

Sep 27, 2025
JY
Jianzhi Yan
🏛️ Harbin Institute of Technology | Pengcheng Laboratory | Shaoguan Research Institute of Data Industry

Existing Chain-of-Thought (CoT) distillation methods rely excessively on large-scale rationale datasets while neglecting rationale quality, often transferring erroneous or low-quality reasoning paths to student models. Method: We propose MoRSD, a model-oriented rationale selection framework that introduces the first multi-dimensional Rationale Difficulty metric—incorporating accuracy, diversity, and difficulty—and a student-model feedback-driven self-guided filtering mechanism to dynamically identify high-value reasoning paths. Contribution/Results: Experiments across seven multiple-choice benchmarks demonstrate that MoRSD achieves an average performance gain of 4.6% using only ~30% of rationales—significantly outperforming full-rationale distillation. This work establishes the critical role of “few but high-quality” reasoning samples in knowledge transfer to compact models, offering a novel paradigm for efficient and robust CoT distillation.

Enhancing performance with fewer but better rationalesImproving reasoning in small language models via distillationSelecting high-quality rationales to avoid noisy information transfer

Existing continuous chain-of-thought (Continuous CoT) methods rely on slow autoregressive generation and suffer significant performance degradation on tasks requiring long reasoning trajectories. This work proposes C-MTP, a novel approach that, for the first time, directly supervises hidden states using the mean of corresponding chain-of-thought embeddings, thereby employing embedding averages as supervision signals to simplify training and eliminate the need for autoregressive decoding. The method outperforms existing direct supervision approaches on short reasoning tasks and matches the performance of indirect supervision methods. However, on long reasoning trajectories spanning hundreds of tokens, all current methods—including C-MTP—experience a performance drop of approximately 65%, revealing a fundamental limitation of contemporary Continuous CoT frameworks in long-horizon reasoning.

complex tasksContinuous Chain-of-Thoughtlong reasoning

Latest Papers

What's happening recently
View more

This study investigates the impact of transferring chain-of-thought (CoT) reasoning from one large language model to another on the recipient model’s inference and generation mechanisms. By establishing a provider–recipient framework and employing techniques such as CoT prefix truncation, forced-answer versus free-generation comparisons, and multi-model, multi-benchmark evaluation, the work reveals that CoT transfer operates through multiple pathways—including answer extraction, reasoning scaffolding, and dependence on the recipient model’s inherent capabilities—rather than a single uniform mechanism. The authors propose using answer consistency in the absence of ground-truth labels as an early stopping signal for reasoning and demonstrate across benchmarks including AIME, MMLU-Pro, and ZebraLogic that partial CoT prompts can effectively guide subsequent reasoning and enhance performance.

answer generationchain-of-thoughtcross-model

This work addresses the challenge of enhancing large language models’ reasoning capabilities in settings with scarce supervision by leveraging unlabeled questions. The authors propose Semi-CoT, a novel framework that extends chain-of-thought (CoT) reasoning from an inference-time prompting strategy to a semi-supervised learning signal. Specifically, the model generates multiple pseudo-reasoning chains for unlabeled questions and employs answer semantic entropy to automatically select high-confidence samples for self-training. Evaluated on AQuA, SVAMP, GSM8K, and MultiArith benchmarks, the approach achieves pseudo-answer accuracy ranging from 91.36% to 100% and yields consistent, albeit modest, performance gains on several tasks, thereby demonstrating the efficacy of semantic entropy–guided semi-supervised CoT learning.

Chain-of-ThoughtLarge Language ModelsPseudo-labeling

This work addresses the challenge in distilling reasoning capabilities from large language models to smaller student models, where inconsistent reasoning path structures generated by the teacher for semantically similar questions hinder effective learning. To resolve this, the authors propose a Dynamic Reasoning Path Compression (D-RPC) distillation mechanism that constructs and maintains a high-order reasoning path repository, guiding the teacher during training to produce reasoning trajectories that are both structurally consistent and diverse in coverage, based on the most relevant stored paths. The approach uniquely integrates PAC-Bayes generalization bound analysis into reasoning distillation, achieving a theoretically optimal trade-off between path repository size and coverage capacity. Experiments across five mathematical and commonsense reasoning benchmarks demonstrate that two distinct student models consistently outperform existing baselines while requiring fewer generated tokens than template-intensive methods.

language model compressionrationale consistencyreasoning distillation

This work addresses the computational challenge of learning from multiple annotators who provide correct yet stylistically diverse chain-of-thought (CoT) rationales. The study is the first to formally characterize the computational complexity of this setting and introduces an efficient active learning algorithm that overcomes the limitations of passive learning. The proposed method requires only a fixed number of CoT examples per annotator, O(log(1/ε) log log(1/ε)) annotators, and Õ(1/ε) supervisory labels on final answers to achieve ε-accurate learning. Notably, its annotation efficiency is independent of the target accuracy ε, substantially enhancing the scalability of learning under heterogeneous CoT supervision.

active learningChain-of-Thoughtcomputational learning

Hot Scholars

XG

Xin Gao

Shanghai AI Laboratory & SJTU
ML、NLP、LLM
LW

Lijun Wu

Shanghai AI Laboratory
MLLLMAI4Science
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence
AM

Amir M. Rahmani

Samueli Chair Professor of Integrative Health, Nursing, CS & EECS, University of California, Irvine
AI in HealthcareWearable ComputingDigital HealthHealth Informatics