Score
Designs and builds training pipelines, datasets, and compact student models that use chain-of-thought (CoT) or rationale traces from one or more teacher models as supervision signals — including generating weak CoT rationales, producing multi-teacher fine-grained labels, augmenting data with rationales, blending CoT supervision losses, and transferring CoT knowledge across architectures. Analyzes and evaluates the fidelity, stability and generalization of distilled reasoning traces (e.g., divergence over traces, recovery after pruning), and develops methods for weak supervision, rationale distillation, and rationale-based transfer to improve the reasoning behavior of smaller or constrained models.
Existing chain-of-thought (CoT) fine-tuning research predominantly focuses on technical implementation, lacking systematic analysis grounded in human cognitive mechanisms. This work bridges that gap by introducing, for the first time, a cognitive-dimension classification framework guided by de Bono’s “Six Thinking Hats” theory—systematically categorizing and reorganizing CoT fine-tuning methods along core human reasoning processes: planning, divergent thinking, intuitive judgment, and reflection. Methodologically, we integrate supervised and reinforcement fine-tuning, explicitly modeling CoT data according to empirically grounded human reasoning patterns. Empirically, we conduct comprehensive evaluations across mainstream benchmarks and model architectures, and release a continuously updated GitHub repository with curated resources. Our study fills a critical void at the intersection of CoT fine-tuning and cognitive science, establishing a scalable theoretical framework and practical paradigm for endowing large language models with human-like reasoning capabilities.
This work investigates efficient distillation of large language models’ (LLMs) chain-of-thought (CoT) reasoning capabilities into small language models (SLMs), balancing computational efficiency and reasoning performance. We conduct large-scale controlled experiments across seven mathematical and commonsense reasoning benchmarks, systematically varying four teacher LLMs, seven student architectures, CoT granularity levels, CoT formatting strategies, and teacher selection criteria. Key findings reveal: (i) a non-monotonic granularity effect in CoT distillation—neither finest nor coarsest granularities yield optimal performance; (ii) minimal impact of CoT formatting on student outcomes; and (iii) no positive correlation between teacher model strength and student performance, highlighting the need to balance teacher diversity and reasoning complexity. Based on these insights, we propose a student-adaptive CoT distillation strategy that significantly improves SLM generalization and stability across multi-task reasoning. All code and datasets are publicly released.
This work addresses the limitations of existing chain-of-thought (CoT) distillation methods, which rely on a single teacher model and are thus prone to its capability biases and catastrophic forgetting, hindering the full realization of student models’ reasoning potential. To overcome this, we propose COMPACT, a multi-teacher collaborative distillation framework that dynamically integrates supervision signals compatible with the student’s evolving capabilities. COMPACT introduces a multidimensional compatibility assessment mechanism: it filters erroneous reasoning paths via graph-based consensus, identifies informative teaching moments through mutual information, and mitigates negative transfer by modulating loss difficulty. By integrating graph-structured analysis, mutual information estimation, and dynamic gradient weighting, COMPACT constructs a compatibility-aware distillation system. Experiments demonstrate that COMPACT achieves state-of-the-art performance across multiple reasoning benchmarks, effectively enhancing small language models’ reasoning abilities while alleviating catastrophic forgetting and preserving their original knowledge.
This work investigates whether long-chain reasoning capabilities can be effectively elicited in base language models using only a small number of high-quality, human-authored chain-of-thought (CoT) examples or lightweight fine-tuning—without resorting to reinforcement learning or large-model distillation. Method: We propose a synergistic approach integrating prompt engineering, multi-round structured editing, and parameter-efficient fine-tuning, leveraging merely 20 high-precision CoT samples—generated by advanced reasoning models and rigorously validated by human experts—to optimize Qwen2.5-32B. Contribution/Results: Our method yields substantial improvements in mathematical and logical reasoning performance; the fine-tuned model surpasses the larger Qwen2.5-Math-72B-Instruct across multiple benchmarks. Crucially, we empirically demonstrate that carefully curated, human-annotated CoT data exhibits exceptional transfer efficacy for reasoning capability, establishing a cost-effective paradigm for unlocking latent reasoning potential in foundational language models.
Existing Chain-of-Thought (CoT) distillation methods rely excessively on large-scale rationale datasets while neglecting rationale quality, often transferring erroneous or low-quality reasoning paths to student models. Method: We propose MoRSD, a model-oriented rationale selection framework that introduces the first multi-dimensional Rationale Difficulty metric—incorporating accuracy, diversity, and difficulty—and a student-model feedback-driven self-guided filtering mechanism to dynamically identify high-value reasoning paths. Contribution/Results: Experiments across seven multiple-choice benchmarks demonstrate that MoRSD achieves an average performance gain of 4.6% using only ~30% of rationales—significantly outperforming full-rationale distillation. This work establishes the critical role of “few but high-quality” reasoning samples in knowledge transfer to compact models, offering a novel paradigm for efficient and robust CoT distillation.
Existing continuous chain-of-thought (Continuous CoT) methods rely on slow autoregressive generation and suffer significant performance degradation on tasks requiring long reasoning trajectories. This work proposes C-MTP, a novel approach that, for the first time, directly supervises hidden states using the mean of corresponding chain-of-thought embeddings, thereby employing embedding averages as supervision signals to simplify training and eliminate the need for autoregressive decoding. The method outperforms existing direct supervision approaches on short reasoning tasks and matches the performance of indirect supervision methods. However, on long reasoning trajectories spanning hundreds of tokens, all current methods—including C-MTP—experience a performance drop of approximately 65%, revealing a fundamental limitation of contemporary Continuous CoT frameworks in long-horizon reasoning.
This study investigates the impact of transferring chain-of-thought (CoT) reasoning from one large language model to another on the recipient model’s inference and generation mechanisms. By establishing a provider–recipient framework and employing techniques such as CoT prefix truncation, forced-answer versus free-generation comparisons, and multi-model, multi-benchmark evaluation, the work reveals that CoT transfer operates through multiple pathways—including answer extraction, reasoning scaffolding, and dependence on the recipient model’s inherent capabilities—rather than a single uniform mechanism. The authors propose using answer consistency in the absence of ground-truth labels as an early stopping signal for reasoning and demonstrate across benchmarks including AIME, MMLU-Pro, and ZebraLogic that partial CoT prompts can effectively guide subsequent reasoning and enhance performance.
This work addresses the challenge of enhancing large language models’ reasoning capabilities in settings with scarce supervision by leveraging unlabeled questions. The authors propose Semi-CoT, a novel framework that extends chain-of-thought (CoT) reasoning from an inference-time prompting strategy to a semi-supervised learning signal. Specifically, the model generates multiple pseudo-reasoning chains for unlabeled questions and employs answer semantic entropy to automatically select high-confidence samples for self-training. Evaluated on AQuA, SVAMP, GSM8K, and MultiArith benchmarks, the approach achieves pseudo-answer accuracy ranging from 91.36% to 100% and yields consistent, albeit modest, performance gains on several tasks, thereby demonstrating the efficacy of semantic entropy–guided semi-supervised CoT learning.
This work addresses the challenge in distilling reasoning capabilities from large language models to smaller student models, where inconsistent reasoning path structures generated by the teacher for semantically similar questions hinder effective learning. To resolve this, the authors propose a Dynamic Reasoning Path Compression (D-RPC) distillation mechanism that constructs and maintains a high-order reasoning path repository, guiding the teacher during training to produce reasoning trajectories that are both structurally consistent and diverse in coverage, based on the most relevant stored paths. The approach uniquely integrates PAC-Bayes generalization bound analysis into reasoning distillation, achieving a theoretically optimal trade-off between path repository size and coverage capacity. Experiments across five mathematical and commonsense reasoning benchmarks demonstrate that two distinct student models consistently outperform existing baselines while requiring fewer generated tokens than template-intensive methods.
This work addresses the computational challenge of learning from multiple annotators who provide correct yet stylistically diverse chain-of-thought (CoT) rationales. The study is the first to formally characterize the computational complexity of this setting and introduces an efficient active learning algorithm that overcomes the limitations of passive learning. The proposed method requires only a fixed number of CoT examples per annotator, O(log(1/ε) log log(1/ε)) annotators, and Õ(1/ε) supervisory labels on final answers to achieve ε-accurate learning. Notably, its annotation efficiency is independent of the target accuracy ε, substantially enhancing the scalability of learning under heterogeneous CoT supervision.