Score
Designs and implements supervised fine‑tuning regimes and curricula that train models to generate explicit step‑by‑step reasoning traces (chain‑of‑thought), including long or visual CoT forms and code‑paired traces; builds or curates expert stepwise explanation datasets, selects progressive learning schedules, and fine‑tunes models to improve multi‑step problem solving, correctness, and the clarity/auditability of generated reasoning.
Existing chain-of-thought (CoT) fine-tuning research predominantly focuses on technical implementation, lacking systematic analysis grounded in human cognitive mechanisms. This work bridges that gap by introducing, for the first time, a cognitive-dimension classification framework guided by de Bono’s “Six Thinking Hats” theory—systematically categorizing and reorganizing CoT fine-tuning methods along core human reasoning processes: planning, divergent thinking, intuitive judgment, and reflection. Methodologically, we integrate supervised and reinforcement fine-tuning, explicitly modeling CoT data according to empirically grounded human reasoning patterns. Empirically, we conduct comprehensive evaluations across mainstream benchmarks and model architectures, and release a continuously updated GitHub repository with curated resources. Our study fills a critical void at the intersection of CoT fine-tuning and cognitive science, establishing a scalable theoretical framework and practical paradigm for endowing large language models with human-like reasoning capabilities.
Large language models (LLMs) lack intrinsic self-refinement capability within a single forward pass, limiting their ability to correct reasoning errors without external feedback or parallel generation. Method: We propose Diversified Chain-of-Thought (DCoT) fine-tuning—a supervised fine-tuning paradigm that constructs structured DCoT datasets by integrating diversity-aware sampling, inter-chain quality comparison, and prompt engineering. This enables the model to generate multiple complementary reasoning chains in one forward pass and perform intra-chain self-correction without external signals or post-hoc aggregation. Contribution/Results: DCoT is the first method to achieve CoT self-refinement *within* a single forward pass, departing from conventional multi-chain parallel generation followed by post-processing. Evaluated across 1.3B–70B models, DCoT consistently outperforms standard CoT baselines—especially on numerically intensive tasks with large state spaces. Human evaluation confirms an intra-chain improvement rate of 68.3%.
The empirical benefits of curriculum learning in post-training inference for large language models (LLMs) lack principled theoretical justification. Method: We propose curriculum strategies based on incrementally increasing reasoning-chain depth or decrementally shortening prompt length, and introduce a state-conditioned autoregressive reasoning tree model. This framework enables curriculum-aware fine-tuning via reinforcement learning under outcome-only reward signals. Contribution: We provide the first theoretical proof that curriculum learning can overcome the exponential sample complexity barrier inherent in tree-structured reasoning—reducing it to polynomial order—and establish polynomial-cost scaling guarantees for test-time inference. Experiments demonstrate substantial improvements in reasoning accuracy, alongside significant reductions in sampling overhead and API query costs.
Existing chain-of-thought (CoT) methods exhibit limited generalization for logic-intensive domain-specific reasoning tasks—such as legal reasoning—that require deep, structured domain knowledge; meanwhile, Monte Carlo Tree Search (MCTS) lacks adaptability to professional reasoning contexts. Method: This paper pioneers the integration of MCTS into domain-specific reasoning via a step-level supervision framework: (i) a knowledge-guided MCTS search space constraint mechanism that explicitly aligns domain rules (e.g., statutory provisions, legal elements) with reasoning steps; and (ii) a learnable reflection-path preference model that enhances self-monitoring and correction of erroneous reasoning trajectories. Contribution/Results: Our approach achieves significant improvements over state-of-the-art baselines across multiple legal reasoning benchmarks. Further analysis reveals a strong positive correlation between fine-grained domain knowledge representation—such as statutory hierarchy and precise legal-element decomposition—and reasoning accuracy, establishing a novel paradigm for expert-level AI reasoning.
Existing synthetic chain-of-thought (CoT) data often relies on teacher models to generate “plausible-sounding” yet unverifiable reasoning steps, leading language models to internalize logical hallucinations. To address this, we propose Execution-Traced CoT: a method that instruments code execution to capture ground-truth program traces and structurally maps them to natural-language reasoning steps—each strictly verifiable via observable program behavior. This enables bidirectional verifiability: forward (execution → reasoning) and backward (reasoning → execution). Using this approach, we construct high-fidelity training data and perform supervised fine-tuning of language models. On code reasoning benchmarks, our method improves prediction accuracy by up to 30 percentage points (output) and 28 percentage points (input), while substantially enhancing logical consistency and trustworthiness in both code generation and explanation.
This work investigates how fine-tuning affects the chain-of-thought (CoT) reasoning capabilities of large language models (LLMs), particularly its detrimental impact on reasoning faithfulness and generalization. Through controlled, comparative experiments across four standard CoT benchmarks, we systematically evaluate prevalent fine-tuning paradigms—including supervised fine-tuning (SFT), RLHF, and Q-LoRA—and quantitatively demonstrate, for the first time, that while fine-tuning improves task accuracy, it consistently degrades CoT faithfulness by a significant margin. To diagnose this phenomenon, we introduce a fine-grained faithfulness evaluation framework and empirically establish that fine-tuning induces systematic shifts in internal reasoning mechanisms. Our core contribution is the first causal identification of fine-tuning as a driver of CoT faithfulness degradation, accompanied by a reproducible diagnostic toolkit—laying foundational groundwork for developing trustworthy, interpretable LLM reasoning systems.
This study investigates whether chain-of-thought (CoT) reasoning traces faithfully reflect a model’s actual internal decision-making process, thereby questioning their reliability as a supervisory and auditing mechanism. To this end, the authors propose a step-level Detect-Classify-Compare framework, integrating multidimensional validation techniques—including answer-commitment agents, Patchscopes, tuned-lens probes, causal ablation, truncation experiments, and donor contamination tests. Experiments across nine models and seven reasoning benchmarks reveal that, on average, only 61.9% of CoT steps align with the model’s internal computations; in 58% of misaligned cases, models generate redundant “reasoning” after the answer has already been determined—a phenomenon termed “hallucinated continuation.” Notably, stronger CoT performance correlates with lower temporal fidelity. This work provides the first systematic evidence of a fundamental disconnect between CoT traces and genuine reasoning dynamics, challenging the core assumption that CoT serves as a faithful reasoning log.
This study investigates how the source of chain-of-thought (CoT) data used in supervised fine-tuning affects model generalization, uncovering a paradox wherein low training loss coincides with poor generalization. By comparing CoT trajectories generated by DeepSeek-R1-0528 and gpt-oss-120b, the work identifies two distinct reasoning patterns—convergent deductive and divergent multi-branch reasoning—as key drivers of generalization differences. Building on this insight, the authors propose a trajectory filtering strategy based on branch frequency, which yields consistent performance gains across five reasoning benchmarks, including AIME25 and BeyondAIME, improving average accuracy by 3.6% (up to 5.5%) and substantially enhancing the model’s reasoning generalization capability.
This work investigates the trade-off between performance and inference cost when large language models are post-trained on compressed reasoning data, a mechanism that remains poorly understood. The study introduces the first taxonomy of compressed chain-of-thought (CoT) reasoning, categorizing it into Explicit, Composed, and Implicit types. Through synthetic compositional reasoning tasks, the authors systematically analyze the effects of compression granularity and data scale using supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and multi-model ablation studies. Key findings reveal that coarse-grained CoT requires larger datasets for compensation, Composed CoT benefits from repeated training, and Implicit CoT is prone to memorization. Moreover, RLVR effectively decouples compression steps learned during SFT, and unidirectional CoT ordering demonstrates superior generalization on long-sequence tasks.
This work addresses the performance degradation of student models in supervised fine-tuning caused by style distribution mismatch when using synthetic data generated by strong teacher models. To mitigate this issue, the authors propose TESSY, a novel framework that introduces a teacher–student collaborative mechanism into the data synthesis process. TESSY alternately generates style and content tokens, effectively decoupling and then fusing the teacher’s advanced reasoning capabilities with the student’s linguistic style to produce training data that leverages the strengths of both. Experimental results demonstrate that, when applied to Qwen3-8B, TESSY improves performance by 11.25% on LiveCodeBench-Pro and 6.68% on OJBench compared to conventional teacher-generated data, successfully balancing reasoning ability and stylistic consistency.
This work addresses the frequent lack of faithfulness in chain-of-thought (CoT) reasoning generated by large language models, where answers are often derived by bypassing the intended reasoning process. Framing CoT causally as Z→X→Y (instruction → reasoning chain → answer), the authors propose the CASE framework: during training, counterfactual data and a selective loss function strengthen the causal influence of the reasoning chain on the final answer; during inference, attention masking blocks the direct path from instruction to answer. Integrating causal alignment with structural constraints, CASE achieves an average 37% relative improvement in CoT faithfulness across three models and four benchmarks, significantly enhancing interpretability, reliability, and cross-dataset generalization while maintaining competitive accuracy.