🤖 AI Summary
Large language models (LLMs) lack intrinsic self-refinement capability within a single forward pass, limiting their ability to correct reasoning errors without external feedback or parallel generation.
Method: We propose Diversified Chain-of-Thought (DCoT) fine-tuning—a supervised fine-tuning paradigm that constructs structured DCoT datasets by integrating diversity-aware sampling, inter-chain quality comparison, and prompt engineering. This enables the model to generate multiple complementary reasoning chains in one forward pass and perform intra-chain self-correction without external signals or post-hoc aggregation.
Contribution/Results: DCoT is the first method to achieve CoT self-refinement *within* a single forward pass, departing from conventional multi-chain parallel generation followed by post-processing. Evaluated across 1.3B–70B models, DCoT consistently outperforms standard CoT baselines—especially on numerically intensive tasks with large state spaces. Human evaluation confirms an intra-chain improvement rate of 68.3%.
📝 Abstract
Requiring a large language model (LLM) to generate intermediary reasoning steps, known as Chain of Thought (CoT), has been shown to be an effective way of boosting performance. Previous approaches have focused on generating multiple independent CoTs, combining them through ensembling or other post-hoc strategies to enhance reasoning. In this work, we introduce a novel approach where LLMs are fine-tuned to generate a sequence of Diverse Chains of Thought (DCoT) within a single inference step, which is fundamentally different from prior work that primarily operate on parallel CoT generations. DCoT allows LLMs to gain the ability to perform within-inference refinement of reasoning chains without requiring external feedback. Through a rigorous set of experiments spanning a wide range of tasks that require various reasoning types, we show that fine-tuning on DCoT improves performance over the CoT baseline across model families and scales (1.3B to 70B). These improvements are particularly impactful for tasks with a large result state space, such as those involving numeric answers. Our work is also significant because both quantitative analyses and manual evaluations reveal the observed gains stem from the models' ability to refine an initial reasoning chain by generating a second, improved chain within the same inference step, demonstrating previously elusive self-improvement. Our code and data are publicly available at https://github.com/UKPLab/acl2025-diverse-cot.