Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs

📅 2024-07-03
📈 Citations: 10
✨ Influential: 1
📄 PDF
🤖 AI Summary
Large language models (LLMs) lack intrinsic self-refinement capability within a single forward pass, limiting their ability to correct reasoning errors without external feedback or parallel generation. Method: We propose Diversified Chain-of-Thought (DCoT) fine-tuning—a supervised fine-tuning paradigm that constructs structured DCoT datasets by integrating diversity-aware sampling, inter-chain quality comparison, and prompt engineering. This enables the model to generate multiple complementary reasoning chains in one forward pass and perform intra-chain self-correction without external signals or post-hoc aggregation. Contribution/Results: DCoT is the first method to achieve CoT self-refinement *within* a single forward pass, departing from conventional multi-chain parallel generation followed by post-processing. Evaluated across 1.3B–70B models, DCoT consistently outperforms standard CoT baselines—especially on numerically intensive tasks with large state spaces. Human evaluation confirms an intra-chain improvement rate of 68.3%.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Diffusion Models for Vision

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Requiring a large language model (LLM) to generate intermediary reasoning steps, known as Chain of Thought (CoT), has been shown to be an effective way of boosting performance. Previous approaches have focused on generating multiple independent CoTs, combining them through ensembling or other post-hoc strategies to enhance reasoning. In this work, we introduce a novel approach where LLMs are fine-tuned to generate a sequence of Diverse Chains of Thought (DCoT) within a single inference step, which is fundamentally different from prior work that primarily operate on parallel CoT generations. DCoT allows LLMs to gain the ability to perform within-inference refinement of reasoning chains without requiring external feedback. Through a rigorous set of experiments spanning a wide range of tasks that require various reasoning types, we show that fine-tuning on DCoT improves performance over the CoT baseline across model families and scales (1.3B to 70B). These improvements are particularly impactful for tasks with a large result state space, such as those involving numeric answers. Our work is also significant because both quantitative analyses and manual evaluations reveal the observed gains stem from the models' ability to refine an initial reasoning chain by generating a second, improved chain within the same inference step, demonstrating previously elusive self-improvement. Our code and data are publicly available at https://github.com/UKPLab/acl2025-diverse-cot.
Problem

Research questions and friction points this paper is trying to address.

Enhancing LLM reasoning via within-inference CoT refinement
Improving performance on tasks with diverse reasoning types
Enabling self-improvement in reasoning chains without external feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuning LLMs for Diverse Chains of Thought
Generating multiple reasoning chains in one step
Enabling self-improvement within single inference
🔎 Similar Papers
No similar papers found.