Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出了一种反向推理模式构建方法,通过两阶段数学数据集训练和细粒度奖励机制,提高大型语言模型的反向思维能力和动态适应性,以解决复杂问题。
📝 Abstract
When facing complex problems, humans tend to try various ideas for different issues. Human thinking patterns exhibit remarkable flexibility in adapting to diverse scenarios. GPT-o1, GPT-o3, and DeepSeek-R1 adopt long chain-of-thought models to address complex problems by increasing reasoning depth, which default to a forward reasoning mode. We conducted statistical analysis on the accuracy of different mathematical problem datasets on models of different scales, and found five reasons for errors: Insufficient solution-space coverage, Computational mistakes, Unverified assumptions, Ignoring constraint conditions, Maximum response length limitation. To address the above issues, we proposed a backward reasoning pattern construction method aimed at enhancing the model's reverse thinking ability and dynamic adaptability. First, we constructed an easy-hard two-stage Math dataset for training large models and gradually improving their inference ability at different difficulty levels. The dataset contains forward reasoning paths as well as backward reasoning paths. And a two-stage supervised fine-tuning process is applied to progressively train the model's backward reasoning capability. Furthermore, a fine-grained reward mechanism is developed, employing smoothed reward signals to strengthen the model's ability to autonomously select thinking modes during the reasoning process, thereby avoiding reward hacking. A linear-decay balanced sampling strategy is designed to maintain a balance between forward and backward reasoning path samples during training, enabling the model to converge quickly and stably. Experimental results show that our method significantly improves reasoning efficiency and accuracy in tasks such as mathematical proofs, offering a flexible and efficient reasoning paradigm for solving complex problems.
Problem

Research questions and friction points this paper is trying to address.

forward reasoning
complex problems
large language models
reasoning depth
error analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

backward reasoning
dynamic adaptability
fine-grained reward mechanism
linear-decay balanced sampling
🔎 Similar Papers
No similar papers found.
X
Xin Liu
China University of Petroleum (East China), Qingdao, China
Y
Yunhai Li
China University of Petroleum (East China), Qingdao, China
C
Chunfu Jia
College of Cryptology and Cyber Science, Nankai University, Tianjin, China
Ziliang Chen
Ziliang Chen
AP, Pengcheng Lab
Machine learningFoundation ModelsMultimodal Embodied Intelligence
J
Jisen Song
China University of Petroleum (East China), Qingdao, China