Score
Designs, trains, and evaluates a model component (the "thinker") whose internal reasoning or action-selection is treated as a policy and optimized with a mixture of supervised learning and reinforcement learning, typically via two-stage workflows (e.g., pretraining on reconstructed thoughts then RL fine-tuning). Implements reward functions, loss terms, and training pipelines to align the thinker’s generated internal states or decision-making behavior to target strategies while retaining task competency under compute or capacity constraints.
This work addresses the limited reasoning capabilities of large language models (LLMs) by deeply integrating reinforcement learning (RL) into the entire “slow-thinking” stepwise reasoning process—both during training and decoding. We introduce the first formalized “slow-thinking” reasoning framework, unifying model-based (e.g., Monte Carlo Tree Search) and model-free (e.g., Proximal Policy Optimization) RL paradigms to render reasoning steps learnable, intervenable, and interpretable. Our method jointly incorporates stepwise reasoning modeling, chain-of-thought fine-tuning, reward modeling over reasoning trajectories, and in-model policy optimization. Evaluated on benchmarks including GSM8K and MMLU, our approach achieves 12–18% absolute improvements in mathematical reasoning and complex problem-solving over strong baselines. Moreover, generated reasoning processes exhibit enhanced reliability, transparency, and traceability—enabling rigorous auditing and human-in-the-loop refinement.
This work investigates the emergence mechanism of “thinking” behaviors in model-free reinforcement learning (RL). Method: We propose the Thought-Markov Decision Process (Thought-MDP) framework—a formal theoretical model that first defines endogenous thinking actions as online policy improvement steps and rigorously proves their equivalence to standard RL updates. We identify policy initialization as the critical condition for thinking emergence and derive a general necessary and sufficient criterion characterizing when model-free RL acquires reasoning capability. Results: Through theoretical modeling, formal proofs, and empirical analysis on open-source large language models (LLMs), we validate that mainstream LLMs satisfy the theoretical predictions. On synthetic reasoning tasks, explicitly incorporating thinking actions significantly improves data efficiency. This work establishes the first empirically testable theoretical foundation for understanding the RL origins of LLM reasoning capabilities.
This work investigates three critical properties of post-trained language models: (1) self-awareness of their own decision-making strategies, (2) cross-domain generalization capability, and (3) alignment between internal reasoning trajectories and final outputs. We propose a three-dimensional evaluation framework and conduct systematic comparisons across supervised fine-tuning (SFT), direct preference optimization (DPO), and group-relative policy optimization (GRPO) models on diverse multi-task benchmarks. Results show that reinforcement learning–based methods—particularly DPO and GRPO—significantly outperform SFT in strategy awareness and cross-task transfer. However, all RL-based models exhibit weak alignment between reasoning paths and outputs, with GRPO showing the most severe inconsistency. To our knowledge, this is the first study to systematically expose the “strong behavior, weak reasoning” tension inherent in current RL-based post-training paradigms. Our findings provide empirical grounding for advancing interpretable AI and trustworthy reasoning modeling, highlighting concrete directions for improving reasoning fidelity in policy-optimized language models.
Reinforcement learning (RL) for reasoning models suffers from weak exploration and narrow reasoning boundaries due to insufficient external knowledge. Method: This paper proposes *Thought-Augmented Reinforcement Learning*—the first framework to dynamically inject generalizable, high-order abstract thought patterns as external guidance signals into policy optimization, enabling adaptive synergy between internal exploration and external steering. It integrates structured thought embedding, adaptive weight modulation, and joint thought-action modeling, and builds a lightweight, efficient training paradigm grounded in policy gradients. Contribution/Results: With only 500 samples, the method enables cross-task and cross-model transfer, significantly improving reasoning interpretability and output readability. It outperforms GRPO by 99%, 41%, and 17% on AIME, AMC, and Minerva Math, respectively, demonstrating dual gains in reasoning performance and generalization capability through explicit thought injection.
Supervising the logical validity of intermediate reasoning steps in multi-step inference remains challenging due to the difficulty of obtaining reliable, fine-grained step-level feedback. Method: This paper proposes a generative judge model that reformulates step-level reward modeling as an interpretable meta-reasoning task. Instead of relying on static annotations or black-box scoring, it employs a reinforcement learning framework that optimizes a generative judgment policy via relative rollout outcomes, producing fine-grained, process-aware step evaluation tokens. Contribution/Results: To our knowledge, this is the first work to cast judging as a generative reasoning task—enabling traceable criteria and fully interpretable judgments. Moreover, it supports online policy optimization and accelerated inference search. Experiments demonstrate significant improvements over existing baselines in intermediate-step accuracy, while also enhancing final answer quality and search efficiency.
This work addresses the lack of reflexivity and multi-step reasoning capabilities in foundational large language models. We propose ThinkTuning—a novel fine-tuning paradigm that, without knowledge distillation, leverages implicit feedback (e.g., thought-trace correction and cognitive guidance) from a same-scale teacher model during inference to establish a classroom-like interactive training mechanism, dynamically eliciting latent reasoning and self-reflective abilities in the student model. Built upon the GRPO framework, ThinkTuning enables interactive reinforcement training, overcoming the limitation of conventional RL methods that merely exploit pre-existing capabilities. Experiments demonstrate substantial improvements across multiple reasoning benchmarks: an average gain of 3.85% over zero-shot baselines; and relative improvements of 2.08%, 2.23%, and 3.99% over standard GRPO on MATH-500, AIME, and GPQA-Diamond, respectively—validating its effectiveness and generalizability in fostering emergent reasoning capabilities.
Current understanding of how reinforcement learning (RL) enhances reasoning capabilities during post-training remains unclear. This study addresses this gap through controlled mathematical reasoning experiments, explicitly disentangling and validating two core mechanisms in post-training: policy selection and policy improvement. Leveraging the Qwen-2.5-1.5B model with diverse supervised fine-tuning (SFT) data and progressively harder RL data, the research demonstrates that diverse SFT data effectively facilitates policy selection, while high-difficulty RL data drives policy improvement. The findings not only clarify the distinct roles of SFT and RL data in activating these mechanisms but also offer actionable pathways for enhancing model reasoning performance.
This work investigates how reasoning-focused fine-tuning reshapes the internal mechanisms of language models, with particular emphasis on its impact on local token-level capabilities and global reasoning temporal structures. We model chain-of-thought reasoning as a Switching Dynamical System (SDS), integrating time-aware contrastive learning with discrete latent state discovery to recover functionally specialized strategy states from activation trajectories. For the first time, we reveal that reasoning fine-tuning induces persistent and structured strategy dynamics, establishing SDS-based analysis as a novel paradigm for mechanistic interpretability. Experiments across models ranging from 1.5B to 32B parameters and four benchmarks demonstrate that fine-tuned models exhibit richer strategy structures; strategy state transfer enhances baseline performance; and SDS-guided dynamic pruning outperforms self-consistency in 11 out of 12 settings, achieving gains up to 12.5 percentage points.
Reinforcement learning in reasoning tasks suffers from sparse outcome supervision and difficulty in credit assignment across intermediate steps, while existing process supervision relies on costly human annotations that hinder scalability. This work proposes a novel paradigm termed “supervision internalization,” which leverages a self-reflection mechanism to automatically identify and correct failed reasoning trajectories, thereby generating fine-grained process-level supervision signals endogenously from only outcome feedback—without requiring external annotations. This approach enables precise credit assignment and significantly improves both policy training efficiency and reasoning performance, offering a scalable pathway toward fine-grained reinforcement learning for complex reasoning tasks.
This work addresses the limited performance of existing image editing models on complex reasoning tasks, which stems primarily from inadequate modeling of planning capabilities. To overcome this, the authors propose the DDA-Thinker framework, introducing a novel Thinker-centric paradigm that decouples the reasoning planner (the Thinker) from the generative module (the Editor), enabling independent optimization of the Thinker while keeping the Editor fixed. The approach incorporates a dual atomic reward mechanism—combining cognitive and visual feedback based on verifiable checklists—and a difficulty-aware curriculum learning strategy, supported by a two-stage data construction pipeline. Experimental results demonstrate that DDA-Thinker significantly outperforms baseline methods on both RISE-Bench and KRIS-Bench, achieving performance on par with powerful closed-source models using only open-source components.
This work addresses the limitations of existing vision-language-action (VLA) models in long-horizon tasks, which rely on static visual contexts and purely textual reasoning, thereby struggling to actively revisit visual inputs to resolve ambiguities. To overcome this, the authors propose an “Image Thinking” reasoning framework that, for the first time, models visual perception as a dynamically invocable reasoning action, enabling on-demand revisiting of environmental images during task execution. The approach combines supervised fine-tuning (SFT) for cold-start initialization with GRPO reinforcement learning, leveraging visual chain-of-thought data to align structured reasoning with tool-use behaviors. Evaluated on the LIBERO and RoboTwin 2.0 benchmarks, the method achieves significant performance gains, reaching a 97.5% success rate on LIBERO tasks and demonstrating substantial improvements in long-horizon robotic manipulation.