policy-optimized thinker

Designs, trains, and evaluates a model component (the "thinker") whose internal reasoning or action-selection is treated as a policy and optimized with a mixture of supervised learning and reinforcement learning, typically via two-stage workflows (e.g., pretraining on reconstructed thoughts then RL fine-tuning). Implements reward functions, loss terms, and training pipelines to align the thinker’s generated internal states or decision-making behavior to target strategies while retaining task competency under compute or capacity constraints.

policy-optimizedthinker

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$122K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

When Can Model-Free Reinforcement Learning be Enough for Thinking?

Jun 20, 2025
JP
Josiah P. Hanna
🏛️ University of Wisconsin – Madison

This work investigates the emergence mechanism of “thinking” behaviors in model-free reinforcement learning (RL). Method: We propose the Thought-Markov Decision Process (Thought-MDP) framework—a formal theoretical model that first defines endogenous thinking actions as online policy improvement steps and rigorously proves their equivalence to standard RL updates. We identify policy initialization as the critical condition for thinking emergence and derive a general necessary and sufficient criterion characterizing when model-free RL acquires reasoning capability. Results: Through theoretical modeling, formal proofs, and empirical analysis on open-source large language models (LLMs), we validate that mainstream LLMs satisfy the theoretical predictions. On synthetic reasoning tasks, explicitly incorporating thinking actions significantly improves data efficiency. This work establishes the first empirically testable theoretical foundation for understanding the RL origins of LLM reasoning capabilities.

How policy initialization affects emergence of thinkingSufficient conditions for learning thinking outside languageWhen does model-free RL enable thinking-like behavior

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

Oct 17, 2025
PS
Pratham Singla
🏛️ Indian Institute of Technology Roorkee | Boston University

This work investigates three critical properties of post-trained language models: (1) self-awareness of their own decision-making strategies, (2) cross-domain generalization capability, and (3) alignment between internal reasoning trajectories and final outputs. We propose a three-dimensional evaluation framework and conduct systematic comparisons across supervised fine-tuning (SFT), direct preference optimization (DPO), and group-relative policy optimization (GRPO) models on diverse multi-task benchmarks. Results show that reinforcement learning–based methods—particularly DPO and GRPO—significantly outperform SFT in strategy awareness and cross-task transfer. However, all RL-based models exhibit weak alignment between reasoning paths and outputs, with GRPO showing the most severe inconsistency. To our knowledge, this is the first study to systematically expose the “strong behavior, weak reasoning” tension inherent in current RL-based post-training paradigms. Our findings provide empirical grounding for advancing interpretable AI and trustworthy reasoning modeling, highlighting concrete directions for improving reasoning fidelity in policy-optimized language models.

Assessing generalization of learned policies across domainsEvaluating reasoning awareness in post-trained language modelsMeasuring alignment between internal reasoning and final outputs

Thought-Augmented Policy Optimization: Bridging External Guidance and Internal Capabilities

May 21, 2025
JW
Jinyang Wu
🏛️ Tsinghua University | Beijing National Research Center for Information Science and Technology | Shanghai Qi Zhi Institute | Shanghai AI Lab

Reinforcement learning (RL) for reasoning models suffers from weak exploration and narrow reasoning boundaries due to insufficient external knowledge. Method: This paper proposes *Thought-Augmented Reinforcement Learning*—the first framework to dynamically inject generalizable, high-order abstract thought patterns as external guidance signals into policy optimization, enabling adaptive synergy between internal exploration and external steering. It integrates structured thought embedding, adaptive weight modulation, and joint thought-action modeling, and builds a lightweight, efficient training paradigm grounded in policy gradients. Contribution/Results: With only 500 samples, the method enables cross-task and cross-model transfer, significantly improving reasoning interpretability and output readability. It outperforms GRPO by 99%, 41%, and 17% on AIME, AMC, and Minerva Math, respectively, demonstrating dual gains in reasoning performance and generalization capability through explicit thought injection.

Bias in RL without external knowledge limits explorationEnhance reasoning models' explainability and output readabilityNeed to balance internal exploration and external guidance

StepWiser: Stepwise Generative Judges for Wiser Reasoning

Aug 26, 2025
WX
Wei Xiong
🏛️ FAIR at Meta | NYU | University of Illinois Urbana-Champaign

Supervising the logical validity of intermediate reasoning steps in multi-step inference remains challenging due to the difficulty of obtaining reliable, fine-grained step-level feedback. Method: This paper proposes a generative judge model that reformulates step-level reward modeling as an interpretable meta-reasoning task. Instead of relying on static annotations or black-box scoring, it employs a reinforcement learning framework that optimizes a generative judgment policy via relative rollout outcomes, producing fine-grained, process-aware step evaluation tokens. Contribution/Results: To our knowledge, this is the first work to cast judging as a generative reasoning task—enabling traceable criteria and fully interpretable judgments. Moreover, it supports online policy optimization and accelerated inference search. Experiments demonstrate significant improvements over existing baselines in intermediate-step accuracy, while also enhancing final answer quality and search efficiency.

Addressing limitations of stepwise feedback without explanationsImproving generalization beyond static datasets in process rewardsSupervising logical validity of intermediate reasoning steps

ThinkTuning: Instilling Cognitive Reflections without Distillation

Aug 11, 2025
AR
Aswin RRV
🏛️ Arizona State University

This work addresses the lack of reflexivity and multi-step reasoning capabilities in foundational large language models. We propose ThinkTuning—a novel fine-tuning paradigm that, without knowledge distillation, leverages implicit feedback (e.g., thought-trace correction and cognitive guidance) from a same-scale teacher model during inference to establish a classroom-like interactive training mechanism, dynamically eliciting latent reasoning and self-reflective abilities in the student model. Built upon the GRPO framework, ThinkTuning enables interactive reinforcement training, overcoming the limitation of conventional RL methods that merely exploit pre-existing capabilities. Experiments demonstrate substantial improvements across multiple reasoning benchmarks: an average gain of 3.85% over zero-shot baselines; and relative improvements of 2.08%, 2.23%, and 3.99% over standard GRPO on MATH-500, AIME, and GPQA-Diamond, respectively—validating its effectiveness and generalizability in fostering emergent reasoning capabilities.

Develop thinking abilities in non-reflective language modelsImprove reasoning without relying solely on reinforcement learningTrain student models using teacher-guided interactive feedback

Latest Papers

What's happening recently
View more

Current understanding of how reinforcement learning (RL) enhances reasoning capabilities during post-training remains unclear. This study addresses this gap through controlled mathematical reasoning experiments, explicitly disentangling and validating two core mechanisms in post-training: policy selection and policy improvement. Leveraging the Qwen-2.5-1.5B model with diverse supervised fine-tuning (SFT) data and progressively harder RL data, the research demonstrates that diverse SFT data effectively facilitates policy selection, while high-difficulty RL data drives policy improvement. The findings not only clarify the distinct roles of SFT and RL data in activating these mechanisms but also offer actionable pathways for enhancing model reasoning performance.

mechanistic understandingpost-trainingreasoning models

This work investigates how reasoning-focused fine-tuning reshapes the internal mechanisms of language models, with particular emphasis on its impact on local token-level capabilities and global reasoning temporal structures. We model chain-of-thought reasoning as a Switching Dynamical System (SDS), integrating time-aware contrastive learning with discrete latent state discovery to recover functionally specialized strategy states from activation trajectories. For the first time, we reveal that reasoning fine-tuning induces persistent and structured strategy dynamics, establishing SDS-based analysis as a novel paradigm for mechanistic interpretability. Experiments across models ranging from 1.5B to 32B parameters and four benchmarks demonstrate that fine-tuned models exhibit richer strategy structures; strategy state transfer enhances baseline performance; and SDS-guided dynamic pruning outperforms self-consistency in 11 out of 12 settings, achieving gains up to 12.5 percentage points.

chain-of-thoughtdynamical systemslatent policy states

Reinforcement learning in reasoning tasks suffers from sparse outcome supervision and difficulty in credit assignment across intermediate steps, while existing process supervision relies on costly human annotations that hinder scalability. This work proposes a novel paradigm termed “supervision internalization,” which leverages a self-reflection mechanism to automatically identify and correct failed reasoning trajectories, thereby generating fine-grained process-level supervision signals endogenously from only outcome feedback—without requiring external annotations. This approach enables precise credit assignment and significantly improves both policy training efficiency and reasoning performance, offering a scalable pathway toward fine-grained reinforcement learning for complex reasoning tasks.

credit assignmentoutcome supervisionprocess supervision

This work addresses the limited performance of existing image editing models on complex reasoning tasks, which stems primarily from inadequate modeling of planning capabilities. To overcome this, the authors propose the DDA-Thinker framework, introducing a novel Thinker-centric paradigm that decouples the reasoning planner (the Thinker) from the generative module (the Editor), enabling independent optimization of the Thinker while keeping the Editor fixed. The approach incorporates a dual atomic reward mechanism—combining cognitive and visual feedback based on verifiable checklists—and a difficulty-aware curriculum learning strategy, supported by a two-stage data construction pipeline. Experimental results demonstrate that DDA-Thinker significantly outperforms baseline methods on both RISE-Bench and KRIS-Bench, achieving performance on par with powerful closed-source models using only open-source components.

complex reasoningimage editing benchmarksplanning module

This work addresses the limitations of existing vision-language-action (VLA) models in long-horizon tasks, which rely on static visual contexts and purely textual reasoning, thereby struggling to actively revisit visual inputs to resolve ambiguities. To overcome this, the authors propose an “Image Thinking” reasoning framework that, for the first time, models visual perception as a dynamically invocable reasoning action, enabling on-demand revisiting of environmental images during task execution. The approach combines supervised fine-tuning (SFT) for cold-start initialization with GRPO reinforcement learning, leveraging visual chain-of-thought data to align structured reasoning with tool-use behaviors. Evaluated on the LIBERO and RoboTwin 2.0 benchmarks, the method achieves significant performance gains, reaching a 97.5% success rate on LIBERO tasks and demonstrating substantial improvements in long-horizon robotic manipulation.

chain-of-thought reasoningembodied intelligencelong-horizon tasks

Hot Scholars

KC

Kewei Chen

Arizona State University, AZ
Alzheimer'sneuroimage (PET MRI)statisticsML/AI
ZL

Zhongtang Luo

Research Assistant, Purdue University
CryptographyNetwork SecurityBlockchain
EB

Engin Bumbacher

University of Teacher Education Vaud
STEM EducationComputational ThinkingInnovative Assessment
FM

Francesco Mondada

Ecole Polytechnique Fédérale de Lausanne
RoboticsEducationMechatronics
GA

Giorgia Adorni

Institute of Information Systems and Networking (ISIN), SUPSI
Artificial IntelligenceComputer Science EducationLearning TechnologiesGenerative AI