Thought-Augmented Policy Optimization: Bridging External Guidance and Internal Capabilities

πŸ“… 2025-05-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Reinforcement learning (RL) for reasoning models suffers from weak exploration and narrow reasoning boundaries due to insufficient external knowledge. Method: This paper proposes *Thought-Augmented Reinforcement Learning*β€”the first framework to dynamically inject generalizable, high-order abstract thought patterns as external guidance signals into policy optimization, enabling adaptive synergy between internal exploration and external steering. It integrates structured thought embedding, adaptive weight modulation, and joint thought-action modeling, and builds a lightweight, efficient training paradigm grounded in policy gradients. Contribution/Results: With only 500 samples, the method enables cross-task and cross-model transfer, significantly improving reasoning interpretability and output readability. It outperforms GRPO by 99%, 41%, and 17% on AIME, AMC, and Minerva Math, respectively, demonstrating dual gains in reasoning performance and generalization capability through explicit thought injection.

Technology Category

Machine Learning: Reinforcement LearningKnowledge Representation and Reasoning: Reasoning with BeliefsCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
πŸ“ Abstract
Reinforcement learning (RL) has emerged as an effective method for training reasoning models. However, existing RL approaches typically bias the model's output distribution toward reward-maximizing paths without introducing external knowledge. This limits their exploration capacity and results in a narrower reasoning capability boundary compared to base models. To address this limitation, we propose TAPO (Thought-Augmented Policy Optimization), a novel framework that augments RL by incorporating external high-level guidance ("thought patterns"). By adaptively integrating structured thoughts during training, TAPO effectively balances model-internal exploration and external guidance exploitation. Extensive experiments show that our approach significantly outperforms GRPO by 99% on AIME, 41% on AMC, and 17% on Minerva Math. Notably, these high-level thought patterns, abstracted from only 500 prior samples, generalize effectively across various tasks and models. This highlights TAPO's potential for broader applications across multiple tasks and domains. Our further analysis reveals that introducing external guidance produces powerful reasoning models with superior explainability of inference behavior and enhanced output readability.
Problem

Research questions and friction points this paper is trying to address.

Bias in RL without external knowledge limits exploration
Need to balance internal exploration and external guidance
Enhance reasoning models' explainability and output readability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Augments RL with external high-level guidance
Balances internal exploration and external guidance
Generalizes thought patterns from few samples
πŸ”Ž Similar Papers
No similar papers found.
J
Jinyang Wu
Department of Automation, Tsinghua University
C
Chonghua Liao
Institution for Interdisciplinary Information Sciences, Tsinghua University
M
Mingkuan Feng
Department of Automation, Tsinghua University
S
Shuai Zhang
Department of Automation, Tsinghua University
Zhengqi Wen
Zhengqi Wen
Tshinghua University
LLM
P
Pengpeng Shao
Beijing National Research Center for Information Science and Technology
Huazhe Xu
Huazhe Xu
Tsinghua University
Embodied AIReinforcement LearningComputer VisionDeep Learning
J
Jianhua Tao
Department of Automation, Tsinghua University, Beijing National Research Center for Information Science and Technology