DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决强化学习中策略-奖励流形不匹配问题,提出DCRL框架,通过三元逻辑提示进化机制和策略-奖励重耦合机制来提升大型语言模型的推理能力。
📝 Abstract
Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing reward systems, such as rule-based and reward-model-based, often exhibit issues such as unstable optimization and reward hacking. In this work, we revisit the general reasoning of LLMs from a geometric perspective, conceptualizing it as a coupled manifold composed of three interdependent sub-manifolds: logical deduction, evaluation, and representation. Based on this perspective, response generation in RL can be interpreted as a decoupling process from the evaluation manifold, while reward estimation corresponds to a decoupling process from the logical deduction manifold. The limitations of rule-based and reward-model RL systems can be geometrically interpreted as the mismatch of policy-reward manifolds during RL process. To address the aforementioned misalignment, we propose Decoupling and Coupling Reinforcement Learning (DCRL) framework, which incorporates two key components: (1) a syllogistic logic-based prompt evolution mechanism that dynamically refines reward rubrics to enhance the expressiveness of the reward manifold; and (2) a policy-reward re-coupling mechanism that jointly updates the reward and policy models, ensuring consistent evaluation and mitigating manifold mismatch during training. Theoretical analysis and extensive experiments across multiple reasoning domains demonstrate that DCRL consistently outperforms both rule-based and reward-model baselines. Notably, a Qwen3-4B model trained under DCRL surpasses a Qwen3-32B baseline and approaches the performance of a Qwen3-235B model, highlighting superior effectiveness and generalization in RL.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Reward Systems
Optimization Stability
Reward Hacking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoupling and Coupling Reinforcement Learning (DCRL)
Policy-Reward Manifold Alignment
Syllogistic Logic-Based Prompt Evolution
Reward Rubrics Refinement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Henan Sun
Henan Sun
Beijing Institute of Technology
Graph Neural Networks (GNNs)Graph Invariant LearningDifferential Privacy
Z
Zehua Li
The Hong Kong University of Science and Technology (Guangzhou)
H
Haitao Hu
The Hong Kong University of Science and Technology (Guangzhou)
Q
Qifan Zhang
Huawei Noah’s Ark Lab
J
Jianfeng Zhang
Huawei Noah’s Ark Lab
Nuo Chen
Nuo Chen
Hong Kong University of Science and Technology
large language modelpre-trainreasoningrole-playing
J
Jia Li
The Hong Kong University of Science and Technology (Guangzhou), The Hong Kong University of Science and Technology