Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

📅 2026-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决RLVR中二元奖励无法区分正确轨迹的问题,提出使用基于梯度对齐的奖励方法(GAR),通过与专家梯度的余弦相似性提供密集奖励。
📝 Abstract
Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosine admits a multiplicative decomposition into prediction-error and activation-pattern factors, providing a concrete characterization of what the alignment signal measures. On Qwen3-4B and Qwen3-8B, GAR consistently improves over GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. Code and data are available at https://github.com/LQgdwind/GAR.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Verifiable Rewards
Chain-of-Thought Reasoning
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient-Aligned Reward (GAR)
truncated backpropagation
cosine similarity
reasoning-aware reward
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Leqi Zheng
Tsinghua University
J
Jinbo Su
Renmin University of China
F
Fang Niu
Tsinghua University
Chaokun Wang
Chaokun Wang
Tsinghua University
DatabaseMultimediaSocial Networks
Weiping Wang
Weiping Wang
School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security
Jiajun Zhang
Jiajun Zhang
Institute of Automation Chinese Academy of Sciences
Natural Language ProcessingLarge Language ModelsMultimodal Information Processing
S
Shannan Yan
Tsinghua University
J
Jie Wu
The Australian National University
Z
Zhaolu Kang
Peking University
R
Rong Fu
University of Macau
H
Hang Zhang
Tsinghua University