Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the signal imbalance in multi-reward reinforcement learning caused by uneven reward activation densities, which constrains the effectiveness of multi-objective training for large language models. To this end, it proposes a density-aware reward aggregation mechanism that reveals the intrinsic relationship between advantage energy and activation density. Specifically, the method dynamically adjusts the weights of sparse rewards through inverse square root density correction. Furthermore, by integrating GDPO normalization, it introduces an adaptive weighting algorithm that requires no modifications to the underlying objectives. Experimental results demonstrate that this approach reduces the number of training steps by 26% on tool-calling tasks and significantly improves length compliance in mathematical reasoning, while maintaining competitive overall performance.
📝 Abstract
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Problem

Research questions and friction points this paper is trying to address.

Multi-reward reinforcement learning
Sparse rewards
Signal imbalance
Reward aggregation
Large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-reward reinforcement learning
Density-Aware Reward Aggregation
Advantage energy
Inverse-square-root density correction
Large language models
Tong Zheng
Tong Zheng
University of Maryland, College Park; Northeastern University
Machine TranslationLanguage ModelingReasoningInference
S
Skylar Zhai
University of Minnesota Twin Cities
Z
Zhan Cheng
University of Wisconsin–Madison
T
TianMing Sha
Stony Brook University
Y
Youling Huang
Dalian University of Technology
S
Shuo Zhou
Beijing Foreign Studies University
S
Shaotong Qi
Southeast University
J
Jingcheng Liang
University of Minnesota Twin Cities
X
Xuwei Ding
University of Wisconsin–Madison
Pengcheng Xu
Pengcheng Xu
Western University
machine learninggenerative modeltransfer learningcomputer vision