Reward Inflation: A Healthy Stimulus for Reinforcement Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that fixed reward magnitudes in reinforcement learning (RL) induce policy saturation, vanishing gradients, and neuronal dormancy, thereby constraining model adaptability and plasticity. To overcome these limitations, this work proposes a reward inflation mechanism that progressively scales reward values during training to introduce implicit recency weighting while preserving gradient signals. An adaptive variant, Fed, is further designed to dynamically regulate the inflation level. Supported by theoretical analysis within deep RL frameworks, this approach achieves temporally dynamic optimization of reward signals. This paper is the first to reveal the critical role of temporal reward modulation. Experiments on ALE and MuJoCo benchmarks demonstrate that moderate reward inflation significantly mitigates neuronal dormancy and enhances generalization, with the adaptive method outperforming fixed configurations. These findings establish a novel paradigm for RL optimization.
📝 Abstract
Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Reward Inflation
Dormant Neurons
Plasticity
Temporal Modulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Inflation
Reinforcement Learning
Recency Weighting
Dormant Neurons
Adaptive Scaling