Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragility of hyperparameter tuning and policy instability in deep reinforcement learning caused by heuristic-dependent reward shaping. Through theoretical analysis and empirical experiments in robotic control, it systematically investigates how zeroth- and first-order information affects policy gradients in stable control scenarios. We rigorously prove that control tasks can be accomplished without first-order reward terms, and that introducing such terms significantly increases policy gradient sensitivity. Based on these findings, we propose a “zeroth-order completeness” principle for reward design. This work provides theoretically grounded yet practically actionable reward design guidelines for robotic reinforcement learning, effectively reducing tuning complexity and enhancing training stability.
📝 Abstract
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.
Problem

Research questions and friction points this paper is trying to address.

reward shaping
stabilization control
policy gradient
deep reinforcement learning
robotic control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Shaping
Policy Gradient
Stabilization Control
Zeroth-Order Information
Deep Reinforcement Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yisheng Zhang
University of California San Diego, La Jolla, CA, USA; Tsinghua University, Beijing, China
T
Tao Wang
University of California San Diego, La Jolla, CA, USA
Sicun Gao
Sicun Gao
UCSD
ReasoningOptimizationAutomation