🤖 AI Summary
This study addresses the lack of theoretical guarantees in action shaping, which prevents the safe removal of deployment-time offsets. We propose an action shaping theorem proving that policies can losslessly absorb linear offsets exactly reproducible by the output layer, revealing that exact reproducibility—rather than network capacity—is the critical factor for absorption. Accordingly, we design a zero-initialized linear head, a learnable gating mechanism, and a learning-based action-value function training scheme, accompanied by a pre-removal diagnostic protocol. Extensive evaluations across twenty tasks demonstrate that removing the shaping head incurs negligible reward degradation, the gating mechanism adapts autonomously, and the approach generalizes broadly to both stochastic and deterministic policies.
📝 Abstract
Reward shaping has a theorem: a potential-based term can be removed without changing the optimal policy. The same practice on the action channel, an offset added in training and dropped at deployment, has no theorem. Nothing cancels an action offset, so the correction is kept at deployment or removed without a guarantee. We call it action shaping and state its principle. A trainable policy absorbs an offset its own output layer can reproduce exactly, which is what we mean by express; what is absorbed can be removed with the return intact. Its minimal instance is a zero-initialized linear head behind a learnable gate, added to an actor that trains through a learned action-value function, with no penalty or schedule. The gate rises and then falls on its own, for deterministic and stochastic actors alike, and on 20 tasks removing the head costs almost nothing. The condition is exact reproduction, not capacity: a nonlinear head with more parameters is not absorbed, and in a paired control, one linear path added to a nonlinear base head restores absorption. Exact reproduction gives the loss a flat direction that gradient noise drifts along, and the offset's amplitude indicates, before removal, what dropping the head will cost. Action shaping thus gains the counterpart of the shaping theorem, a condition for absorption, together with the mechanism behind it and a diagnostic that reads it. Policies absorb what they can express, and only that.