🤖 AI Summary
This study addresses the critic overestimation problem caused by fixed minimum aggregation in off-policy reinforcement learning by proposing the GeZo-SAC algorithm. The method introduces a Zonotope polyhedral geometric representation, computing a geometric width via predicted generators to serve as an adaptive pessimism offset. Combined with a Log-sum-exp mechanism, this enables state-action-dependent dynamic value aggregation. The core innovation lies in adaptively modulating the degree of pessimism through geometric information during training while preserving the standard SAC architecture during inference, thereby balancing efficiency and accuracy. Evaluated on MuJoCo benchmarks, GeZo-SAC achieves the highest average returns, significantly reduces actuator energy consumption, and suppresses the overestimation frequency to near-zero levels.
📝 Abstract
Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.