A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critic overestimation problem caused by fixed minimum aggregation in off-policy reinforcement learning by proposing the GeZo-SAC algorithm. The method introduces a Zonotope polyhedral geometric representation, computing a geometric width via predicted generators to serve as an adaptive pessimism offset. Combined with a Log-sum-exp mechanism, this enables state-action-dependent dynamic value aggregation. The core innovation lies in adaptively modulating the degree of pessimism through geometric information during training while preserving the standard SAC architecture during inference, thereby balancing efficiency and accuracy. Evaluated on MuJoCo benchmarks, GeZo-SAC achieves the highest average returns, significantly reduces actuator energy consumption, and suppresses the overestimation frequency to near-zero levels.
📝 Abstract
Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.
Problem

Research questions and friction points this paper is trying to address.

overestimation bias
off-policy actor-critic
critic disagreement
pessimism adaptation
locomotion learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zonotope
Soft Actor-Critic
Adaptive Pessimism
Critic Disagreement
Locomotion Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Panagiotis Roditis
Robotics Institute, Athena Research Center, Marousi, Greece; HERON - Hellenic Robotics Center of Excellence, Athens, Greece; School of Electrical & Computer Engineering, NTUA, Greece
P
Panagiotis P. Filntisis
Robotics Institute, Athena Research Center, Marousi, Greece; HERON - Hellenic Robotics Center of Excellence, Athens, Greece
Petros Maragos
Petros Maragos
Professor of Electrical and Computer Engineering, National Technical University of Athens
computer visionsignal processingspeech&languagemachine learningrobotics