🤖 AI Summary
This study addresses the high sample complexity of reinforcement learning in physical domains and its difficulty in exploiting geometric symmetries. We propose a policy learning method based on group-symmetry-induced homomorphisms, integrating group theory with equivariant neural networks. For the first time, we rigorously prove that incorporating symmetry substantially reduces the number of interactions required to identify optimal policies in both finite-horizon Markov Decision Processes and continuous state spaces, providing formal theoretical guarantees. High-dimensional robotic simulation experiments demonstrate that this symmetry-aware approach significantly improves sample efficiency and control performance. By establishing principled bounds on interaction complexity, this work offers a novel pathway toward data-efficient robot control.
📝 Abstract
Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.