🤖 AI Summary
This work addresses the stabilization control of an inverted pendulum under vertical excitation by proposing a reinforcement learning approach that incorporates the Lyapunov Characteristic Exponent (LCE) as a physics-informed dense reward. For the first time within a reinforcement learning framework, LCE is employed to guide policy optimization, enabling the recovery of the classical Kapitza pendulum stabilization mechanism—reliant on high-frequency oscillations—and further achieving a novel stable state wherein the pendulum remains perfectly upright without any oscillation. Experimental results demonstrate that the proposed method effectively steers the agent toward discovering control policies superior to conventional mechanisms, substantially enhancing both system stability and control accuracy.
📝 Abstract
We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted pendulum with vertical motion. With LCE, the agent not only successfully found the oscillatory motion known as the Kapitza pendulum but also damped the pendulum's pivoting, leaving it in a strictly upright position.