Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum

📅 2026-07-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the stabilization control of an inverted pendulum under vertical excitation by proposing a reinforcement learning approach that incorporates the Lyapunov Characteristic Exponent (LCE) as a physics-informed dense reward. For the first time within a reinforcement learning framework, LCE is employed to guide policy optimization, enabling the recovery of the classical Kapitza pendulum stabilization mechanism—reliant on high-frequency oscillations—and further achieving a novel stable state wherein the pendulum remains perfectly upright without any oscillation. Experimental results demonstrate that the proposed method effectively steers the agent toward discovering control policies superior to conventional mechanisms, substantially enhancing both system stability and control accuracy.
📝 Abstract
We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted pendulum with vertical motion. With LCE, the agent not only successfully found the oscillatory motion known as the Kapitza pendulum but also damped the pendulum's pivoting, leaving it in a strictly upright position.
Problem

Research questions and friction points this paper is trying to address.

Lyapunov Exponent
Reinforcement Learning
Inverted Pendulum
Stabilization
Kapitza Pendulum
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lyapunov exponent
physics-informed reinforcement learning
dense reward
Kapitza pendulum
stabilization