🤖 AI Summary
To address the poor training stability and slow convergence of Adaptive Neuro-Fuzzy Inference System (ANFIS) policies in reinforcement learning, this paper proposes an on-policy actor-critic framework based on Proximal Policy Optimization (PPO) for end-to-end training of ANFIS controllers. Departing from the off-policy paradigm of Deep Q-Networks (DQN), the approach preserves the interpretability of fuzzy rules while significantly enhancing training robustness and convergence efficiency. Extensive experiments across multiple random seeds in the CartPole-v1 environment demonstrate that the proposed ANFIS-PPO agent consistently achieves the maximum return of 500 after 20,000 policy updates (zero variance), outperforming the ANFIS-DQN baseline. The key contribution lies in the first integration of PPO’s gradient clipping and trust-region optimization mechanisms into joint parameter learning of ANFIS, thereby unifying interpretability with high control performance.
📝 Abstract
We present a reinforcement learning method for training neuro-fuzzy controllers using Proximal Policy Optimization (PPO). Unlike prior approaches that used Deep Q-Networks (DQN) with Adaptive Neuro-Fuzzy Inference Systems (ANFIS), our PPO-based framework leverages a stable on-policy actor-critic setup. Evaluated on the CartPole-v1 environment across multiple seeds, PPO-trained fuzzy agents consistently achieved the maximum return of 500 with zero variance after 20000 updates, outperforming ANFIS-DQN baselines in both stability and convergence speed. This highlights PPO's potential for training explainable neuro-fuzzy agents in reinforcement learning tasks.