On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the poor training stability and slow convergence of Adaptive Neuro-Fuzzy Inference System (ANFIS) policies in reinforcement learning, this paper proposes an on-policy actor-critic framework based on Proximal Policy Optimization (PPO) for end-to-end training of ANFIS controllers. Departing from the off-policy paradigm of Deep Q-Networks (DQN), the approach preserves the interpretability of fuzzy rules while significantly enhancing training robustness and convergence efficiency. Extensive experiments across multiple random seeds in the CartPole-v1 environment demonstrate that the proposed ANFIS-PPO agent consistently achieves the maximum return of 500 after 20,000 policy updates (zero variance), outperforming the ANFIS-DQN baseline. The key contribution lies in the first integration of PPO’s gradient clipping and trust-region optimization mechanisms into joint parameter learning of ANFIS, thereby unifying interpretability with high control performance.

Technology Category

Multiagent Systems: Multiagent LearningIntelligent Robots: Behavior Learning & ControlMachine Learning: Reinforcement Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
We present a reinforcement learning method for training neuro-fuzzy controllers using Proximal Policy Optimization (PPO). Unlike prior approaches that used Deep Q-Networks (DQN) with Adaptive Neuro-Fuzzy Inference Systems (ANFIS), our PPO-based framework leverages a stable on-policy actor-critic setup. Evaluated on the CartPole-v1 environment across multiple seeds, PPO-trained fuzzy agents consistently achieved the maximum return of 500 with zero variance after 20000 updates, outperforming ANFIS-DQN baselines in both stability and convergence speed. This highlights PPO's potential for training explainable neuro-fuzzy agents in reinforcement learning tasks.
Problem

Research questions and friction points this paper is trying to address.

Optimizing ANFIS policies using PPO method
Improving stability and convergence in neuro-fuzzy controllers
Demonstrating PPO's superiority over DQN for ANFIS training
Innovation

Methods, ideas, or system contributions that make the work stand out.

PPO-based on-policy actor-critic setup
Training neuro-fuzzy controllers with PPO
Stable ANFIS policies outperform DQN baselines
University of Cincinnati
K
Kaaustaaub Shankar
College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45219, USA
W
Wilhelm Louw
College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45219, USA
Kelly Cohen
Kelly Cohen
Brian Rowe Endowed Chair in Aerospace Engineering, University of Cincinnati
UAV & roboticsAdvanced Air MobilityIntelligent autonomous systemsGenetic fuzzy AI & predictive modelingFeedback flow con