🤖 AI Summary
This study addresses the performance degradation of reinforcement learning policies under out-of-distribution (OOD) magnitude shifts by proposing a Causal Symbolic BAPR (CS-BAPR) framework. Methodologically, it integrates the Soft Actor-Critic algorithm with Bayesian robust estimation to establish six training stabilization configurations. Furthermore, Neural Arithmetic Units (NAU) and Kolmogorov-Arnold Networks (KAN) are introduced as actor heads, achieving architectural innovation through a neural additive-multiplicative correction mechanism. By leveraging these components, this work effectively mitigates policy deterioration in OOD scenarios, significantly enhancing both model generalization capability and training stability.
📝 Abstract
Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The practical method combines six training-stabilization settings with alternative actor heads: a Neural Addition Unit (NAU) with a Neural Multiplication Unit (NMU)-inspired quadratic correction, a Kolmogorov-Arnold Network (KAN), or a multilayer perceptron (MLP) with rectified linear unit (ReLU) or hyperbolic-tangent activations.