Arithmetic Actor Heads and Training Stabilization for Out-of-Distribution Reinforcement Learning

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of reinforcement learning policies under out-of-distribution (OOD) magnitude shifts by proposing a Causal Symbolic BAPR (CS-BAPR) framework. Methodologically, it integrates the Soft Actor-Critic algorithm with Bayesian robust estimation to establish six training stabilization configurations. Furthermore, Neural Arithmetic Units (NAU) and Kolmogorov-Arnold Networks (KAN) are introduced as actor heads, achieving architectural innovation through a neural additive-multiplicative correction mechanism. By leveraging these components, this work effectively mitigates policy deterioration in OOD scenarios, significantly enhancing both model generalization capability and training stability.
📝 Abstract
Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The practical method combines six training-stabilization settings with alternative actor heads: a Neural Addition Unit (NAU) with a Neural Multiplication Unit (NMU)-inspired quadratic correction, a Kolmogorov-Arnold Network (KAN), or a multilayer perceptron (MLP) with rectified linear unit (ReLU) or hyperbolic-tangent activations.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Out-of-Distribution
Magnitude Shifts
Policy Deterioration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Out-of-Distribution Reinforcement Learning
Arithmetic Actor Heads
Neural Addition Unit
Kolmogorov-Arnold Network
Training Stabilization
🔎 Similar Papers
No similar papers found.