Role-Adaptive Policy Optimization for Offline Reinforcement Learning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the optimization imbalance in offline reinforcement learning arising from the coupling of policy execution and value guidance roles. We propose RAPO (Role-based Adaptive Policy Optimization), a framework that adaptively modulates update coefficients by distinguishing between execution and bootstrapping roles, effectively decoupling TD3+BC and optimizing the temperature parameter in IQL. To our knowledge, this work is the first to enable independent regulation of policy coefficients serving distinct purposes, integrating deep reinforcement learning with advantage-weighted extraction techniques to enhance existing algorithms. Extensive evaluations on the D4RL benchmark demonstrate that RAPO significantly outperforms mainstream baselines, achieving particularly notable performance improvements among TD3+BC variants.
📝 Abstract
Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.
Problem

Research questions and friction points this paper is trying to address.

Offline Reinforcement Learning
Policy Regularization
Critic Bootstrapping
Role Coupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Reinforcement Learning
Role-Adaptive Policy Optimization
Policy Regularization
Actor Decoupling
Differentiable Policy Update
🔎 Similar Papers
S
Seonvin Cho
Department of Electronic Engineering, Hanyang University
S
Soohyun Choi
Department of Electronic Engineering, Hanyang University
Songnam Hong
Songnam Hong
Hanyang University
Machine LearningInformation TheoryOptimization