Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing expert guidance with autonomous exploration in on-policy reinforcement learning, where existing methods often converge to suboptimal solutions due to fragmented optimization objectives. To overcome this, we propose an adaptive expert-guidance mechanism that treats the expert intervention weight as a learnable parameter, jointly optimized with the policy under a unified on-policy objective. Built upon Proximal Policy Optimization (PPO) and an alternating control framework, our approach enables automatic decay of expert influence without requiring auxiliary components or complex scheduling heuristics. Extensive evaluations across 34 tasks demonstrate significant improvements in sample efficiency over strong baselines, with notably low hyperparameter sensitivity. Crucially, the expert weight naturally diminishes to zero as performance improves, facilitating a smooth transition from reliance on expert demonstrations to independent policy execution, ultimately yielding policies that surpass the guiding expert.
📝 Abstract
With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. The challenge then becomes balancing expert guidance against learning from rewards. Existing methods set the expert's influence through a blending weight, a schedule, or an evaluation-driven curriculum. Alternatively, they adapt it with additional learned components such as critics over expert actions or auxiliary agents. However, none optimizes it using the same on-policy objective as the policy itself. We propose a method in which the learner and the expert alternate control within each training episode, and the expert's share of control is a single learnable parameter optimized jointly with the policy. The learner benefits from the expert early in training, but its share of control declines as the learner becomes more competent, until eventually vanishing completely. This leaves the learner acting alone and better than the suboptimal expert. We evaluate our method on 34 tasks across two benchmarks, spanning discrete and continuous action spaces, using both learned and model-based experts. Our method improves sample efficiency over guided and unguided baselines while requiring minimal hyperparameter variation. The expert's share decays to zero as the learner improves, vanishing when the expert is no longer useful.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Reinforcement Learning
Expert Guidance
Sample Efficiency
Suboptimal Expert
Control Balancing
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Reinforcement Learning
Adaptive Expert Guidance
Sample Efficiency
Joint Optimization
Learnable Control Parameter
🔎 Similar Papers
No similar papers found.