Experience-Efficient Model-Free Deep Reinforcement Learning Using Pre-Training

📅 2025-10-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the low sample efficiency and high training cost of reinforcement learning in physics simulation environments, this paper proposes PPOPT: a pretraining-based, model-agnostic Proximal Policy Optimization (PPO) algorithm. Its core innovation is a segmented neural network architecture wherein intermediate layers are jointly pretrained across multiple similar physical environments to acquire transferable dynamics representations; downstream tasks then require only fine-tuning of input and output layers, drastically reducing interaction samples needed in the target environment. Experiments demonstrate that PPOPT achieves faster convergence and more stable policies under extremely limited samples (<10⁵ interaction steps), outperforming standard PPO in final performance while incurring significantly lower computational overhead than model-based methods. The implementation is publicly available.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchComputer Vision: Low Level & Physics-based VisionMachine Learning: Imitation Learning & Inverse Reinforcement Learning

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSystems and Infrastructure for Web, Mobile and WoT: Experiences and lessons learnt from Web-based algorithms and system deploymentsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
We introduce PPOPT - Proximal Policy Optimization using Pretraining, a novel, model-free deep-reinforcement-learning algorithm that leverages pretraining to achieve high training efficiency and stability on very small training samples in physics-based environments. Reinforcement learning agents typically rely on large samples of environment interactions to learn a policy. However, frequent interactions with a (computer-simulated) environment may incur high computational costs, especially when the environment is complex. Our main innovation is a new policy neural network architecture that consists of a pretrained neural network middle section sandwiched between two fully-connected networks. Pretraining part of the network on a different environment with similar physics will help the agent learn the target environment with high efficiency because it will leverage a general understanding of the transferrable physics characteristics from the pretraining environment. We demonstrate that PPOPT outperforms baseline classic PPO on small training samples both in terms of rewards gained and general training stability. While PPOPT underperforms against classic model-based methods such as DYNA DDPG, the model-free nature of PPOPT allows it to train in significantly less time than its model-based counterparts. Finally, we present our implementation of PPOPT as open-source software, available at github.com/Davidrxyang/PPOPT.
Problem

Research questions and friction points this paper is trying to address.

Leveraging pretraining for efficient model-free reinforcement learning
Reducing computational costs in physics-based environments
Improving training stability with small interaction samples
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses pretrained neural network for policy architecture
Leverages transferable physics from pretraining environment
Achieves efficiency with small training samples
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruoxing Yang
Department of Computer Science, Georgetown University