Evaluation of Deep Reinforcement Learning Algorithms for Portfolio Optimisation

📅 2023-07-15
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates the applicability of deep reinforcement learning (DRL) to portfolio optimization under market impact and dynamic regime switching. We construct a simulated trading environment based on geometric Brownian motion and the Bertsimas–Lo market impact model, with the Kelly criterion as the objective function. First, we systematically reveal DRL’s high sensitivity to reward noise—a previously underexplored challenge. Second, we propose PPO-GAE+HMM, a novel framework integrating proximal policy optimization with generalized advantage estimation and a hidden Markov model for latent market-state inference and adaptive policy adjustment, achieving a 27% return improvement in multi-regime settings. Third, we empirically validate that the PPO clipping mechanism is critical for policy stability. In static environments, PPO-GAE asymptotically approaches the theoretical optimum (error <3%), yet suffers from low sample efficiency, requiring over two million training steps for convergence.
📝 Abstract
We evaluate benchmark deep reinforcement learning (DRL) algorithms on the task of portfolio optimisation under a simulator. The simulator is based on correlated geometric Brownian motion (GBM) with the Bertsimas-Lo (BL) market impact model. Using the Kelly criterion (log utility) as the objective, we can analytically derive the optimal policy without market impact and use it as an upper bound to measure performance when including market impact. We found that the off-policy algorithms DDPG, TD3 and SAC were unable to learn the right Q function due to the noisy rewards and therefore perform poorly. The on-policy algorithms PPO and A2C, with the use of generalised advantage estimation (GAE), were able to deal with the noise and derive a close to optimal policy. The clipping variant of PPO was found to be important in preventing the policy from deviating from the optimal once converged. In a more challenging environment where we have regime changes in the GBM parameters, we found that PPO, combined with a hidden Markov model (HMM) to learn and predict the regime context, is able to learn different policies adapted to each regime. Overall, we find that the sample complexity of these algorithms is too high, requiring more than 2m steps to learn a good policy in the simplest setting, which is equivalent to almost 8,000 years of daily prices.
Problem

Research questions and friction points this paper is trying to address.

Evaluating deep reinforcement learning for portfolio optimization
Assessing algorithm performance with noisy rewards and market impact
Overcoming high sample complexity in real-world applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Use correlated Brownian motion for data simulation
Apply PPO with generalized advantage estimation
Combine PPO with hidden Markov model
🔎 Similar Papers
No similar papers found.