🤖 AI Summary
This study investigates the efficacy and robustness of deep reinforcement learning (DRL) for multi-asset trading. To address the limitations of static, supervised-learning-based strategies, we propose a risk-aware action masking mechanism integrated with two state-of-the-art DRL algorithms—Double Deep Q-Network (DDQN) and Proximal Policy Optimization (PPO)—and evaluate their performance systematically on daily-frequency data from 2019 to 2023 across the S&P 500 index, Bitcoin, and three major FX pairs. Our work constitutes the first cross-asset comparative analysis of risk-adjusted returns between DDQN and PPO, enabling dynamic, context-sensitive decision-making that actively avoids adverse market regimes. Empirical results demonstrate that the DRL strategies achieve an average 42% improvement in annualized Sharpe ratio and a 31% reduction in maximum drawdown relative to a buy-and-hold benchmark, thereby validating their superior adaptive decision-making capability and real-world market resilience.
📝 Abstract
The paper explores the use of Deep Reinforcement Learning (DRL) in stock market
trading, focusing on two algorithms: Double Deep Q-Network (DDQN) and Proximal Policy
Optimization (PPO) and compares them with Buy and Hold benchmark. It evaluates these
algorithms across three currency pairs, the S&P 500 index and Bitcoin, on the daily data in the
period of 2019-2023. The results demonstrate DRL's effectiveness in trading and its ability to
manage risk by strategically avoiding trades in unfavorable conditions, providing a substantial
edge over classical approaches, based on supervised learning in terms of risk-adjusted returns.