🤖 AI Summary
This paper addresses the S&P 500 at-the-money option hedging problem without assuming an explicit pricing model. We propose a deep reinforcement learning framework based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, which takes six-dimensional real-time market features as input and learns hedging policies end-to-end. The method incorporates a risk-aware reward function, explicit transaction cost constraints, and a walk-forward training scheme to enhance out-of-sample generalization. Evaluated over a near-17-year out-of-sample period (2004–2024) using high-frequency data, our approach significantly outperforms Black–Scholes delta hedging: it achieves an 18.3% higher annualized return and a 24.6% improvement in Sharpe ratio, while demonstrating superior robustness under high-volatility and high-transaction-cost regimes. Our key contribution is the first systematic, long-horizon empirical application of a model-free deep RL framework to options hedging—demonstrating its feasibility and superiority as a viable alternative to classical model-based approaches.
📝 Abstract
This paper explores the application of deep Q-learning to hedging at-the-money options on the S&P~500 index. We develop an agent based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, trained to simulate hedging decisions without making explicit model assumptions on price dynamics. The agent was trained on historical intraday prices of S&P~500 call options across years 2004--2024, using a single time series of six predictor variables: option price, underlying asset price, moneyness, time to maturity, realized volatility, and current hedge position. A walk-forward procedure was applied for training, which led to nearly 17~years of out-of-sample evaluation. The performance of the deep reinforcement learning (DRL) agent is benchmarked against the Black--Scholes delta-hedging strategy over the same period. We assess both approaches using metrics such as annualized return, volatility, information ratio, and Sharpe ratio. To test the models' adaptability, we performed simulations across varying market conditions and added constraints such as transaction costs and risk-awareness penalties. Our results show that the DRL agent can outperform traditional hedging methods, particularly in volatile or high-cost environments, highlighting its robustness and flexibility in practical trading contexts. While the agent consistently outperforms delta-hedging, its performance deteriorates when the risk-awareness parameter is higher. We also observed that the longer the time interval used for volatility estimation, the more stable the results.