🤖 AI Summary
This study addresses the limitations of traditional static portfolio optimization in handling sequential decision-making, tail risk, and market frictions such as transaction costs. To this end, it proposes a deep reinforcement learning–based bi-objective dynamic optimization framework that jointly maximizes expected return and minimizes downside risk by integrating three risk measures—variance, Conditional Value-at-Risk (CVaR), and Entropic Value-at-Risk (EVaR)—while explicitly incorporating transaction costs and position constraints. The approach innovatively applies deep reinforcement learning to multi-objective portfolio optimization, modeling asset return uncertainty through a combination of GARCH(1,1), extreme value theory, and t-copula. Scenario generation employs quasi-Monte Carlo simulation, and the policy is trained using the Proximal Policy Optimization (PPO) algorithm. Empirical validation on equity index data from ten countries across pre-, mid-, and post-pandemic periods demonstrates that the proposed method significantly outperforms benchmarks such as NSGA-II in terms of risk–return trade-off, control of extreme downside risk, and scalability to high-dimensional portfolios.
📝 Abstract
Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dynamics, and practical investment constraints. Existing reliability based portfolio optimization approaches primarily rely on static optimization frameworks and often fail to capture sequential decision making, tail risk, and market frictions such as transaction costs. To address these limitations, we propose a deep reinforcement learning framework for multi-objective reliability based portfolio optimization (MORP-DRL). The proposed framework jointly optimizes expected return and downside risk using three complementary risk measures: variance, Conditional Value-at-Risk (CVaR), and Entropic Value-at-Risk (EVaR). To model uncertainty and heavy-tailed market behavior, asset returns are represented using GARCH(1,1), Extreme Value Theory, and a t-copula dependence structure, while realistic scenarios are generated through quasi-Monte Carlo simulation. A Proximal Policy Optimization (PPO) based strategy is developed under practical constraints including transaction costs and portfolio bounds, and is benchmarked against NSGA-II. Experiments on ten global equity indices across pre-COVID, COVID, and post-COVID market regimes demonstrate that MORP-DRL achieves competitive risk-return performance, reduced downside risk during periods of market stress, and scalability to high-dimensional portfolio settings.