end-to-end portfolio learning

Designs and trains parametric allocation policies that map market-state features to portfolio weights, implementing end-to-end policy-based optimization to directly optimize a portfolio objective. Builds parametric portfolio rules (e.g., neural or linear policies) that learn allocation rules from data, generalize allocations across assets, and are evaluated versus simple rule benchmarks.

end-to-endportfoliolearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of constructing end-to-end cross-asset futures timing strategies that outperform traditional rule-based approaches. To overcome the limitations of conventional two-stage paradigms—separate prediction followed by portfolio optimization—the authors propose a unified framework that directly maps market states to portfolio weights. They implement this approach using LSTM and Transformer architectures, innovatively training the models with a differentiable Sharpe ratio as the loss function. Empirical evaluation on 16 highly liquid CME futures contracts demonstrates that the proposed Transformer-based strategy significantly outperforms standard benchmarks—including equal-weight, risk parity, and time-series momentum portfolios—in out-of-sample tests. Notably, the strategy achieves superior performance with lower trading frequency and maintains robustness under moderate transaction costs.

AI modelscross-asset futuresend-to-end learning

This study addresses the neglect of parameter uncertainty in traditional parametric portfolio strategies, which leads to overestimated expected utility and underestimated risk. It introduces a Bayesian approach into the parametric framework by placing prior distributions on strategy coefficients, thereby constructing a Bayesian mean-variance optimization model that explicitly accounts for estimation risk. The decision rule is further refined to incorporate posterior uncertainty. Theoretical analysis reveals that utility loss is positively related to both posterior uncertainty and signal strength. Empirical results based on 242 signals and six factors from 1973 to 2023 demonstrate that the proposed method significantly improves the Sharpe ratio, reduces portfolio turnover and tail risk, and yields monotonically increasing investor welfare with risk aversion, exhibiting particularly robust performance during financial crises.

Bayesian decisionestimation riskexpected utility

This study investigates the effective trade-off between return and risk in portfolio allocation by conducting the first fair and systematic empirical comparison between deep reinforcement learning (DRL) and the classical mean-variance optimization (MVO) framework within a unified experimental setup. The authors develop a model-free DRL agent and evaluate its performance through backtesting on historical market data, while implementing and adapting MVO as a benchmark. Results demonstrate that DRL significantly outperforms MVO across key metrics—including Sharpe ratio, maximum drawdown, and absolute returns—thereby validating its practical advantages and application potential in dynamic asset allocation. This work provides a novel methodological foundation for intelligent investment research and decision-making.

Asset AllocationDeep Reinforcement LearningMean-Variance Optimization

Reinforcement-Learning Portfolio Allocation with Dynamic Embedding of Market Information

Jan 29, 2025
JH
Jinghai He
🏛️ University of California at Berkeley | Shanghai Jiao Tong University

To address the insufficient robustness of investment decisions in financial markets—characterized by high dimensionality, non-stationarity, and low signal-to-noise ratios—this paper proposes a reinforcement learning (RL) portfolio optimization framework integrating dynamic state embedding and online meta-learning. The framework employs a generative autoencoder to compress the temporal state space and disentangles time-series features to enhance representation stability. Crucially, it pioneers the coupling of online meta-learning with RL policy updates, enabling automatic risk timing and dynamic position adjustment under market volatility. Empirical evaluation on S&P 500 constituents shows that the model achieves significantly higher Sharpe ratios compared to equal-weighted (EW), mean-variance (MV), and baseline RL portfolios, with gains up to 37% during stress periods. Ablation studies confirm both cross-algorithm robustness and the necessity of each core module.

Investment Decision MakingMarket VolatilityPortfolio Allocation

This paper investigates the applicability of deep reinforcement learning (DRL) to portfolio optimization under market impact and dynamic regime switching. We construct a simulated trading environment based on geometric Brownian motion and the Bertsimas–Lo market impact model, with the Kelly criterion as the objective function. First, we systematically reveal DRL’s high sensitivity to reward noise—a previously underexplored challenge. Second, we propose PPO-GAE+HMM, a novel framework integrating proximal policy optimization with generalized advantage estimation and a hidden Markov model for latent market-state inference and adaptive policy adjustment, achieving a 27% return improvement in multi-regime settings. Third, we empirically validate that the PPO clipping mechanism is critical for policy stability. In static environments, PPO-GAE asymptotically approaches the theoretical optimum (error <3%), yet suffers from low sample efficiency, requiring over two million training steps for convergence.

Assessing algorithm performance with noisy rewards and market impactEvaluating deep reinforcement learning for portfolio optimizationOvercoming high sample complexity in real-world applications

Latest Papers

What's happening recently
View more

This study addresses the neglect of parameter and return uncertainty in traditional portfolio optimization by proposing a Distributional Portfolio Optimization (DPO) framework. The approach unifies portfolio weights, asset returns, and model parameters under a joint probability measure, integrating Bayesian inference, distributionally robust optimization (DRO), chance constraints, and distributional reinforcement learning. Key theoretical contributions include a Wasserstein–CVaR duality, a no-randomization theorem, Bayesian credible-radius-calibrated Wasserstein DRO, Gaussian conservative bounds, and a distributional Bellman contraction property under risk translation. Empirically, the method achieves near-oracle tail risk performance in factor models—exceeding it by only 3–7 basis points—without requiring a validation set. However, in out-of-sample backtests on the DJIA, it does not significantly outperform benchmark strategies such as equal weighting, Black–Litterman, or Ledoit–Wolf in terms of Sharpe ratio.

Distributional Portfolio OptimizationFinancial Risk ManagementPortfolio Optimization

This work addresses the challenge of accurately identifying optimal policies and active constraints when both are unknown. It proposes a unified simulation-grid-based dual framework that integrates Fenchel duality, Doob martingale compensation, complementary slackness, and occupancy-measure weighting. By decomposing residuals via conditional budget identities and leveraging Bellman curvature, the method constructs a tight policy region without requiring a reference solution. It simultaneously estimates policy error, certifies active constraint facets, and provides joint verification of value bounds and policy distance. Empirical results demonstrate its ability to achieve full coverage in auditing external policy errors, deliver zero false positives in active facet detection, maintain tightness in 50-dimensional asset stress tests, and reveal that dual-learning accuracy becomes the performance bottleneck in high dimensions.

binding constraintsconstrained dynamic portfoliosduality

Traditional feature importance methods struggle to interpret combinatorial investment decisions where prediction and optimization are tightly coupled, and they fail to elucidate how macroeconomic conditions influence investment outcomes. This work proposes a novel prediction–optimization–explanation framework that, for the first time, integrates gradient-guided counterfactual sample generation with portfolio optimization to construct economically meaningful “what-if” scenarios. By jointly modeling the prediction and optimization processes, the method flexibly generates macroeconomic scenarios tailored to specific investment objectives, effectively identifying critical conditions—such as those that narrow strategy return gaps, trigger diversification, or enable excess returns. Empirical results demonstrate that the proposed framework substantially enhances both the interpretability and robustness of portfolio strategies.

decision pipelinesexplainabilitymacroeconomic conditions

研究通过强化学习选择优化器的方法,解决不同问题和阶段下最优优化方法的选择问题,采用序列决策制定策略以适应性选择最合适的优化器。

Derivative-free OptimizersGradient-based OptimizersOptimizers

This work proposes FPILOT, a novel framework that introduces model predictive control into financial reinforcement learning to overcome the limitations of static trading policies. Existing reinforcement learning agents typically employ fixed strategies during inference and cannot dynamically adapt based on price forecasts. In contrast, FPILOT leverages multi-step unconditional price predictions to construct return targets and performs real-time optimization of any pretrained policy at inference time—without requiring retraining. The approach is particularly effective in enhancing stochastic policies and achieves significant improvements in both cumulative returns and risk-adjusted performance metrics—including Sharpe, Sortino, and Calmar ratios—on the TradeMaster DJ30 benchmark. Moreover, the gains scale consistently with the quality of the underlying price predictions.

inference-time optimizationportfolio managementprice forecasting

Hot Scholars

LA

Laura Alessandretti

Technical University of Denmark
Complex NetworksHuman behaviorComputational Social Science
QW

Qiuqi Wang

Assistant Professor, Georgia State University
Quantitative Risk Management
TJ

Tim J. Boonen

University of Hong Kong
Actuarial sciencemathematical economicsmathematical finance
YC

Yuyu Chen

GSM, Peking university
ECONOMICS