🤖 AI Summary
This study addresses the widespread reliance by institutional investors on marketing-driven backtests to evaluate structured investment strategies, despite concerns about their out-of-sample validity. Leveraging a global dataset of 1,726 structured products issued by institutional firms, the paper systematically examines the translatability of backtested performance into live results through peer benchmarking, macro-factor regime identification, and direct comparison between backtested and realized returns. The findings reveal that backtested returns are primarily driven by common factor conditions prevailing prior to strategy inception rather than genuine strategy-specific skill. Notably, products launched following periods of strong factor performance exhibit significantly degraded out-of-sample results. Building on these insights, the study proposes a novel evaluation framework that dynamically adjusts the credibility assigned to backtests based on the extremity of prevailing factor conditions at product launch, offering a more prudent approach to assessing structured strategies.
📝 Abstract
Institutional allocators often evaluate structured strategies on the basis of marketed backtests -- hypothetical track records constructed by applying a strategy's rules to historical data prior to any live trading, also referred to as pro-forma performance. It is unclear how much of that signal survives once the strategy is actually traded. Using 1,726 commercially distributed structured strategies from ten global institutions, this paper shows that raw pro-forma performance has only limited portability into the live period and weakens sharply once live outcomes are measured relative to peer and external benchmarks. The evidence indicates that marketed backtests predominantly reflect the common factor regime present before launch rather than strategy-specific skill. Strategies launched after unusually strong bucket-factor conditions experience materially worse subsequent deterioration. For allocators, the implication is practical: backtests should be judged relative to appropriate peer benchmarks, and the discount applied to them should increase when launch occurs after an extreme factor run.