Score
Measuring and evaluating investment or trading strategies after accounting for risk, transaction costs, and practical constraints to determine economically meaningful alpha. This includes constructing and testing metrics and experiments (e.g., delta-hedged returns, calendar spreads) that reflect real-world frictions and statistical uncertainty.
Traditional risk-adjusted return metrics (e.g., Sharpe ratio) suffer from poor robustness in dynamic markets and limited predictive power for future performance. To address this, we propose AlphaSharpe—a novel framework that implicitly encodes financial domain knowledge into large language models (LLMs), then integrates evolutionary operators (crossover/mutation) with a generalization-aware evaluation mechanism to automatically evolve more robust and forward-looking performance metrics. Evaluated on real-world financial time series, AlphaSharpe achieves three key advances: (1) the evolved metrics exhibit threefold improvement in out-of-sample performance prediction accuracy; (2) portfolio backtests demonstrate twofold gains in empirical performance, significantly outperforming the Sharpe ratio; and (3) the proposed generalization scoring mechanism ensures cross-regime stability and metric transferability. Altogether, AlphaSharpe establishes an interpretable, iterative, and intelligent paradigm for automated, knowledge-informed metric discovery in asset management.
This study addresses the frequent failure of quantitative trading strategies when transitioning from backtesting to live trading, often due to overfitting, selection bias, and shifts in market regimes. To mitigate these issues, the authors propose a three-stage robust evaluation framework: first identifying stable parameter regions in-sample, then applying walk-forward analysis (WFA) with rolling windows and strict information isolation, and finally locking parameters in out-of-sample testing with no further optimization allowed. The framework innovatively incorporates a defense-in-depth mechanism—featuring cliff vetoes, circuit breakers, and strategy fuses—and employs multi-objective optimization to reveal the trade-off between Sharpe ratio and maximum drawdown. Empirical results on USDJPY M5 data demonstrate the framework’s effectiveness in detecting overfitting, with four Alpha strategy variants exhibiting rank reversals under different optimization objectives, thereby highlighting the inherent tension between risk-adjusted returns and tail-risk control.
This study addresses the limitation of conventional e-commerce A/B tests, which often overlook the long-term impact of interventions on profitability across an inventory item’s full lifecycle due to short experimental windows. To overcome this, the authors propose Stock Lifetime Value (SLV), a novel metric that aggregates the expected profit of current inventory over its entire sales horizon within short-term experiments, thereby enabling more accurate assessment of long-term profitability. SLV uniquely integrates inventory constraints and seasonal lifecycle dynamics into the A/B testing framework, combining causal inference with financial mapping to support both item-level and user-level experimentation while aligning with annual financial reporting. Empirical validation at Zalando demonstrates that SLV effectively predicts actual profits over an 18-month horizon, enhances pricing algorithm performance, and delivers interpretable estimates of annual financial impact.
Traditional alpha factor mining—via manual construction or genetic programming—faces inherent limitations in expressing researchers’ qualitative intent, ensuring interpretability, and guaranteeing empirical validity. Method: This paper proposes a human-in-the-loop alpha mining paradigm grounded in large language models (LLMs). Through structured prompt engineering, it translates quantitative researchers’ domain intuition into executable instructions, enabling an interactive factor synthesis framework with multi-round feedback and end-to-end alpha generation. A dedicated backtesting validation module is integrated to enforce interpretability and robustness. Contribution/Results: Experiments across multiple markets and time horizons demonstrate that the generated alphas achieve significantly higher information ratios (IR) than baseline methods, while exhibiting superior generalization capability and empirical robustness.
To address the poor stability and weak adaptability of deep learning models in quantitative investing, this paper proposes an LLM-driven multi-agent collaborative framework for automated discovery and dynamic ensemble optimization of multimodal (numerical, textual, and chart-based) alpha factors. The method innovatively integrates large language models (LLMs), multi-agent systems, and a dynamic weight gating mechanism to establish a market-state-aware, adaptive strategy generation paradigm. Empirically evaluated on the Chinese A-share market, the framework significantly outperforms state-of-the-art baselines: it achieves a 23.6% improvement in Sharpe ratio and a 31.2% reduction in maximum drawdown, effectively balancing return enhancement and risk control. By enabling interpretable, robust, and adaptive decision-making, the proposed framework establishes a novel paradigm for AI-powered quantitative investment.
This study addresses the long-standing lack of systematic measurement of “implementation risk” in quantitative investment backtesting—the performance discrepancies arising from differences in backtesting engine implementations. The work formally defines this risk for the first time and proposes four metrological metrics alongside a taxonomy of five failure modes. These are derived from parallel execution of 15 benchmark strategies across five open-source backtesting engines, incorporating transaction cost modeling, non-overlapping stratified asset buckets, and source code defect analysis. Experiments reveal that while engine outputs converge under zero-cost assumptions, performance divergence can reach up to 3.71% when transaction costs are introduced. Crucially, however, the relative ranking of strategy efficacy remains unchanged across engines (conclusion stability index = 1), indicating that implementation risk affects performance attribution but does not alter investment decisions.
Traditional risk metrics struggle to capture the structural fragility of systematic investment strategies under shifts in market regimes. This study proposes a “Minimum Regime Performance” (MRP) framework that, for the first time, quantifies strategy decay risk as a measurable indicator by evaluating the lowest risk-adjusted return a strategy achieves across distinct historical market regimes. By delineating historical regimes and conducting cross-regime performance analysis, the research validates the efficacy of MRP across a broad sample of factor-based strategies. The findings reveal a trade-off between long-term Sharpe ratios and strategy resilience, offering investors a practical tool to identify and manage the risk of performance deterioration stemming from regime transitions.
This study addresses the widespread reliance by institutional investors on marketing-driven backtests to evaluate structured investment strategies, despite concerns about their out-of-sample validity. Leveraging a global dataset of 1,726 structured products issued by institutional firms, the paper systematically examines the translatability of backtested performance into live results through peer benchmarking, macro-factor regime identification, and direct comparison between backtested and realized returns. The findings reveal that backtested returns are primarily driven by common factor conditions prevailing prior to strategy inception rather than genuine strategy-specific skill. Notably, products launched following periods of strong factor performance exhibit significantly degraded out-of-sample results. Building on these insights, the study proposes a novel evaluation framework that dynamically adjusts the credibility assigned to backtests based on the extremity of prevailing factor conditions at product launch, offering a more prudent approach to assessing structured strategies.
This study addresses the common oversight in existing reinforcement learning trading environments, which often neglect or oversimplify transaction costs, leading to strategies that fail in real-world deployment. Building upon the Almgren-Chriss framework and the square-root market impact law, this work proposes three open-source, Gymnasium-compatible trading environments that, for the first time, systematically incorporate empirically validated nonlinear market impact models. These environments support modular cost structures, exponentially decaying permanent impact, and fine-grained logging. Integrated with FinRL-Meta extensions and Optuna-based hyperparameter optimization, five state-of-the-art deep reinforcement learning algorithms are evaluated on NASDAQ-100 data. Results demonstrate that adopting the proposed model reduces average daily trading costs from $200,000 to $8,000 and turnover from 19% to 1%; hyperparameter optimization further cuts costs by up to 82%, with algorithm performance shown to be highly sensitive to the fidelity of cost modeling.
This study evaluates whether equity anomaly strategies proposed in the academic literature since the 21st century retain practical profitability when applied to non-microcap portfolios. Through a systematic examination of approximately 200 long–short anomaly portfolios—incorporating extensive backtesting, subsample analyses across time periods and market capitalization groups, and zero-investment portfolio frameworks—the analysis rigorously accounts for transaction costs and sample selection biases. The large-scale empirical investigation reveals, for the first time, that post-2005, these anomalies generate an average monthly return of only seven basis points in non-microcap stocks, which effectively vanishes after appropriate adjustments. These findings suggest that the public equity market has become highly efficient, thereby challenging conventional views on the efficacy of factor-based investing.