🤖 AI Summary
This study addresses the overreliance of existing AI trading strategies on historical backtesting, which limits the validation of their generalizability under real-world market frictions and future conditions. We propose a progressive realism evaluation protocol that establishes a unified three-stage benchmark encompassing backtesting, paper trading, and live execution to systematically assess machine learning, reinforcement learning, and large language model agents in cryptocurrency trading. By quantifying performance degradation from backtesting to live deployment, this work reveals, for the first time, the differential robustness of various AI approaches under temporal extrapolation and execution frictions. Furthermore, we open-source both the evaluation framework and the live trading platform interfaces, providing a reproducible quantitative foundation for assessing the real-world validity of AI-driven trading systems.
📝 Abstract
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at https://github.com/Starlien95/Awesome-TradingAI.