🤖 AI Summary
Existing time-series generation (TSG) methods lack systematic, domain-specific evaluation for cryptocurrency markets—characterized by 24/7 trading, high volatility, and rapid regime shifts. Method: We introduce CTBench, the first comprehensive, cryptocurrency-specific benchmark, covering 452 tokens and establishing a dual-task evaluation framework integrating forecasting utility and statistical arbitrage. It assesses eight models—including LSTM, GAN, VAE, diffusion models, and linear baselines—across five dimensions (forecasting accuracy, trading profitability, risk robustness, etc.) and 13 quantitative metrics. Contribution/Results: CTBench demonstrates fine-grained discriminative power across four empirically identified market regimes. It is the first to empirically reveal the trade-off between generative fidelity and real-world trading performance. By providing reproducible evaluation protocols and model performance rankings, CTBench establishes a standardized assessment foundation and practical model selection guidance for crypto-quantitative research.
📝 Abstract
Synthetic time series are essential tools for data augmentation, stress testing, and algorithmic prototyping in quantitative finance. However, in cryptocurrency markets, characterized by 24/7 trading, extreme volatility, and rapid regime shifts, existing Time Series Generation (TSG) methods and benchmarks often fall short, jeopardizing practical utility. Most prior work (1) targets non-financial or traditional financial domains, (2) focuses narrowly on classification and forecasting while neglecting crypto-specific complexities, and (3) lacks critical financial evaluations, particularly for trading applications. To address these gaps, we introduce extsf{CTBench}, the first comprehensive TSG benchmark tailored for the cryptocurrency domain. extsf{CTBench} curates an open-source dataset from 452 tokens and evaluates TSG models across 13 metrics spanning 5 key dimensions: forecasting accuracy, rank fidelity, trading performance, risk assessment, and computational efficiency. A key innovation is a dual-task evaluation framework: (1) the emph{Predictive Utility} task measures how well synthetic data preserves temporal and cross-sectional patterns for forecasting, while (2) the emph{Statistical Arbitrage} task assesses whether reconstructed series support mean-reverting signals for trading. We benchmark eight representative models from five methodological families over four distinct market regimes, uncovering trade-offs between statistical fidelity and real-world profitability. Notably, extsf{CTBench} offers model ranking analysis and actionable guidance for selecting and deploying TSG models in crypto analytics and strategy development.