FinVerse: Financial Time-Series Benchmark

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing time series forecasting benchmarks predominantly rely on uniform error metrics to evaluate heterogeneous financial sequences, which often fail to capture a model’s practical utility in real-world investment decisions. To address this limitation, this work proposes FinVerse—the first domain-specific benchmark for financial time series forecasting—that tailors evaluation metrics to the economic semantics of each sequence. It establishes a multidimensional assessment framework encompassing 78 metrics across 11 categories and conducts a systematic evaluation of 43 state-of-the-art models on a curated set of 60,000 economically meaningful core series, drawn from a large-scale dataset comprising 116,897 sequences and 171 million observations. Empirical results demonstrate that models excelling under generic metrics do not necessarily perform well in financial decision-making contexts, thereby validating the necessity and efficacy of domain-aware evaluation.
📝 Abstract
As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful standardized comparisons, but they often evaluate heterogeneous series with uniform error-based metrics. Strong performance under such metrics does not necessarily imply that a model's forecasts will support the best real-world decisions across domains. For example, in stock forecasting, correctly predicting whether a price will rise or fall can be more directly relevant to realized returns than minimizing point-wise forecast error alone. To this end, we introduce FinVerse, a finance-domain time-series forecasting benchmark that takes a first step toward more realistic evaluation. The released FinVerse data artifact contains 116,897 financial time series with 171.1M observations, of which 60,232 series with 17.4M observations are selected as evaluated targets based on their economic relevance to financial decisions. Unlike generic forecasting benchmarks that primarily emphasize uniform point-forecast or probabilistic accuracy, FinVerse defines 11 metric families comprising 78 evaluation metrics and assigns the most appropriate evaluation metrics to each individual time series based on its underlying economic meaning. Our analysis of 43 public time-series forecasting foundation models shows that strong performance under generic forecasting criteria does not necessarily translate into useful financial forecasts. This finding highlights the need for domain-aware benchmarks that evaluate models under objectives closer to real-world decision making.
Problem

Research questions and friction points this paper is trying to address.

time-series forecasting
financial benchmark
domain-aware evaluation
decision-oriented metrics
forecasting accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

domain-aware benchmark
financial time-series forecasting
economic relevance
heterogeneous evaluation metrics
decision-oriented evaluation