🤖 AI Summary
This study addresses the high computational cost of heavy architectures in temporal flow matching models and the insufficient exploration of linear latent space mappings with respect to prior distributions. We propose a lightweight flow matching method based on a kernel-induced latent space. By leveraging Mercer feature bases for linear mapping, time series are embedded into low-dimensional vectors via invertible transformations, decoupling training data and accommodating non-stationary priors. A multilayer perceptron then efficiently learns the resulting flow distribution. Evaluated across five benchmark datasets, our approach achieves Continuous Ranked Probability Score (CRPS) performance comparable to or exceeding that of TSFlow, while reducing training memory consumption by approximately 4.7× and per-epoch training time by 3.5–4.4×.
📝 Abstract
Recent work has shown that probabilistic flow matching for time series forecasting benefits from a data-matched prior. The resulting prior introduces local correlations, which a sequential architecture usually absorbs: a recurrent neural network (RNN), a structured state-space model (S4), or a Transformer. However, such a backbone costs GPU memory and time per epoch. A cheaper alternative is MLP-based latent-space flow matching: embed the time series via an invertible map to a single latent vector and learn the flow there, so a tabular MLP can treat the series as a set of features. The relationship between the prior and the choice of linear latent map is understudied in conditional flow matching (CFM) forecasting, yet we found it strongly affects performance. Fixed transforms such as Fourier or discrete cosine (DCT) are only well-conditioned for Ornstein--Uhlenbeck priors, while a principal-component (PCA) map fit to the data is a strong but training-set-dependent reference sensitive to train--test shift. Instead, we propose to use the Mercer eigenbasis of the prior kernel: it diagonalises the centred covariance exactly, decouples from training data, and adapts to non-stationary and periodic priors. On five GluonTS benchmarks (ETTh1, ETTh2, Weather, Electricity, Traffic) under a shared protocol with TSFlow, the resulting MLP matches or beats it on CRPS at about $4.7\times$ less training memory and $3.5\times$--$4.4\times$ less time per epoch.