KALEIDO: Input-Space Adaptation of a Vision Model for Time-Series Forecasting Through Gated Fold Geometries

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the distributional discrepancy between rendered and natural images when visual models are directly applied to time series forecasting. To bridge this gap, this work proposes leveraging rendered geometry as a controllable axis, generating folded geometric structures via periodicity detection, and introducing a convex combination gating mechanism to adapt the input space of an ImageNet-pretrained masked autoencoder. By fine-tuning only 0.05% of the parameters—specifically the LayerNorm layers—the proposed approach achieves zero-shot time series forecasting without requiring dataset-specific hyperparameters. Extensive experiments demonstrate that this method reduces mean squared error by 13% on the LTSF benchmark compared to baselines, while improving MASE and CRPS by 7.4% and 19.3%, respectively, on GIFT-Eval. These results establish an efficient and generalizable framework for cross-modal transfer in time series analysis.
📝 Abstract
Time-series foundation models buy zero-shot forecasting with large temporal corpora; a vision model needs none, since a natural image implicitly embeds the patterns a forecaster must model, and an ImageNet-pretrained masked autoencoder forecasts a series by inpainting a rendering of it. A rendered series is not a natural image, however, and closing that gap takes temporal-aware adaptation. We show that the rendering geometry - how the series is folded and drawn - is a controllable, mixable axis for it. Kaleido detects the dominant periods, renders a rule-generated set of fold geometries, combines the inpaintings with a convex per-position gate fit on validation only, and fuses the result with the zero-shot output at one fixed share, with no per-dataset hyperparameter beyond the baseline's published settings. Training only LayerNorm (0.05%), Kaleido lowers MSE by 13% against the published zero-shot baseline on LTSF and, frozen, by 6.6%; on GIFT-Eval it improves the baseline by 7.4% in MASE and 19.3% in CRPS.
Problem

Research questions and friction points this paper is trying to address.

time-series forecasting
vision model adaptation
rendering geometry
domain gap
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Input-Space Adaptation
Gated Fold Geometries
Masked Autoencoder
Time-Series Forecasting
Parameter-Efficient Fine-Tuning
🔎 Similar Papers
No similar papers found.