🤖 AI Summary
This work addresses the common disconnect between forecasting and interpretability in existing time series models, which typically fail to jointly generate accurate numerical predictions and verifiable causal explanations within a unified framework. To bridge this gap, we propose ReasonCast, a novel approach that fine-tunes large language models to enable end-to-end, simultaneous generation of forecasts and self-explanations through a single autoregressive pass, producing both predictive outputs and coherent reasoning chains. We introduce ReasonTS-Bench, the first benchmark dataset tailored for this joint task, and develop a multi-task training paradigm to support it. Experimental results demonstrate that ReasonCast outperforms both conventional time series models and general-purpose large language models in prediction accuracy while generating high-quality, logically sound, and verifiable textual explanations.
📝 Abstract
Most time series (TS) models are specialized for a single task, either understanding (i.e., returning text answers about a TS) or generation (i.e., returning a numeric forecast). Only recently have unified models begun to handle the two within a single architecture. Even these models, however, produce the two outputs as task-separated paths and cannot predict a series and explain why that prediction arises within a single coherent response. In this paper, we argue for a task-fused model that jointly produces 1) prediction (generation) and 2) selfexplanation (understanding), thereby integrating 1) numerical TS forecasting and 2) interpretable text reasoning within a single response. To enable the systematic study of this capability, we present both a benchmark and a recipe that jointly address the two tasks. The benchmark, ReasonTS-Bench, identifies five fundamental patterns underlying TS and enables the joint evaluation of both tasks. ReasonCast, our recipe for finetuning any LLM to perform both tasks jointly, yields a model that generates a reasoning chain and a forecast together in a single autoregressive pass. Extensive experiments show that ReasonCast outperforms both LLMs and TS models on prediction accuracy while producing verifiable, causal reasoning. Code is available at: https://github.com/seunghan96/reasoncast.