🤖 AI Summary
This study addresses the limitation that existing time series and language models are predominantly developed in isolation, lacking a unified framework for comprehension and prediction. To bridge this gap, this work proposes an interleaved global residual attention mechanism that aligns pretrained language models with time series foundation models into a shared representation space, thereby fusing bimodal knowledge. Furthermore, a unified prompting scheme and a stable joint training strategy are designed, which, combined with large-scale instruction-tuning data, effectively balance the model's understanding and generation capabilities. The proposed approach achieves unified modeling of time series and language, demonstrating performance on multiple benchmarks comparable to that of larger general-purpose models as well as task-specific counterparts.
📝 Abstract
We present TimeBraid, a series of unified time-series and language models that align pretrained language models and pretrained time-series foundation models through interleaved global residual attention layers. Each model inherits knowledge, instruction following, and reasoning from one side, continuous-signal perception and zero-shot forecasting from the other, and fuses the two in a shared representation space where both modalities are understood and generated. We study the design choices that make such unified modeling work: where to align the two representation spaces, how to ground language in temporal structure, how to balance understanding with generation, and how to keep joint optimization stable. The resulting recipe combines a unified prompting scheme for diverse time-series and text tasks, stabilized joint training, and supervision from 2.2M curated series--text pairs and 4.9M instruction-tuning samples. Across benchmarks spanning time-series perception, understanding, reasoning, and both context-aided and unimodal forecasting, TimeBraid remains competitive with far larger general-purpose models and task-specific counterparts.