LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of supervised fine-tuning quality and the difficulty of adapting reinforcement learning rewards to open-ended generation in time series description tasks. To this end, we propose LineupRL, a framework that introduces a novel verifiable reward mechanism based on description-sequence matching. Specifically, it employs a frozen large language model as a verifier to identify target data from candidate sequences for generation optimization, eliminating the need for manually designed scoring criteria while effectively resisting reward hacking. Experimental results demonstrate that the proposed method comprehensively outperforms SFT and RL baselines across multiple benchmarks. Notably, a 3B-parameter model surpasses its 72B distillation source model, achieving significant improvements in trend capture and numerical naming capabilities.
📝 Abstract
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.
Problem

Research questions and friction points this paper is trying to address.

Time Series Captioning
Supervised Fine-Tuning
Reinforcement Learning
Reward Design
Open-ended Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Verifiable Rewards
Time Series Captioning
Caption-to-Series Identification
Reward Hacking Resistance
Vision Language Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.