Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过FWBench工具评估了语言模型在成本限制下选择和使用时间序列预测进行决策的能力,测试了包括小型语言模型在内的十种配置。
📝 Abstract
Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.
Problem

Research questions and friction points this paper is trying to address.

Time-series foundation models
decision quality
forecast cost
cost constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

FWBench
Time-series forecasts
Decision quality
Cost constraints
Language models
🔎 Similar Papers
S
Shunya Nagashima
Neurogica Inc.