Score
Designs and implements simulation platforms and benchmark frameworks that generate repeatable, fixed-day scenarios and synthetic time-series inputs (e.g., wind or price signals) under controlled random seeds; these systems support configurable experimental factors such as delayed task-completion feedback and extensible test scenarios. Builds reproducible test harnesses and benchmarks that enable methodical, comparable evaluation of online controllers and other algorithms.
为解决风电预测在系统评估方面的不足,本文提出WPBench,一个综合性的基准测试平台,通过整合26个公开数据集和19种代表性模型来全面、公平地评估风电预测方法。
This study addresses the frequent failure of time-series foundation models in predictive control due to their inability to accurately model responses to intervention actions. Using heat pump control as a representative scenario, we conduct closed-loop experiments integrating model predictive control with zero-shot forecasting. Our findings reveal a critical insight: low prediction error does not necessarily translate to effective control performance. Furthermore, we establish contextual excitation as a fundamental prerequisite for achieving model controllability. Specifically, sufficient contextual excitation is essential for recovering input-response relationships and ensuring system controllability. Preliminary closed-loop evaluations further demonstrate the practical potential of short context windows in real-world control applications.
This study addresses the limitations of existing time series forecasting benchmarks, which evaluate a narrow scope and fail to capture structured anomalies or system-level failures in real-world scenarios. This challenge is further compounded by the unauditable pretraining data of foundation models, which impedes rigorous generalization assessment. To overcome these issues, this work proposes a scenario-based stress testing framework that transcends conventional noise perturbation by incorporating causal and system-level failure modeling. By integrating semantic scenarios, explicit failure operators, and graded difficulty levels, the framework jointly evaluates historical inputs, future targets, and multidimensional failure metrics. Ultimately, this research establishes a novel paradigm combining attribution capability with deployment relevance, effectively revealing model failure mechanisms under specific operational conditions.
Existing market simulators struggle to simultaneously ensure controllability, plausibility, and cross-market/multi-frequency adaptability of synthetic financial data, hindering quantitative model development and robust evaluation. To address this, we propose the Retrieval-Augmented Financial Market Simulator (RA-FMS), the first framework integrating macro-level trend modeling and micro-level agent behavior via a retrieval-augmented diffusion architecture. RA-FMS enables causal “what-if” scenario generation and cross-market trend synthesis. We further design an automated model optimization framework grounded in simulated stress testing. Methodologically, RA-FMS unifies conditional diffusion modeling, cross-sectional information retrieval, and causal prompting for fine-grained control. Empirical results demonstrate that RA-FMS significantly enhances downstream quantitative models’ generalization under high-volatility regimes and improves stability of risk-adjusted returns. By providing interpretable, intervenable, and reproducible synthetic data, RA-FMS establishes a foundational infrastructure for trustworthy financial AI.
Time-series foundation models (TSFMs) and large language model–driven time-series models (TSLLMs) suffer from scarcity of high-quality, diverse real-world time-series data. Method: We systematically investigate the role of synthetic data in pretraining, fine-tuning, and evaluation of TSFMs/TSLLMs. We integrate generative approaches—including GANs, VAEs, diffusion models, LLM-based generation, and prompt-driven synthesis—while explicitly modeling temporal characteristics such as periodicity, abrupt changes, and multi-scale dependencies. Contribution/Results: We propose the first methodology framework for synthetic data in time-series AI, mapping generation strategies to model capability improvements. We categorize seven mainstream synthetic methods and four core application scenarios, identify six critical research gaps, and advocate a future paradigm emphasizing scalability, debiasing, and fidelity. This work delivers the first comprehensive roadmap for synthetic-data–driven time-series AI.
Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.
This work proposes the first falsifiable and reproducible synthetic experimental framework for systematically comparing the coordination efficacy of centralized planning and polycentric market mechanisms within a unified simulated economic environment. The framework integrates input-output networks, heterogeneous firms, capacity constraints, and endogenous pricing, leveraging agent-based modeling, adversarial stress testing, and structural shock analysis. Experimental results demonstrate that computational planners consistently achieve lower welfare losses across training, holdout, and adversarial scenarios, thereby validating the framework’s effectiveness. This approach establishes a methodological prototype for empirical calibration and mechanism design research in comparative economic systems.
本文介绍了CityLearn v3,一个可配置的仿真和评估框架,用于在真实条件下研究可再生能源社区(REC)的控制问题,包括成员、资产变化及数据质量等。
研究通过FWBench工具评估了语言模型在成本限制下选择和使用时间序列预测进行决策的能力,测试了包括小型语言模型在内的十种配置。
This study addresses the limitation of existing generative model benchmarks, which focus solely on code compilation or structural similarity and fail to verify engineering compliance. We construct a text-to-executable Simulink model generation benchmark spanning ten domains and propose an executable system archive-based method for aligning with engineering specifications. Furthermore, we design a hierarchical automated evaluation mechanism alongside a six-dimensional dynamic response metric system, enabling end-to-end assessment from deliverability and executability to engineering qualification. Experimental results demonstrate that the best-performing agent achieves a score of only 42.86, confirming a significant disconnect between structural similarity and engineering performance. These findings reveal critical capability bottlenecks in current large language models regarding the generation of qualified systems that satisfy multidimensional engineering requirements.