build reproducible simulators

Designs and implements simulation platforms and benchmark frameworks that generate repeatable, fixed-day scenarios and synthetic time-series inputs (e.g., wind or price signals) under controlled random seeds; these systems support configurable experimental factors such as delayed task-completion feedback and extensible test scenarios. Builds reproducible test harnesses and benchmarks that enable methodical, comparable evaluation of online controllers and other algorithms.

buildreproduciblesimulators

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Financial Wind Tunnel: A Retrieval-Augmented Market Simulator

Mar 23, 2025
BC
Bokai Cao
🏛️ The Hong Kong University of Science and Technology | IDEA Research | International Digital Economy Academy

Existing market simulators struggle to simultaneously ensure controllability, plausibility, and cross-market/multi-frequency adaptability of synthetic financial data, hindering quantitative model development and robust evaluation. To address this, we propose the Retrieval-Augmented Financial Market Simulator (RA-FMS), the first framework integrating macro-level trend modeling and micro-level agent behavior via a retrieval-augmented diffusion architecture. RA-FMS enables causal “what-if” scenario generation and cross-market trend synthesis. We further design an automated model optimization framework grounded in simulated stress testing. Methodologically, RA-FMS unifies conditional diffusion modeling, cross-sectional information retrieval, and causal prompting for fine-grained control. Empirical results demonstrate that RA-FMS significantly enhances downstream quantitative models’ generalization under high-volatility regimes and improves stability of risk-adjusted returns. By providing interpretable, intervenable, and reproducible synthetic data, RA-FMS establishes a foundational infrastructure for trustworthy financial AI.

Enhances downstream model performance in volatile market conditionsGenerates controllable synthetic financial data for model testingIntegrates macro and micro market patterns via retrieval-augmented diffusion

Empowering Time Series Analysis with Synthetic Data: A Survey and Outlook in the Era of Foundation Models

Mar 14, 2025
XL
Xu Liu
🏛️ Salesforce AI Research | National University of Singapore | Squirrel Ai Learning | Hong Kong University of Science and Technology

Time-series foundation models (TSFMs) and large language model–driven time-series models (TSLLMs) suffer from scarcity of high-quality, diverse real-world time-series data. Method: We systematically investigate the role of synthetic data in pretraining, fine-tuning, and evaluation of TSFMs/TSLLMs. We integrate generative approaches—including GANs, VAEs, diffusion models, LLM-based generation, and prompt-driven synthesis—while explicitly modeling temporal characteristics such as periodicity, abrupt changes, and multi-scale dependencies. Contribution/Results: We propose the first methodology framework for synthetic data in time-series AI, mapping generation strategies to model capability improvements. We categorize seven mainstream synthetic methods and four core application scenarios, identify six critical research gaps, and advocate a future paradigm emphasizing scalability, debiasing, and fidelity. This work delivers the first comprehensive roadmap for synthetic-data–driven time-series AI.

Addressing dataset challenges for time series analysis modelsExploring synthetic data for scalable and unbiased alternativesReviewing synthetic data's role in model training and evaluation

CTBench: Cryptocurrency Time Series Generation Benchmark

Aug 03, 2025
YA
Yihao Ang
🏛️ National University of Singapore | Harbin Institute of Technology

Existing time-series generation (TSG) methods lack systematic, domain-specific evaluation for cryptocurrency markets—characterized by 24/7 trading, high volatility, and rapid regime shifts. Method: We introduce CTBench, the first comprehensive, cryptocurrency-specific benchmark, covering 452 tokens and establishing a dual-task evaluation framework integrating forecasting utility and statistical arbitrage. It assesses eight models—including LSTM, GAN, VAE, diffusion models, and linear baselines—across five dimensions (forecasting accuracy, trading profitability, risk robustness, etc.) and 13 quantitative metrics. Contribution/Results: CTBench demonstrates fine-grained discriminative power across four empirically identified market regimes. It is the first to empirically reveal the trade-off between generative fidelity and real-world trading performance. By providing reproducible evaluation protocols and model performance rankings, CTBench establishes a standardized assessment foundation and practical model selection guidance for crypto-quantitative research.

Existing TSG methods fail in volatile cryptocurrency marketsNo comprehensive benchmark exists for crypto time series generationPrior work lacks crypto-specific evaluations and financial metrics

Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.

computational constraintsdomain shiftlimited supervision

Learning to Simulate: Generative Metamodeling via Quantile Regression

Nov 29, 2023
LH
L. Hong
🏛️ Fudan University | City University of Hong Kong | The Hong Kong University of Science and Technology

Conventional metamodeling is constrained by pre-specified single-output statistics (e.g., mean), limiting its applicability in real-time decision-making where arbitrary statistics must be computed on demand. Method: We propose a generative metamodeling paradigm—designed to serve as a fast surrogate for simulators—that efficiently generates random samples approximating the true conditional distribution given any input. To this end, we formally define generative metamodeling and introduce Quantile Regression-based Generative Metamodeling (QRGMM), a novel algorithm grounded in quantile regression. We provide theoretical guarantees on its conditional distribution convergence and establish its optimal convergence rate. Results: Experiments across diverse real-time decision tasks demonstrate that QRGMM significantly outperforms existing generative models: it achieves 100×–1000× faster inference while preserving high distributional fidelity, thereby overcoming the flexibility bottleneck inherent in traditional single-statistic surrogates.

Generative metamodels need to preserve conditional distributions accuratelySlow stochastic simulations hinder real-time decision-making speedTraditional metamodels limit flexibility by using single output statistics

Latest Papers

What's happening recently
View more

This work proposes the first falsifiable and reproducible synthetic experimental framework for systematically comparing the coordination efficacy of centralized planning and polycentric market mechanisms within a unified simulated economic environment. The framework integrates input-output networks, heterogeneous firms, capacity constraints, and endogenous pricing, leveraging agent-based modeling, adversarial stress testing, and structural shock analysis. Experimental results demonstrate that computational planners consistently achieve lower welfare losses across training, holdout, and adversarial scenarios, thereby validating the framework’s effectiveness. This approach establishes a methodological prototype for empirical calibration and mechanism design research in comparative economic systems.

agent-based marketcomputational planningeconomic coordination

Real-world forecasting benchmarks are often constrained by outcome delays, the rarity of tail events, and the infeasibility of evaluating counterfactuals. To address these limitations, this work proposes a simulation-based forecasting benchmark built on the turn-based strategy game Freeciv. The framework generates forecasting tasks from fixed snapshots of world states and automatically evaluates predictions against subsequent simulated trajectories. It supports continuous or binary forecasts over arbitrary time horizons, as well as conditional and causal intervention queries, enabling repeatable assessment of rare or disruptive events. The authors release a complete benchmark pipeline—including a suite of forecasting questions, a scoring mechanism, and associated datasets—and demonstrate its validity through evaluations with predictive models and an anonymous human pilot study.

counterfactual reasoningforecasting benchmarkprobabilistic reasoning

This study addresses the current lack of human-centered, interpretable, and responsible evaluation criteria for AI in modeling and simulation. The authors propose the first multidimensional benchmark framework specifically designed to assess large language models (LLMs) through a human-centric lens, leveraging an open-source system dynamics AI platform to systematically evaluate performance across qualitative modeling, quantitative modeling, and model discussion tasks—emphasizing human-AI collaboration rather than replacement. The framework incorporates critical capabilities such as causal reasoning, iterative model refinement, and behavioral explanation, while embedding ethical and accountability considerations. Empirical results indicate that existing AI tools perform relatively well in qualitative tasks and model discussions but remain limited in causal reasoning and quantitative error correction; furthermore, different LLMs exhibit distinct strengths, with no single model emerging as universally superior.

AI for Modeling and SimulationBenchmarkingBias in AI

This study addresses the cold-start forecasting challenge in newly commissioned photovoltaic (PV) power plants, where the absence of historical generation data hinders accurate prediction. To overcome this limitation, the authors propose a zero-shot forecasting approach that integrates physics-informed synthetic history with covariate-aware time series foundation models. Specifically, plant metadata and meteorological covariates are leveraged to generate physically guided synthetic time series, which serve as contextual conditioning during inference. The framework uniquely combines five state-of-the-art time series foundation models—including TabPFN-TS and Chronos-2—with this synthesis strategy and incorporates multiple feedback mechanisms. Evaluated across 440 real-world PV sites, the method substantially outperforms conventional baselines by 1.7–2× in accuracy. Notably, TabPFN-TS achieves a mean absolute error (MAE) of 0.514 kWh·kWp⁻¹·d⁻¹ under real feedback, while Chronos-2 demonstrates the most robust performance under self-feedback.

cold-startphotovoltaic forecastingsynthetic history

This work addresses the challenge of enabling multi-megawatt AI/HPC facilities to respond to grid dispatch signals within seconds while precisely regulating GPU power consumption. The authors propose a three-layer predictive control architecture that coordinates regulation across millisecond, second, and hour timescales, augmented by a deterministic safe-island bypass mechanism to achieve rapid closed-loop response from grid commands to GPU power. The study innovatively demonstrates on real hardware that AI supercomputers can serve as flexible grid loads and introduces a real-time PUE correction mechanism to ensure that scheduling commitments are honored at the metering level. Experimental results on a three-GPU V100 platform show an end-to-end response latency of 97.2 ms—6.9× faster than Nordic fast frequency reserve requirements—and reduce cooling-related efficiency penalties by 2.5–5.8 percentage points across six national grid replay scenarios.

AI supercomputersdata-center demandgrid-responsive control