Score
Designs and implements synthetic cloud workload generators that produce CPU, memory and other resource-usage traces whose statistical and temporal properties match real datasets. This includes parameterizing temporal workload characteristics, controlling distributional and autocorrelation properties, and exporting workloads for simulation and evaluation experiments.
Cloud analytics system evaluation commonly relies on static benchmarks (e.g., TPC-H/TPC-DS), which fail to capture key statistical characteristics of real production workloads—such as performance metric distributions, operator frequencies, and temporal query patterns. Meanwhile, existing real-world execution traces lack reproducible SQL queries and database metadata. Method: We formulate the novel problem of *statistically grounded synthetic workload generation*, introducing three core techniques: (1) multi-objective optimization–driven component selection, (2) progressive timestamp modeling, and (3) LLM-enhanced statistical fidelity augmentation—all operating via recombination of benchmark queries and database objects. Contribution/Results: Evaluated on real cloud traces, our approach reduces statistical approximation error by up to 6× over state-of-the-art methods, significantly improving evaluation authenticity, reproducibility, and ecosystem compatibility.
Existing TPC-style benchmarks (e.g., TPC-H/DS) fail to capture two critical characteristics of real-world cloud data warehouse workloads—query repetitiveness and string-intensive operations—leading to inaccurate system evaluation. To address this, we propose Redbench, a novel synthetic benchmark generation framework grounded in empirical cloud workload trace analysis. Rather than relying on superficial execution metrics, Redbench models intrinsic workload signals—including query pattern distributions, repetition cycles, and string-operation intensity—via integrated techniques for query pattern extraction, repetitiveness modeling, and string-aware enhancement. It enables end-to-end, reproducible, and customizable synthesis of realistic workloads from production traces. Experimental evaluation demonstrates that Redbench-generated workloads significantly improve fidelity and effectively expose performance disparities across four major commercial cloud data warehouses under diverse optimization strategies, thereby providing a high-fidelity, reproducible foundation for rigorous system assessment and optimization.
To address the lack of low-cost, high-fidelity simulation tools for cloud computing research, this paper proposes a lightweight, modular cloud workload simulator. The simulator leverages real trace data from Google’s 12.5K-node cluster and is optimized to run efficiently on commodity desktop machines. Implemented in Scala, it employs parallelized trace parsing and event-driven simulation to accurately model job-, task-, and node-level behaviors, including dynamic resource scheduling and fine-grained resource utilization patterns. Its core contribution lies in a novel simulation framework that jointly achieves high fidelity—preserving key statistical properties of real traces—and scalability—significantly reducing computational overhead. The framework is open-sourced, providing researchers with a reliable, accessible experimental platform for cloud scheduling, performance analysis, and system optimization.
This study addresses the strong dynamism of cloud environments, which demands accurate workload prediction to enable efficient autoscaling. The authors propose a Python-based graphical simulation framework that integrates workload generation, XGBoost and LSTM prediction models, and an MAPE-driven autoscaling mechanism into an end-to-end three-stage simulation pipeline. Notably, this framework uniquely combines an intuitive GUI with multi-level validation—internal, intermediate, and external—to balance realism and accessibility. Experimental results demonstrate that the synthetically generated workloads closely resemble real-world data, with Kolmogorov–Smirnov test p-values of 0.19 for CPU and 0.14 for memory usage. The GUI introduces only modest overhead (1.4×–4.67×) and has received expert endorsement, effectively filling a critical gap in cloud autoscaling research tooling.
Resource management across the IoT/Edge/Cloud continuum faces a fundamental trade-off among prediction accuracy, real-time responsiveness, and deployment lightweightness; existing approaches rely either on runtime sampling or static rules, failing to reconcile these requirements. This paper proposes a metadata-driven runtime workload profiling framework. First, it formally defines the mapping between static metadata and dynamic runtime behavior. Second, it introduces a clustering-based mechanism for extracting semantically significant metadata features—without requiring execution traces or online profiling. Third, it integrates lightweight feature engineering with regression modeling to achieve low-overhead, high-accuracy resource demand prediction. Evaluated across Alibaba’s ML workloads and Google’s cluster traces, the framework maintains high prediction accuracy even under partial data anonymization, substantially outperforming conventional static prediction and online profiling baselines. It achieves superior real-time performance, prediction fidelity, and practical deployability.
This work addresses the limitations of existing high-performance computing workload datasets, which often lack long-term temporal coverage, real memory usage traces, and diverse user origins, thereby hindering accurate evaluation of scheduling algorithms. Leveraging comprehensive job logs from the French IN2P3 Computing Center in 2024, this study presents a publicly available, high-fidelity production-grade dataset encompassing 44 million jobs submitted by 1,000 users across up to 312 machines, with concurrent execution of 46,000 threads and aggregate memory consumption of 105 TB. The dataset substantially enhances realism and timeliness, enabling in-depth analysis of user behavior patterns ranging from weekly to seasonal scales, and provides a critical foundation for more realistic simulation and performance evaluation of scheduling algorithms.
Cloud performance is influenced by multi-scale, time-varying factors, and existing decomposition methods struggle to effectively capture their intermittent behavior and complex periodic patterns. This work proposes both a hybrid (expert-informed) and a fully automated time series decomposition approach that, for the first time, stably extracts multi-scale trend and seasonal components from a single performance trace. These components are directly leveraged for Serverless function performance prediction and AWS resource scheduling. The proposed method significantly outperforms baseline approaches, achieving prediction MAPE as low as 1.8% (hybrid) and 2.1% (fully automated), reducing latency variability on AWS by over 60%, and decreasing peak latency by up to 10%, thereby offering high-precision decision support for cloud resource provisioning.
This study addresses the unclear mechanisms by which AI data centers under shared GPU architectures dynamically influence grid power, particularly the challenge of simultaneously mitigating aggregate power fluctuations and meeting short-term ramping demands. The authors propose a modeling framework calibrated with real-world traces that integrates workload arrival patterns, queuing dynamics, scheduling policies, and GPU power characteristics to systematically analyze the impact of varying batch-to-inference task mixtures. Their findings reveal that power fluctuations follow a U-shaped trend while short-term ramping exhibits a hump-shaped pattern across mixture ratios. At moderate mixture levels, queued batch jobs effectively fill inference idle periods, substantially reducing power variability while preserving essential ramping capability. This work is the first to demonstrate the feasibility of decoupling power fluctuation from ramping requirements through load composition tuning, offering a theoretical foundation for green, coordinated scheduling in AI data centers.