๐ค AI Summary
This study addresses the limitation of existing cloud workflow resource provisioning approaches that overlook the influence of DAG topology on task intensity prediction. To this end, it proposes a graph learning-based framework for resource intensity prediction. Methodologically, the work systematically compares multiple modeling strategies, including graph neural networks and handcrafted topological feature extraction, and quantitatively evaluates the contribution of topological features to CPU and memory intensity through large-scale benchmarking. The results demonstrate that integrating handcrafted features with graph-native models significantly enhances prediction accuracy, achieving low mean absolute error. By validating the critical role of topological information in precisely forecasting task resource demands, this research provides effective support for workflow scheduling and resource optimization in cloud environments.
๐ Abstract
Efficient resource provisioning for large-scale workflows on cloud infrastructures is a critical performance engineering challenge. These workflows are often structured as directed acyclic graphs (DAGs), where under-provisioning can cause critical bottlenecks and over-provisioning leads to unnecessary costs. Accurate, task-level prediction of resource intensity (e.g., CPU load and memory usage) is essential for mitigating these issues. While task-level features are commonly used for prediction, the performance impact of the workflow's overall topological structure is often overlooked or assumed. The central question of our work is: To what extent does what part of the DAG topology influence task-level resource intensity, and what is the most effective way to model this influence?
This paper presents a comprehensive benchmark to systematically quantify the impact of graph topology on task intensity prediction. We evaluate and compare a spectrum of modeling approaches. Our findings demonstrate that topology is a critical feature for accurate prediction. Models incorporating important topological information, even through simple handcrafted features, significantly outperform baseline models. We show that graph-native models provide the highest accuracy, achieving low mean absolute errors for both CPU and memory predictions, and can still be combined with simple topological features that they do not learn for better performance.