DCcluster-Opt: Benchmarking Dynamic Multi-Objective Optimization for Geo-Distributed Data Center Workloads

📅 2025-10-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
A lack of benchmark platforms capable of jointly modeling time-varying environmental factors (e.g., grid carbon intensity, electricity pricing, weather), fine-grained data center physical characteristics (CPU/GPU/memory/HVAC power consumption), and geographically distributed network dynamics (latency, data transfer costs) hinders reproducible research on green AI workload scheduling. Method: We propose the first high-fidelity open-source simulation benchmark—GreenSim—that integrates real-world multi-regional datasets with physics-aware modeling of carbon intensity, meteorology, network latency, and HVAC heat recovery. It features a modular multi-objective reward function unifying carbon emissions, energy consumption, SLA compliance, and water usage optimization, alongside Gymnasium API support and built-in RL and rule-based baseline controllers. Contribution/Results: GreenSim significantly enhances reproducibility, enables fair algorithmic comparison, and accelerates validation of sustainable computing strategies for geo-distributed task scheduling.

Technology Category

Machine Learning: Efficient ML / Green AISearch and Optimization: Sampling/Simulation-based SearchPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSystems and Infrastructure for Web, Mobile and WoT: Sustainability and carbon-aware systems for Web, mobile, and WoTGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
The increasing energy demands and carbon footprint of large-scale AI require intelligent workload management in globally distributed data centers. Yet progress is limited by the absence of benchmarks that realistically capture the interplay of time-varying environmental factors (grid carbon intensity, electricity prices, weather), detailed data center physics (CPUs, GPUs, memory, HVAC energy), and geo-distributed network dynamics (latency and transmission costs). To bridge this gap, we present DCcluster-Opt: an open-source, high-fidelity simulation benchmark for sustainable, geo-temporal task scheduling. DCcluster-Opt combines curated real-world datasets, including AI workload traces, grid carbon intensity, electricity markets, weather across 20 global regions, cloud transmission costs, and empirical network delay parameters with physics-informed models of data center operations, enabling rigorous and reproducible research in sustainable computing. It presents a challenging scheduling problem where a top-level coordinating agent must dynamically reassign or defer tasks that arrive with resource and service-level agreement requirements across a configurable cluster of data centers to optimize multiple objectives. The environment also models advanced components such as heat recovery. A modular reward system enables an explicit study of trade-offs among carbon emissions, energy costs, service level agreements, and water use. It provides a Gymnasium API with baseline controllers, including reinforcement learning and rule-based strategies, to support reproducible ML research and a fair comparison of diverse algorithms. By offering a realistic, configurable, and accessible testbed, DCcluster-Opt accelerates the development and validation of next-generation sustainable computing solutions for geo-distributed data centers.
Problem

Research questions and friction points this paper is trying to address.

Optimizing dynamic workload scheduling across geo-distributed data centers
Balancing carbon emissions, energy costs, and service level agreements
Addressing time-varying environmental factors and data center physics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simulates geo-distributed data centers with real-world datasets
Models multi-objective optimization for dynamic task scheduling
Provides Gymnasium API with baseline reinforcement learning controllers
Antonio Guillen-Perez
Antonio Guillen-Perez
Research Scientist @ Hewlett Packard Labs HPE
Machine LearningReinforcement LearningDeep LearningMulti Agent SystemsRobotics
A
Avisek Naug
Hewlett Packard Enterprise
V
Vineet Gundecha
Hewlett Packard Enterprise
S
Sahand Ghorbanpour
Hewlett Packard Enterprise
R
Ricardo Luna Gutierrez
Hewlett Packard Enterprise
Ashwin Ramesh Babu
Ashwin Ramesh Babu
Senior Research Scientist @ Hewlett Packard Labs
Computer VisionSelf-Supervised LearningAdversarial AttacksReinforcement Learningrenewable
M
Munther Salim
Hewlett Packard Enterprise
S
Shubhanker Banerjee
Hewlett Packard Enterprise
E
Eoin H. Oude Essink
Hewlett Packard Enterprise
D
Damien Fay
Hewlett Packard Enterprise
S
Soumyendu Sarkar
Hewlett Packard Enterprise