A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers

📅 2026-07-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the surge in GPU-intensive workloads in AI data centers, which has led to soaring energy consumption and electricity costs, thereby intensifying stress on power distribution networks. Conventional approaches that decouple workload prediction from scheduling optimization struggle to effectively minimize operational losses. To overcome this limitation, the paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that jointly models prediction and scheduling. By integrating differentiable convex optimization, PTS directly generates optimal scheduling decisions from input features and employs an over-commitment loss function that combines electricity costs with penalties for load shedding. This design enables gradient-based training tailored to downstream objectives, breaking away from the traditional two-stage paradigm. The proposed method significantly reduces operational costs while enhancing system safety.
📝 Abstract
The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.
Problem

Research questions and friction points this paper is trying to address.

AI data centers
workload scheduling
power distribution networks
prediction error
operational loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predict-Then-Schedule
differentiable convex optimization
workload scheduling
AI data centers
end-to-end optimization
🔎 Similar Papers
No similar papers found.