Closed-Loop Decision-Focused Learning for User-Aware Cloud Orchestration under Uncertainty

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiencies and contention during peak periods caused by time-varying cloud workloads by proposing CL-DFL, an end-to-end closed-loop decision-focused learning framework. The approach formulates heterogeneous task scheduling as a multi-objective combinatorial optimization problem under uncertainty, enabling joint optimization of resource awareness and scheduling decisions. CL-DFL integrates spatiotemporal forecasting based on MTGNN, zeroth-order decision-focused learning via Tree-structured Parzen Estimator (TPE), and GRPO-driven cooperative local search, further enhanced by the GNeuro-PLS strategy to improve robustness under heterogeneous loads. Experimental results on four real-world datasets demonstrate that the proposed method significantly outperforms existing approaches, achieving a superior trade-off among SLA violation rate, user satisfaction, and resource utilization, while maintaining robust performance even under highly saturated conditions.
📝 Abstract
Time-varying cloud workloads often cause resource under-utilization during off-peak periods and resource contention during peak periods. Existing prediction-then-optimization (PTO) frameworks suffer from two-stage decoupling, hindering the balance among violation rate, user satisfaction, and resource utilization. We formulate heterogeneous job scheduling as a multi-objective combinatorial optimization problem (MOCOP) under uncertain constraints and propose a closed-loop decision-focused learning (CL-DFL) framework for cloud orchestration. CL-DFL integrates a Multivariate Time-series Graph Neural Network (MTGNN)-based spatio-temporal predictor with a zeroth-order decision-focused learning (DFL) mechanism based on the tree-structured Parzen estimator (TPE). This integration establishes an end-to-end (E2E) feedback pathway between resource perception and scheduling decisions. Furthermore, we develop the GNeuro-PLS strategy by incorporating group relative policy optimization (GRPO) into cooperative local search to improve robustness under heterogeneous workloads. Extensive experiments on four real-world datasets demonstrate that CL-DFL achieves superior trade-offs among violation rate, user satisfaction, and resource utilization. It effectively controls overload risks under regular workloads and maintains resilience under highly saturated scenarios compared with state-of-the-art baselines.
Problem

Research questions and friction points this paper is trying to address.

cloud orchestration
uncertainty
user satisfaction
resource utilization
violation rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop Decision-Focused Learning
Multi-objective Combinatorial Optimization
MTGNN
Zeroth-order Optimization
User-aware Orchestration
🔎 Similar Papers
No similar papers found.
Dongbin Jiao
Dongbin Jiao
Lanzhou University
Network Resource OptimizationUAV NetworksLow-altitude EconomyEvolutionary Learning
X
Xubo Zhang
School of Information Science and Engineering, Lanzhou University, Lanzhou, 730000, P. R. China
H
Huakang Lin
Division of Computer Science and Engineering, College of Engineering, Louisiana State University, Baton Rouge, LA 70803, USA
K
Ke Shang
School of Artificial Intelligence, Shenzhen University, Shenzhen, 518060, P. R. China
Shi Yan
Shi Yan
Eindhoven University of Technology
Optical communicationfiber opticsSignal processing