Score
Designs, builds, and analyzes architectures, deployment strategies, and resource-management policies to minimize operational and compute costs by creating consumption and compute cost models, cost-reduction strategies, and cost-aware placement, scheduling, and scaling controllers. Develops and applies cost-sensitive learning, cost-aware training, and optimization algorithms that trade off performance and monetary/resource expenditure.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
Traditional scaling law estimation suffers from high computational costs due to the absence of efficient budget allocation strategies. This work proposes a novel approach that, for the first time, integrates surrogate-guided pruning into scaling law modeling by combining the Successive Halving algorithm with both parametric and non-parametric surrogate models. This integration enables proactive allocation of computational resources and efficient construction of loss-compute Pareto frontiers. The method substantially improves resource utilization efficiency, achieving relative performance gains of up to 2.84% on real datasets and 5.47% on synthetic datasets, while reducing computational costs by as much as 98.7%.
This work addresses the server rack deployment optimization problem in large-scale data centers, aiming to jointly improve utilization of space, power, and cooling resources while satisfying current demand coverage, enabling future elastic scalability, and supporting dynamic online decision-making. Method: We propose the first Single-Sample Online Approximation (SSOA) method, establishing a multi-stage stochastic optimization framework for real-time re-optimization with provable performance guarantees. The approach integrates integer programming modeling with explicit constraints on physical resources—rack space, power capacity, and thermal dissipation—and supports human-in-the-loop interactive deployment. Contribution/Results: The solution has been deployed at scale across Microsoft’s global data center infrastructure. It achieves annual operational cost savings exceeding $10 million, significantly improves Power Usage Effectiveness (PUE), and reduces greenhouse gas emissions.
High electricity costs severely hinder the sustainability of multi-site high-performance computing (HPC) systems. Method: This paper proposes TARDIS, the first scheduler integrating power-aware graph neural network (GNN)–based job power consumption prediction with a spatiotemporal cooperative scheduling framework across multiple HPC centers. TARDIS models dynamic job power profiles via GNNs and jointly optimizes task placement across time (leveraging time-of-use electricity pricing) and space (exploiting geographic price differentials) using time-varying electricity price modeling and multi-objective integer programming. Contribution/Results: Unlike conventional single-site or single-dimensional schedulers, TARDIS achieves substantial cost reduction in trace-driven simulations: up to 18% savings in single-center temporal optimization and 10–20% in multi-center scenarios—while maintaining stable throughput and application performance. The approach enables scalable, cost-efficient, and sustainable HPC operations.
To address resource supply-demand imbalances—manifesting as shortages and surpluses—across globally distributed heterogeneous computing clusters, this paper proposes a resource rationing mechanism grounded in real-world market economics. Methodologically, it introduces a periodic simulated-clock auction framework integrating utilization-driven reserve-price setting, long-term resource quota modeling, and supply-demand equilibrium pricing, enabling dynamic price signals to guide users’ autonomous job placement decisions. Its key contribution lies in being the first to systematically embed microeconomic market mechanisms into large-scale distributed resource allocation, replacing static quota or immediate-scheduling paradigms. Evaluated on the Google experimental market, the mechanism significantly incentivizes user migration toward underutilized clusters: resource utilization variance decreases by 32%, and shortage rate drops by 41%. These results empirically validate that price-based incentives can effectively drive system-level behavioral optimization and achieve global resource equilibrium.
为解决大规模资源投资问题中的调度难题,提出iScheduler框架,利用强化学习驱动迭代优化,加速求解并支持快速重配置。
Dynamic workloads in cloud environments often lead to resource over-provisioning, creating a challenging trade-off between cost and response latency. This work proposes a novel approach that integrates LSTM-based predictive auto-scaling with a game-theoretic heuristic for task scheduling, uniquely unifying time-series load forecasting and game-driven real-time decision-making within a single framework. The method achieves substantial reductions in resource costs while maintaining response times comparable to those of conventional heuristic algorithms, matching the performance of purely machine learning–based solutions. By synergistically combining predictive analytics with strategic scheduling, the proposed approach simultaneously optimizes both cost efficiency and service quality, offering a practical and effective solution for dynamic cloud resource management.
This work addresses the challenge of coupling slow resource provisioning with fast scheduling decisions in network resource allocation, where the former is constrained by switching costs and the latter must satisfy dynamically evolving budget constraints. To this end, we propose the first bilevel online learning framework that integrates Online Convex Optimization (OCO) with Constrained Markov Decision Processes (CMDPs). The upper level performs budget allocation via OCO with switching costs, while the lower level executes state-dependent safe scheduling based on CMDPs. A novel dual feedback mechanism propagates sensitivity information of budget multipliers across layers to enforce cross-level constraint coupling. Additionally, we introduce a budget-adaptive safe exploration strategy to handle dynamic constraints. Theoretical analysis shows that the proposed method achieves near-optimal cumulative regret while satisfying cross-level constraints with high probability, offering dual guarantees on both performance and feasibility.
This work addresses the surge in GPU-intensive workloads in AI data centers, which has led to soaring energy consumption and electricity costs, thereby intensifying stress on power distribution networks. Conventional approaches that decouple workload prediction from scheduling optimization struggle to effectively minimize operational losses. To overcome this limitation, the paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that jointly models prediction and scheduling. By integrating differentiable convex optimization, PTS directly generates optimal scheduling decisions from input features and employs an over-commitment loss function that combines electricity costs with penalties for load shedding. This design enables gradient-based training tailored to downstream objectives, breaking away from the traditional two-stage paradigm. The proposed method significantly reduces operational costs while enhancing system safety.
This work addresses the physical constraints—such as energy availability, cooling capacity, and network bandwidth—that challenge the sustainable operation of AI infrastructure, which traditional software-level optimizations alone cannot resolve. The authors propose a joint compute-network optimization framework that explicitly incorporates carbon intensity, water usage, and power capacity as hard constraints within a closed-loop system to co-schedule computing and optical networking resources. A key innovation is the introduction of the “Feasible Sovereign Operating Region” (FSOR), which transforms infeasible solutions into precise decision signals for infrastructure expansion or load curtailment. By integrating task scheduling with optical circuit routing through a scenario-driven approach and embedding multidimensional sustainability constraints, the framework significantly reduces environmental impact, demonstrating its effectiveness in enhancing the sustainability of AI infrastructure under real-world physical limitations.