cost optimization

Designs, builds, and analyzes architectures, deployment strategies, and resource-management policies to minimize operational and compute costs by creating consumption and compute cost models, cost-reduction strategies, and cost-aware placement, scheduling, and scaling controllers. Develops and applies cost-sensitive learning, cost-aware training, and optimization algorithms that trade off performance and monetary/resource expenditure.

costoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.

computing continuumconfiguration spaceheterogeneous infrastructure

Traditional scaling law estimation suffers from high computational costs due to the absence of efficient budget allocation strategies. This work proposes a novel approach that, for the first time, integrates surrogate-guided pruning into scaling law modeling by combining the Successive Halving algorithm with both parametric and non-parametric surrogate models. This integration enables proactive allocation of computational resources and efficient construction of loss-compute Pareto frontiers. The method substantially improves resource utilization efficiency, achieving relative performance gains of up to 2.84% on real datasets and 5.47% on synthetic datasets, while reducing computational costs by as much as 98.7%.

compute budget allocationefficient estimationlearning curves

Online Rack Placement in Large-Scale Data Centers

Jan 22, 2025
SB
Saumil Baxi
🏛️ Microsoft | Massachusetts Institute of Technology

This work addresses the server rack deployment optimization problem in large-scale data centers, aiming to jointly improve utilization of space, power, and cooling resources while satisfying current demand coverage, enabling future elastic scalability, and supporting dynamic online decision-making. Method: We propose the first Single-Sample Online Approximation (SSOA) method, establishing a multi-stage stochastic optimization framework for real-time re-optimization with provable performance guarantees. The approach integrates integer programming modeling with explicit constraints on physical resources—rack space, power capacity, and thermal dissipation—and supports human-in-the-loop interactive deployment. Contribution/Results: The solution has been deployed at scale across Microsoft’s global data center infrastructure. It achieves annual operational cost savings exceeding $10 million, significantly improves Power Usage Effectiveness (PUE), and reduces greenhouse gas emissions.

Data Center OptimizationEnergy EfficiencyServer Deployment

Power-Aware Scheduling for Multi-Center HPC Electricity Cost Optimization

Mar 14, 2025
AH
Abrar Hossain
🏛️ The University of Toledo | Amazon Web Services | The University of Texas at Arlington

High electricity costs severely hinder the sustainability of multi-site high-performance computing (HPC) systems. Method: This paper proposes TARDIS, the first scheduler integrating power-aware graph neural network (GNN)–based job power consumption prediction with a spatiotemporal cooperative scheduling framework across multiple HPC centers. TARDIS models dynamic job power profiles via GNNs and jointly optimizes task placement across time (leveraging time-of-use electricity pricing) and space (exploiting geographic price differentials) using time-varying electricity price modeling and multi-objective integer programming. Contribution/Results: Unlike conventional single-site or single-dimensional schedulers, TARDIS achieves substantial cost reduction in trace-driven simulations: up to 18% savings in single-center temporal optimization and 10–20% in multi-center scenarios—while maintaining stable throughput and application performance. The approach enables scalable, cost-efficient, and sustainable HPC operations.

Minimizes electricity costs in HPC systemsOptimizes job scheduling across multiple HPC centersPredicts job power consumption using GNN

Using a market economy to provision compute resources across planet-wide clusters

May 23, 2009
MS
M. Stokely
🏛️ Google | Stanford University

To address resource supply-demand imbalances—manifesting as shortages and surpluses—across globally distributed heterogeneous computing clusters, this paper proposes a resource rationing mechanism grounded in real-world market economics. Methodologically, it introduces a periodic simulated-clock auction framework integrating utilization-driven reserve-price setting, long-term resource quota modeling, and supply-demand equilibrium pricing, enabling dynamic price signals to guide users’ autonomous job placement decisions. Its key contribution lies in being the first to systematically embed microeconomic market mechanisms into large-scale distributed resource allocation, replacing static quota or immediate-scheduling paradigms. Evaluated on the Google experimental market, the mechanism significantly incentivizes user migration toward underutilized clusters: resource utilization variance decreases by 32%, and shortage rate drops by 41%. These results empirically validate that price-based incentives can effectively drive system-level behavioral optimization and achieve global resource equilibrium.

Balancing supply-demand via simulated clock auctionsMarket-based provisioning for heterogeneous compute resourcesReducing shortages-surpluses by incentivizing resource-efficient behavior

Latest Papers

What's happening recently
View more

Dynamic workloads in cloud environments often lead to resource over-provisioning, creating a challenging trade-off between cost and response latency. This work proposes a novel approach that integrates LSTM-based predictive auto-scaling with a game-theoretic heuristic for task scheduling, uniquely unifying time-series load forecasting and game-driven real-time decision-making within a single framework. The method achieves substantial reductions in resource costs while maintaining response times comparable to those of conventional heuristic algorithms, matching the performance of purely machine learning–based solutions. By synergistically combining predictive analytics with strategic scheduling, the proposed approach simultaneously optimizes both cost efficiency and service quality, offering a practical and effective solution for dynamic cloud resource management.

Cloud OrchestrationCost OptimizationHybrid Framework

This work addresses the challenge of coupling slow resource provisioning with fast scheduling decisions in network resource allocation, where the former is constrained by switching costs and the latter must satisfy dynamically evolving budget constraints. To this end, we propose the first bilevel online learning framework that integrates Online Convex Optimization (OCO) with Constrained Markov Decision Processes (CMDPs). The upper level performs budget allocation via OCO with switching costs, while the lower level executes state-dependent safe scheduling based on CMDPs. A novel dual feedback mechanism propagates sensitivity information of budget multipliers across layers to enforce cross-level constraint coupling. Additionally, we introduce a budget-adaptive safe exploration strategy to handle dynamic constraints. Theoretical analysis shows that the proposed method achieves near-optimal cumulative regret while satisfying cross-level constraints with high probability, offering dual guarantees on both performance and feasibility.

bi-level online optimizationconstrained Markov decision processcross-level constraints

This work addresses the surge in GPU-intensive workloads in AI data centers, which has led to soaring energy consumption and electricity costs, thereby intensifying stress on power distribution networks. Conventional approaches that decouple workload prediction from scheduling optimization struggle to effectively minimize operational losses. To overcome this limitation, the paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that jointly models prediction and scheduling. By integrating differentiable convex optimization, PTS directly generates optimal scheduling decisions from input features and employs an over-commitment loss function that combines electricity costs with penalties for load shedding. This design enables gradient-based training tailored to downstream objectives, breaking away from the traditional two-stage paradigm. The proposed method significantly reduces operational costs while enhancing system safety.

AI data centersoperational losspower distribution networks

This work addresses the physical constraints—such as energy availability, cooling capacity, and network bandwidth—that challenge the sustainable operation of AI infrastructure, which traditional software-level optimizations alone cannot resolve. The authors propose a joint compute-network optimization framework that explicitly incorporates carbon intensity, water usage, and power capacity as hard constraints within a closed-loop system to co-schedule computing and optical networking resources. A key innovation is the introduction of the “Feasible Sovereign Operating Region” (FSOR), which transforms infeasible solutions into precise decision signals for infrastructure expansion or load curtailment. By integrating task scheduling with optical circuit routing through a scenario-driven approach and embedding multidimensional sustainability constraints, the framework significantly reduces environmental impact, demonstrating its effectiveness in enhancing the sustainability of AI infrastructure under real-world physical limitations.

Compute-Network OptimizationEnvironmental ConstraintsSovereign AI

Hot Scholars

SW

Siwei Wang

National University of Defense Technology
Large-graph studymulti-view fusionmulti-view clustering
DT

Dzmitry Tsetserukou

Associate Professor, Skolkovo Institute of Science and Technology (Skoltech)
RoboticsHapticsUAV SwarmAI
WC

Wei Chen

Microsoft Research Asia
Game TheoryMachine LearningRanking
SG

Shrestha Ghosh

University of Tübingen
Knowledge BasesNatural Language ProcessingInformation RetrievalInformation Extraction
HY

Hongzhi Yin

Professor and ARC Future Fellow, University of Queensland
Recommender SystemGraph LearningSpatial-temporal PredictionEdge Intelligence