🤖 AI Summary
This study addresses the resource fragmentation and suboptimal allocation issues arising from Multi-Instance GPU (MIG) placement in GPU cloud environments. We propose a hybrid architecture that integrates a lightweight greedy strategy with constrained-neighborhood Binary Integer Linear Programming (BILP). The method rapidly responds to placement requests via greedy rules, invoking BILP optimization over local neighborhoods and triggering live VM migrations only when necessary. This design avoids the prohibitive computational overhead of global optimization while achieving efficient resource consolidation. Experimental results demonstrate that, compared to the optimal baseline, the proposed approach reduces migration volume by 29%–49%, maintains fewer active hosts, and achieves minimal energy consumption, yielding solution quality that closely approximates the theoretical optimum.
📝 Abstract
The extensive use of GPUs in cloud computing, accelerated by the spread of large language model (LLM) services, and the growing need for multitenancy have driven the development of innovative solutions for efficient GPU resource management. Multi-Instance GPU (MIG) technology from NVIDIA enables shared GPU usage in cloud data centers by providing isolated instances, which are offered as MIG-backed virtual GPUs (vGPUs). However, MIG placement rules often lead to fragmentation and suboptimal resource allocation. In this work, we formally model the MIG-aware virtual machine (VM) placement as a binary integer linear programming (BILP) problem aimed at maximizing request acceptance, consolidating resources, and reducing migration overhead. Building upon this formulation, we propose OrigaMIG, a MIG-aware placement optimizer. OrigaMIG places each arriving request at once with a lightweight greedy rule and calls the solver only when a request cannot be placed or the GPUs of a class become fragmented. Each call solves the BILP on a small neighborhood of physical machines (PMs), keeps the rest of the data center fixed, and executes the solution as live migrations. We compare OrigaMIG with the default placement and with two state-of-the-art MIG-aware policies in simulation, across six loads, eleven variants of the workload and the hardware, and data centers of up to 4096 PMs. On 256 A100 and A30 GPUs in 128 PMs, OrigaMIG keeps fewer PMs active than the stronger policy in 29 of 30 paired runs while migrating 29% to 49% fewer VMs, and it uses the least energy per admitted GPU-hour at every load. Against the default placement, it keeps up to 12.5% fewer PMs active and admits up to 7.8 percentage points more of the requested GPU memory. On small data centers, it stays within 4.1% of the optimal number of active PMs.