Queue-Theoretic Admission Control for Multi-Tenant GPU Clusters

πŸ“… 2026-07-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the lack of predictable latency and formal guarantees for workloads in multi-tenant GPU clusters, where existing systems rely on heuristic policies without theoretical foundations. The paper presents the first formulation of cluster admission control as a multi-class, multi-resource queueing network, integrating an M/G/k queueing model with a vector bin-packing reduction. It introduces an effective service capacity metric, \(k_{\text{eff}}\), to identify bottleneck dimensions and distinguish schedulable from infeasible loads. Theoretically, it establishes the existence of a schedulable subset, proves that waiting time follows an \(O(1/(1-\rho))\) scaling law, and shows that optimal admission ordering under multidimensional resources is NP-hard. Experiments based on Kueue validate the accuracy of Little’s Law and demonstrate that Erlang-C provides a conservative estimate of waiting times.
πŸ“ Abstract
GPU cluster operators cannot predict how long pending workloads will wait for admission. Existing systems use greedy heuristics with no formal wait time guarantees. We formalize GPU cluster admission as a multi-class, multi-resource queueing network and prove a structural decomposition: the pending queue partitions into quotable workloads (bounded wait time under stability) and unfeasible workloads (no finite bound without reconfiguration). For quotable workloads, we model each cluster queue as an M/G/k system where the effective server count k is determined by a vector packing reduction; under an explicit stochastic domination assumption, we establish O(1/(1-rho)) wait time scaling. We prove that optimal admission ordering is NP-hard under multi-dimensional resource demands via reduction from vector bin packing. We validate on Kueue, the standard Kubernetes workload queuing system, using CPU, memory, and GPU (via Dynamic Resource Allocation) resources. The vector k_eff correctly identifies bottleneck resource dimensions, Little's Law holds exactly, and the Erlang-C approximation consistently overestimates observed wait times in the conservative direction.
Problem

Research questions and friction points this paper is trying to address.

admission control
GPU clusters
queueing theory
wait time guarantees
multi-tenant
Innovation

Methods, ideas, or system contributions that make the work stand out.

queueing theory
multi-tenant GPU clusters
vector packing
admission control
M/G/k queue
πŸ”Ž Similar Papers
No similar papers found.