🤖 AI Summary
This study addresses the performance degradation of latency-sensitive tasks caused by interference in GPU sharing by proposing Mosaic, a kernel-level interference prediction framework. Mosaic explicitly models multidimensional interference sources—including thread block placement, memory hierarchy contention, and streaming multiprocessor (SM) resource competition—thereby overcoming the limitations of conventional coarse-grained metrics. By integrating analytical models with lightweight learning-based approaches, the framework achieves precise kernel-level interference prediction. Furthermore, it is incorporated into the scheduler to perform online kernel admission control, ensuring strict latency service-level objectives (SLOs). Experimental results demonstrate that Mosaic reduces prediction error by an order of magnitude while maximizing best-effort task throughput under stringent latency constraints.
📝 Abstract
GPUs are increasingly in demand for AI workloads, yet often remain substantially underutilized, motivating workload colocation. However, colocation introduces interference that can degrade latency-critical workloads. Existing approaches mitigate interference using either heuristic-based scheduling or interference predictors. Heuristics rely on coarse-grained metrics that overlook important interference sources, while predictors often depend on simulator-specific or similarly coarse metrics. However, GPU interference is complex and arises from multiple mechanisms, including thread-block placement, memory hierarchy contention, and intra-SM resource contention. We present Mosaic, a kernel-level interference predictor that explicitly models these mechanisms using a combination of analytical and lightweight learned models. Across four GPU architectures, Mosaic reduces prediction error by up to an order of magnitude compared to prior predictors. We integrate Mosaic into a scheduler, MosaicSched, that performs online kernel admission control and selects between full-GPU colocation and SM partitioning to maximize best-effort throughput while satisfying latency SLOs. Across all workloads, MosaicSched keeps the p99 latency below or very close to the target SLO.