Mosaic: GPU Sharing with Latency Guarantees through Kernel-Level Interference Prediction

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of latency-sensitive tasks caused by interference in GPU sharing by proposing Mosaic, a kernel-level interference prediction framework. Mosaic explicitly models multidimensional interference sources—including thread block placement, memory hierarchy contention, and streaming multiprocessor (SM) resource competition—thereby overcoming the limitations of conventional coarse-grained metrics. By integrating analytical models with lightweight learning-based approaches, the framework achieves precise kernel-level interference prediction. Furthermore, it is incorporated into the scheduler to perform online kernel admission control, ensuring strict latency service-level objectives (SLOs). Experimental results demonstrate that Mosaic reduces prediction error by an order of magnitude while maximizing best-effort task throughput under stringent latency constraints.
📝 Abstract
GPUs are increasingly in demand for AI workloads, yet often remain substantially underutilized, motivating workload colocation. However, colocation introduces interference that can degrade latency-critical workloads. Existing approaches mitigate interference using either heuristic-based scheduling or interference predictors. Heuristics rely on coarse-grained metrics that overlook important interference sources, while predictors often depend on simulator-specific or similarly coarse metrics. However, GPU interference is complex and arises from multiple mechanisms, including thread-block placement, memory hierarchy contention, and intra-SM resource contention. We present Mosaic, a kernel-level interference predictor that explicitly models these mechanisms using a combination of analytical and lightweight learned models. Across four GPU architectures, Mosaic reduces prediction error by up to an order of magnitude compared to prior predictors. We integrate Mosaic into a scheduler, MosaicSched, that performs online kernel admission control and selects between full-GPU colocation and SM partitioning to maximize best-effort throughput while satisfying latency SLOs. Across all workloads, MosaicSched keeps the p99 latency below or very close to the target SLO.
Problem

Research questions and friction points this paper is trying to address.

GPU sharing
workload colocation
interference prediction
latency SLO
kernel-level interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU sharing
interference prediction
kernel admission control
SM partitioning
latency SLO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.