Dissecting How Die Scaling Breaks GPU Fine-grained Scheduling

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对GPU因晶圆缩放导致的计算单元和内存访问不对称问题,提出了一种轻量级表征方法来揭示每颗芯片的具体拓扑结构及内存亲和性,并据此改进了细粒度调度策略,以提升性能。
📝 Abstract
Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning. The former creates chip-specific compute topologies, while the latter causes non-uniform memory access. These asymmetries are substantial. Topology-oblivious compute unit allocation can lead to up to 1.33x performance variation, while remote accesses increase HBM latency by up to 67% and nearly double L2 latency. However, these asymmetries are hidden behind the GPU's logical resource abstractions and can vary across chips. We develop lightweight characterization methods to uncover per-chip compute topology and memory affinity. We then use the discovered information to make existing fine-grained scheduling asymmetry-aware, considering not only how many resources are allocated but also which physical resources are assigned. Across full-GPU kernel execution, intra-application multiplexing, and inter-application co-location, asymmetry-aware scheduling improves mainstream kernels by up to 1.22x, multiplexed LLM inference by up to 14.3%, and avoids up to 1.33x performance variation.
Problem

Research questions and friction points this paper is trying to address.

Die Scaling
GPU Fine-grained Scheduling
Non-uniform Memory Access
Performance Variation
Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Die Scaling
Asymmetry-aware Scheduling
Compute Topology
Memory Affinity
Lightweight Characterization
🔎 Similar Papers
No similar papers found.