Score
Designs and implements parallel execution and data-mapping strategies for multi-device tensor computations and tensor-core hardware, including tensor model parallelism and combined tensor-and-model parallel schemes. Builds and tunes tensor kernels and operators using tensorization, tiling, decomposition, fusion and other kernel-level optimizations, programs and optimizes tensor-core codepaths, and analyzes or simulates tensor-network performance to validate and optimize throughput and resource use.
为解决现代AI工作负载中张量计算的调度瓶颈,提出FIBER架构,通过解耦线程与寄存器所有权实现动态并行和细粒度调度。
This work investigates whether Tensor Cores deliver practical acceleration for memory-bound kernels—such as STREAM Scale, SpMV, and stencil computations—challenging recent studies that overestimate their performance in such scenarios. Method: The authors adopt a dual approach: (i) theoretically deriving the upper bound of double-precision speedup under GPU microarchitectural constraints (e.g., memory bandwidth and instruction scheduling overhead), and (ii) empirically validating across V100, A100, and H100 GPUs using both CUDA and WMMA APIs on representative memory-bound kernels. Contribution/Results: The analysis reveals a strict theoretical speedup ceiling of 1.33× for double-precision operations; all empirical measurements fall at or below this bound. The study refutes the efficacy of Tensor Cores in memory-bottlenecked workloads, attributing prior overestimations to neglect of bandwidth saturation and scheduling latency. It establishes the first principled theoretical foundation for Tensor Core applicability boundaries, providing a critical criterion for heterogeneous resource scheduling in GPU-accelerated computing.
Tensor accelerators (e.g., NVIDIA Tensor Cores) are increasingly prevalent in CPUs and GPUs, yet their programmability remains limited: existing kernel libraries target only traditional ML and scientific computing workloads, failing to support non-ML linear matrix transform workloads such as image processing. This paper proposes an equivalence-saturation–based flexible tensor instruction selection mechanism, enabling general-purpose, schedulable compilation for tensor hardware. Integrating with the Halide domain-specific language and compiler, our approach retains full compatibility with existing scheduling primitives while substantially broadening the programmability of tensor accelerators. Evaluated on an NVIDIA RTX 4070, our framework achieves a 6.1× speedup on image processing pipelines—including downsampling—demonstrating, for the first time, systematic performance acceleration of tensor hardware in non-ML domains.
This work addresses the performance limitations of data-intensive applications in edge computing, which often suffer from poor memory locality or excessive memory footprint. To overcome these challenges, the authors propose a hardware-software co-design that integrates a tensor memory engine into the general-purpose CPU datapath. This engine enables runtime dynamic reconfiguration of memory layouts to significantly enhance cache locality, without requiring application code modifications, offloading computation to memory, or incurring memory overhead. By preserving the CPU’s role as the primary compute unit, the approach effectively improves overall system performance. Notably, this study presents the first practical implementation of runtime memory layout reconfiguration on commercial SoC/FPGA platforms, demonstrating both innovation and real-world applicability.
Although stencil computations are traditionally considered memory-bound, they exhibit significant acceleration on Tensor Cores—an apparent contradiction that demands explanation. This work proposes the first applicability analysis framework for stencil computations on Tensor Cores, leveraging systematic performance modeling to quantify the computational redundancy introduced by time tiling and its impact on arithmetic intensity. Based on this analysis, the study derives precise criteria for achieving effective acceleration and identifies the corresponding “sweet spot” region in the design space. Furthermore, it reveals that sparse Tensor Cores can substantially expand this design space. Experimental validation on state-of-the-art implementations such as DRStencil and EBISU confirms the model’s accuracy, offering theoretical guidance for the efficient optimization of stencil computations on NVIDIA GPUs.
Existing scientific computing codes are difficult to efficiently port to specialized architectures such as AMD AI Engine, often requiring extensive manual refactoring. This work proposes a tensor abstraction–based compilation approach that automatically elevates generic loops to tensor semantics by parsing lightweight OpenMP annotations, and constructs an end-to-end compilation pipeline to map computations onto the AI Engine execution model. The method significantly reduces programming complexity through minimal OpenMP directives and enables CPU–NPU cooperative scheduling. Experimental results on six scientific and AI kernel benchmarks show that the NPU achieves higher energy efficiency than a multi-core CPU at float32 precision; for two kernels, cooperative execution yields a 40% performance improvement and 15% energy reduction.
This work investigates whether handcrafted PTX kernels can outperform the WMMA API for mixed-precision GEMM on NVIDIA L4 GPUs (Ada architecture). We design double-buffered GEMM kernels leveraging PTX instructions such as cp.async, ldmatrix, and mma.sync, and conduct a systematic comparison against WMMA implementations across FP16, INT8, and INT4 precisions, supported by in-depth hardware analysis using Nsight Compute. For the first time on the L4 platform, we quantify the performance limits achievable with hand-optimized PTX, revealing that memory behavior—not Tensor Core utilization—dominates large-matrix performance. Our experiments demonstrate speedups of 1.4×–1.8× for INT8 and 2.9×–4.3× for INT4 over WMMA; compared to the FP16 WMMA baseline at N=8192, peak improvements reach 34.4× for INT8 and 98.7× for INT4.
为解决GPU编程中性能优化与安全性问题,提出Exo-GPU语言,通过将并行性和同步性作为顺序代码注解处理,保证功能等效同时便于性能调优。
This work addresses the limited flexibility in distributed programming for large language model scaling and the inefficiency of existing tensor compilers in handling the complex memory hierarchies of heterogeneous clusters. To overcome these challenges, the authors propose a scalable block-level compiler featuring a novel three-tier hierarchical abstraction—Core, Device, and Task—that uniformly supports diverse parallelization strategies, automatically optimizes intra- and inter-node communication, and enables efficient code generation across both NVIDIA and AMD platforms. When integrated into vLLM, the compiler achieves 5%–30% end-to-end inference speedup and over 10% improvement in training model FLOPs utilization (MFU), translating to approximately 500,000 GPU hours saved per month. The system has been deployed in enterprise settings, delivering over 20% inference performance gains.
This work addresses the limited scalability of tensor parallelism in large model online inference, where non-scalable overheads hinder near-linear cluster performance scaling. The authors propose Albireo, a system that eliminates such bottlenecks without modifying model architecture by overlapping scheduling with computation, employing sequence-parallel sampling, and optimizing KV cache management. Albireo further introduces the concept of “empirically optimal tensor parallelism degree” to guide parallelism strategy selection. Experimental results demonstrate that, compared to vLLM, Albireo achieves up to 1.9× higher throughput, 48% lower latency, 28% improved GPU utilization, and 54% reduced energy consumption, with a twofold throughput gain observed in production environments.