Score
Designs, implements, and integrates PyTorch C++/CUDA extensions and custom operators, including build-system and runtime integration such as profiling hooks, cross-device data transfer abstractions, and distribution-agnostic interfaces. Implements and performance-tunes GPU kernels and batched tensor operations to raise tensor-core utilization (including TF32 activation), minimize SM cycles and dynamic instruction overhead, and improve IPC and overall throughput.
CUDA Graphs in PyTorch suffer from deployment challenges and high overhead—sometimes even yielding negative speedup—due to static graph constraints and redundant host-to-device parameter copies. This paper proposes a compiler-level, fully automatic optimization framework that requires no user code modification. First, it introduces a novel cost-benefit–driven dynamic graph selection mechanism that adaptively enables or bypasses graph capture based on runtime characteristics. Second, it eliminates redundant kernel parameter copying overhead—a previously unaddressed bottleneck. Third, it extends the scope of automatic static graph capture and reuse to support more complex ML workflows. The framework tightly integrates PyTorch 2’s compilation stack, CUDA Graphs’ hardware capabilities, DAG-structured program analysis, and runtime heuristic decision-making. Evaluated across diverse ML benchmarks, it consistently outperforms PyTorch 2, completely eliminates negative speedup, and delivers 1.3×–2.1× average end-to-end speedup.
This work addresses the challenge of automatically generating high-performance GPU tiled kernels from high-level tensor algebra expressions, thereby alleviating the burden of manual optimization. The authors propose an end-to-end compilation framework that integrates layer-wise lowering, expression rewriting, automated schedule search, reduction fusion, and tiling optimizations. For the first time, this framework automatically discovers high-efficiency kernels—comparable to FlashAttention-3—from the mathematical specification of attention operators, while introducing a novel scheduler that preserves program structural regularity. Evaluated on GH200 and RTX 5090 GPUs, the generated kernels achieve up to 23% and 42% higher throughput, respectively, and match or surpass hand-optimized cuDNN kernels across multiple long-sequence configurations.
为解决GPU编程中性能优化与安全性问题,提出Exo-GPU语言,通过将并行性和同步性作为顺序代码注解处理,保证功能等效同时便于性能调优。
This work addresses the performance bottleneck in modern deep learning caused by the overhead of frequent GPU kernel launches due to numerous small-scale tensor operations. To mitigate this, the authors propose a persistent GPU kernel runtime system that continuously schedules computational tasks via a host-side task queue and dynamically loads operators through NVRTC-based just-in-time compilation and function pointer injection, thereby eliminating redundant kernel launches. The design introduces an innovative dual-slot aliasing mechanism for concurrent updates and a generic tensor abstraction, enabling transparent integration with PyTorch via TorchDispatch. Evaluated on representative workloads dominated by small operations, the system achieves up to 15.3× speedup over standard PyTorch, substantially improves GPU utilization, and maintains strong ecosystem compatibility.
To address the challenges of explicit data movement, asynchronous management, and high programming complexity in tensor computations on modern GPUs—particularly NVIDIA’s Hopper architecture—this paper introduces Cypress, a task-driven tensor programming model. Cypress introduces a novel task-level abstraction with sequential semantics, coupled with declarative memory/device mapping and fully automatic compiler scheduling. This enables coordinated offloading to asynchronous hardware units—including the Tensor Memory Accelerator (TMA) and Tensor Cores—while eliminating application-level explicit synchronization, manual data transfers, and concurrency control. Implemented atop a CUDA backend with warp-specialized kernel generation, Cypress achieves 88%–106% of cuBLAS performance on GEMM and 80%–98% of state-of-the-art implementations on Flash Attention. The model significantly improves developer productivity and hardware utilization without sacrificing performance.
为解决现代AI工作负载中张量计算的调度瓶颈,提出FIBER架构,通过解耦线程与寄存器所有权实现动态并行和细粒度调度。
This study addresses the significant end-to-end latency incurred by host-side scheduling overhead when executing Triton kernels within PyTorch. To mitigate this issue, we propose Trident, a compiler backend built upon Torch-MLIR that unifies guard evaluation, argument preparation, and host execution into a single natively compiled code module. This approach eliminates redundant runtime overhead, falling back to the Python layer only when a new specialization is triggered. When integrated with ATen operator optimizations, Trident achieves 1.47× and 1.68× end-to-end speedups over eager mode and torch.compile, respectively, on large language model inference tasks.
This study addresses the absence of systematic comparisons between tile-based programming models such as Triton and cuTile by constructing a controlled benchmark for standardized evaluation on NVIDIA B200 GPUs. Methodologically, we design a unified test suite comprising 45 operators and integrate PyTorch reference implementations, auto-tuning, Roofline modeling, and GPU profiling techniques for comprehensive bottleneck diagnosis. Furthermore, this work presents the first quantitative comparison of token efficiency for LLM-generated code. Our findings reveal that cuTile excels in Tensor Core-intensive kernels, whereas Triton demonstrates superior performance and greater token efficiency in irregular and memory-bandwidth-bound scenarios.
This study addresses the lack of automated Processing-in-Memory (PIM) offloading support in deep learning frameworks and the associated data movement bottlenecks by proposing a compiler-based automated host-PIM placement strategy. Leveraging MLIR and PyTorch compiler technologies, this method overcomes the limitations of fixed operator lists by dynamically profiling the computational intensity and memory access characteristics of loop nests generated through progressive loop order lowering. This enables profile-guided optimization (PGO)-driven automatic offloading decisions. Evaluated across diverse PIM configurations, the proposed approach achieves speedups of 5.1× and 3.6× over CPU-only execution for GPT-J-6B and LLaMA-7B, respectively.
This study addresses the performance limitations of GPU kernels generated by existing compilers and the shortcomings of LLM-based optimization approaches, which often overlook model structure and lack end-to-end verification. We propose a multi-agent collaborative framework that treats compiled models as structured artifacts, focusing specifically on Triton sub-kernel optimization. This work introduces a novel schedule-aware agent search mechanism, integrated with vendor library call protection and a four-stage gated cascaded verification pipeline encompassing static analysis, correctness checking, and performance gating to ensure both safety and efficacy. Evaluated on the KernelBench benchmark, our method achieves average speedups of 1.40×, 1.15×, and 1.07× at the L1, L2, and L3 levels, respectively, compared to torch.compile, significantly enhancing inference efficiency.