KEET: Explaining Performance of GPU Kernels Using LLM Agents

📅 2026-05-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
📝 Abstract
Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU architecture, kernel developers need to spend significant time analyzing and comparing profiles in the tool's graphical interface to identify and understand kernel performance bottlenecks. Large Language Models (LLMs) have shown promise in understanding complex data and generating natural language explanations. In this paper, we propose the Kernel Execution Explanation Toolkit (KEET), an LLM-based agentic framework for interpreting Nsight Compute profiles to generate useful and data-grounded natural language explanations of performance issues in GPU kernels, and suggestions for optimizations. We evaluate \toolname using several CUDA kernels of varying complexity on NVIDIA H100 GPUs. We find that the generated explanations, when provided as context, improve the quality of LLM code optimization and multiple-choice question answering in downstream tasks. We further demonstrate that the tool can be used to interpret performance data from large sets of profiles to improve the quality of optimization suggestions.
Problem

Research questions and friction points this paper is trying to address.

GPU kernels
performance profiling
Nsight Compute
performance bottlenecks
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM agents
GPU kernel optimization
performance profiling
Nsight Compute
natural language explanation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Joshua H. Davis
Joshua H. Davis
PhD Student, University of Maryland
high-performance computingprogramming models
K
Klaudiusz Rydzy
Department of Computer Science, University of Maryland, College Park, MD, USA
S
Srinivasan Ramesh
NVIDIA, Inc., Santa Clara, CA, USA
A
Aadit Nilay
Department of Computer Science, University of Maryland, College Park, MD, USA
Daniel Nichols
Daniel Nichols
Doctoral Student, University of Maryland, College Park
computer sciencehigh performance computingdeep learning
S
Swapna Raj
NVIDIA, Inc., Santa Clara, CA, USA
Nikhil Jain
Nikhil Jain
Nvidia
Parallel Computing
Abhinav Bhatele
Abhinav Bhatele
Associate Professor of Computer Science, University of Maryland, College Park
Parallel Systems and SoftwareDistributed AIHPCMLforSys