SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization

📅 2025-08-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current search-based LLM approaches lack hardware awareness, hindering near-optimal GPU kernel performance optimization. This paper introduces the first hardware-aware LLM framework for automated spatial optimization targeting disaggregated architectures. Our method integrates memory access pattern analysis, fine-grained architectural modeling, and curated historical performance logs with real-time feedback to guide LLMs in collaboratively generating optimal swizzling strategies. Evaluated on ten representative ML and scientific computing kernels, our approach achieves up to 2.06× speedup on nine kernels—accompanied by a 70% improvement in L2 cache hit rate—and reduces GEMM optimization time from two weeks to under five minutes. The framework significantly enhances both hardware-software co-optimization efficiency and generalizability across diverse workloads and architectures.

Technology Category

Machine Learning: Hardware-aware MLSearch and Optimization: Learning to SearchPlanning, Routing, and Scheduling: Optimization of Spatio-temporal Systems

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd workGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models (LLMs) have shown progress in GPU kernel performance engineering using inefficient search-based methods that optimize around runtime. Any existing approach lacks a key characteristic that human performance engineers rely on for near-optimal utilization -- hardware-awareness. By leveraging the workload's specific memory access patterns, architecture specifications, filtered profiling logs, and reflections on historical performance, we can make software-level optimizations that are tailored to the underlying hardware. SwizzlePerf automatically generates spatial optimizations for GPU kernels on disaggregated architectures by giving LLMs explicit hardware-awareness. For a GEMM kernel, SwizzlePerf takes less than 5 minutes to generate the same hardware-specific optimal swizzling pattern that took expert performance engineers 2 weeks to find. On a suite of 10 diverse ML and Science kernels, SwizzlePerf can generate swizzling patterns for 9 of the kernels that achieve up to a 2.06x speedup and 70% improvement in L2 hit rate. This work is the first of many steps toward systematically creating hardware-aware LLM performance engineering agents.
Problem

Research questions and friction points this paper is trying to address.

Optimizing GPU kernel performance using hardware-aware LLMs
Addressing inefficiency in search-based runtime optimization methods
Automating spatial optimizations for disaggregated GPU architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hardware-aware LLMs for GPU optimization
Automatic spatial optimization generation
Leveraging memory patterns and architecture specs
🔎 Similar Papers
No similar papers found.