KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of training data and the inherent trade-off between correctness and performance when large language models generate GPU kernels. To tackle these challenges, we propose a co-evolutionary framework comprising two synergistic models: a Proposer and a Coder. Methodologically, we introduce a frontier-driven module to construct an automatic curriculum learning system that continuously generates high-quality training data. Furthermore, we design a Correctness-Aware Group Relative Policy Optimization (CA-GRPO) algorithm to facilitate alternating reinforcement learning between the two models. Experimental results demonstrate that our approach significantly outperforms advanced models such as Claude-4.5 on the KernelBench benchmark, achieving pass@1 scores of 75.8% for CUDA and 77.2% for Triton. These findings confirm that the proposed framework effectively achieves simultaneous improvements in both kernel generation correctness and execution performance.
📝 Abstract
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.
Problem

Research questions and friction points this paper is trying to address.

GPU kernel generation
large language models
training data scarcity
correctness-performance trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-evolution framework
GPU kernel generation
Frontier-driven module generation
CA-GRPO
Automatic curriculum
🔎 Similar Papers
No similar papers found.
C
Changxin Ke
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
R
Rui Zhang
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Z
Zixiang Fang
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Z
Zhenghong Li
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Yuanbo Wen
Yuanbo Wen
Institute of Computing Technology, Chinese Academy of Sciences
Machine Learning System
J
Jiashuo Shen
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
S
Shuo Wang
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Jiaming Guo
Jiaming Guo
Institute of Computing Technology, Chinese Academy of Sciences
Artificial intelligenceReinforcement Learning
L
Ling Li
University of Chinese Academy of Sciences; Intelligent Software Research Center, Institute of Software, CAS
Q
Qi Guo
State Key Lab of Processors, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences
Yunji Chen
Yunji Chen
Institute of Computing Technology, Chinese Academy of Sciences
processor architecturemicroarchitecturemachine learning