🤖 AI Summary
This study addresses the scarcity of training data and the inherent trade-off between correctness and performance when large language models generate GPU kernels. To tackle these challenges, we propose a co-evolutionary framework comprising two synergistic models: a Proposer and a Coder. Methodologically, we introduce a frontier-driven module to construct an automatic curriculum learning system that continuously generates high-quality training data. Furthermore, we design a Correctness-Aware Group Relative Policy Optimization (CA-GRPO) algorithm to facilitate alternating reinforcement learning between the two models. Experimental results demonstrate that our approach significantly outperforms advanced models such as Claude-4.5 on the KernelBench benchmark, achieving pass@1 scores of 75.8% for CUDA and 77.2% for Triton. These findings confirm that the proposed framework effectively achieves simultaneous improvements in both kernel generation correctness and execution performance.
📝 Abstract
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.