KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
This study addresses the scarcity of training data and the inherent trade-off between correctness and performance when large language models generate GPU kernels. To tackle these challenges, we propose a co-evolutionary framework comprising two synergistic models: a Proposer and a Coder. Methodologically, we introduce a frontier-driven module to construct an automatic curriculum learning system that continuously generates high-quality training data. Furthermore, we design a Correctness-Aware Group Relative Policy Optimization (CA-GRPO) algorithm to facilitate alternating reinforcement learning between the two models. Experimental results demonstrate that our approach significantly outperforms advanced models such as Claude-4.5 on the KernelBench benchmark, achieving pass@1 scores of 75.8% for CUDA and 77.2% for Triton. These findings confirm that the proposed framework effectively achieves simultaneous improvements in both kernel generation correctness and execution performance.