🤖 AI Summary
This work addresses the insufficient rigor of existing LLM-generated GPU kernel verification methods, which often fail to detect silent errors such as improper NaN handling, cross-shape failures, and accumulated precision deviations. To overcome this limitation, we propose the first tolerance-agnostic, contract-level verification framework, integrating twelve multidimensional adversarial check gates alongside double-precision oracle comparison and mixed-precision (fp16/fp32) analysis to stringently validate kernel correctness. As a key application, we present the first native Blackwell backpropagation CUDA kernels for the Gated Linear Recurrent (GDN) family. Among 2,638 kernels accepted by current systems, our framework identifies 39.5% as entirely incorrect and 62.1% as violating correctness contracts. Using verified kernels, we successfully train five GDN models, whose correctness is confirmed through four independent validation layers, substantially enhancing both verification rigor and training efficiency.
📝 Abstract
Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run, break when the shape changes, or accumulate in fp16 where the reference keeps an fp32 total. We build the instrument that checks correctness properly: a contract-grade verifier of twelve adversarial gates, each a property a correct kernel must satisfy, several of them tolerance-free, so no choice of threshold can explain a failure away. Aimed outward, the verifier audits 2,638 machine-generated kernels that a public system's own harness had already accepted as correct. It finds 39.5% broken beyond any tolerance argument and 62.1% carrying at least one violation. The field's standard test accepts 1,487 kernels the verifier rejects, against only 14 the other way. We defend the finding four independent ways: a 7/7 positive control, a threshold-calibration sweep, 98.5% agreement with the reference benchmark's own correctness code, and a stratified hand-audit. Aimed inward, the verifier judges a kernel of our own: the first native Blackwell tcgen05 training backward for the gated-linear-recurrence (GDN) family, including the reverse-state stage the field still runs on a fallback. We establish its correctness independently, against a double-precision oracle, and train five family members through it. The correctness signal behind reported progress in kernel generation is far weaker than the numbers suggest, and a set of tolerance-free contracts would close most of the gap.