๐ค AI Summary
This study addresses the challenges of elusive race conditions and false positives arising from numerical tolerances in the verification of AI-generated GPU kernels. We propose a cross-language methodology that integrates dynamic testing with formal verification. Specifically, binary instrumentation and perturbation injection are employed to expose concurrency defects, while agent-based code rewriting simplifies kernel logic. Subsequently, mathematical equivalence is rigorously proven within the F*/Pulse framework, enabling tolerance-free verification of computational correctness. The proposed approach is successfully applied to multi-framework GEMM implementations and large-scale kernels, where it identifies and automatically repairs four latent defects. Ultimately, this work provides a reliable assurance mechanism for validating the safety and correctness of high-performance code generated by artificial intelligence systems.
๐ Abstract
AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain"reduced-concurrency"versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.