🤖 AI Summary
This work proposes a human-in-the-loop framework for reliably generating efficient and correct GPU kernel code. To address the unreliability of large language models (e.g., Codex, Claude Code) in kernel optimization, we introduce a decoupled architecture comprising an evaluation harness and an optimization controller. The harness handles compilation, correctness verification, vendor-aligned benchmarking, and archival, while the controller leverages performance profiling to guide the LLM in generating candidate kernels under human-imposed constraints and high-quality reference implementations. Evaluated on NVIDIA Blackwell B200, our approach achieves significant speedups across multiple operators, outperforming the FlashInfer baseline by average latency reductions of 1.62×, 18.05×, 29.68×, 1.12×, and 13.70× across five operators, respectively. These results demonstrate the efficacy of integrating expert knowledge with LLM capabilities.
📝 Abstract
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be reliably constrained, validated, profiled, and selected. This paper presents a harness-centered system for LLM-driven GPU kernel optimization in the MLSys 2026 FlashInfer AI Kernel Generation Contest on NVIDIA Blackwell B200 GPUs. The system separates an evaluation harness from a profile-backed optimization controller: the harness enforces compilation, correctness, official-aligned timing, and artifact archival, while the controller turns profiler and workload evidence into bounded candidate-generation decisions. Human-authored skills capture operator constraints, references, profiling procedures, and promotion rules, while Codex and Claude Code agents generate candidate kernels inside those constraints. Across five operator definitions, the retained official-aligned artifacts achieved mean-latency speedups over supplied FlashInfer baselines of 1.62x, 18.05x, 29.68x, 1.12x, and 13.70x. The Agent-Assisted kernels outperform the Full-Agent artifacts across the evaluated definitions, indicating that expert-provided optimization directions, high-quality references, and workload context remain critical for reliable AI-driven kernel optimization.