Harness Engineering for LLM-Driven GPU Kernel Generation

📅 2026-07-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a human-in-the-loop framework for reliably generating efficient and correct GPU kernel code. To address the unreliability of large language models (e.g., Codex, Claude Code) in kernel optimization, we introduce a decoupled architecture comprising an evaluation harness and an optimization controller. The harness handles compilation, correctness verification, vendor-aligned benchmarking, and archival, while the controller leverages performance profiling to guide the LLM in generating candidate kernels under human-imposed constraints and high-quality reference implementations. Evaluated on NVIDIA Blackwell B200, our approach achieves significant speedups across multiple operators, outperforming the FlashInfer baseline by average latency reductions of 1.62×, 18.05×, 29.68×, 1.12×, and 13.70× across five operators, respectively. These results demonstrate the efficacy of integrating expert knowledge with LLM capabilities.
📝 Abstract
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be reliably constrained, validated, profiled, and selected. This paper presents a harness-centered system for LLM-driven GPU kernel optimization in the MLSys 2026 FlashInfer AI Kernel Generation Contest on NVIDIA Blackwell B200 GPUs. The system separates an evaluation harness from a profile-backed optimization controller: the harness enforces compilation, correctness, official-aligned timing, and artifact archival, while the controller turns profiler and workload evidence into bounded candidate-generation decisions. Human-authored skills capture operator constraints, references, profiling procedures, and promotion rules, while Codex and Claude Code agents generate candidate kernels inside those constraints. Across five operator definitions, the retained official-aligned artifacts achieved mean-latency speedups over supplied FlashInfer baselines of 1.62x, 18.05x, 29.68x, 1.12x, and 13.70x. The Agent-Assisted kernels outperform the Full-Agent artifacts across the evaluated definitions, indicating that expert-provided optimization directions, high-quality references, and workload context remain critical for reliable AI-driven kernel optimization.
Problem

Research questions and friction points this paper is trying to address.

GPU kernel generation
large language models
code validation
performance profiling
AI-driven optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

harness-centered system
LLM-driven kernel generation
profile-backed optimization
GPU kernel optimization
human-in-the-loop AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yue Shui
Baidu, Inc.
C
Chenyu Ma
Baidu, Inc.
H
Hangfei Xu
Baidu, Inc.
S
Shengzhao Wen
Baidu, Inc.
Y
Yanpeng Wang
Baidu, Inc.