🤖 AI Summary
This study addresses the significant efficiency gap between LLM-generated GPU kernels and expert implementations, where the specific impact of design guidance on performance remains difficult to quantify. To investigate this, we construct a diagnostic benchmark and propose a hierarchical design guidance framework that establishes an evaluation mechanism mapping algorithmic insights to low-level optimizations. This approach systematically examines the capacity of LLM agents to translate multi-granularity expert knowledge into efficient kernels. Comprehensive evaluations are conducted on NVIDIA B200 hardware using paired experiments and multidimensional metrics. The results demonstrate that incorporating structured guidance elevates code correctness to 98.5% and achieves a geometric mean speedup of 2.49×, substantially narrowing the performance disparity between LLM-generated kernels and expert-crafted implementations.
📝 Abstract
GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from $1.69\times$ to $2.49\times$. Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.