D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant efficiency gap between LLM-generated GPU kernels and expert implementations, where the specific impact of design guidance on performance remains difficult to quantify. To investigate this, we construct a diagnostic benchmark and propose a hierarchical design guidance framework that establishes an evaluation mechanism mapping algorithmic insights to low-level optimizations. This approach systematically examines the capacity of LLM agents to translate multi-granularity expert knowledge into efficient kernels. Comprehensive evaluations are conducted on NVIDIA B200 hardware using paired experiments and multidimensional metrics. The results demonstrate that incorporating structured guidance elevates code correctness to 98.5% and achieves a geometric mean speedup of 2.49×, substantially narrowing the performance disparity between LLM-generated kernels and expert-crafted implementations.
📝 Abstract
GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from $1.69\times$ to $2.49\times$. Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.
Problem

Research questions and friction points this paper is trying to address.

LLM Agents
GPU Kernels
Expert Design Guidance
Benchmark
Code Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diagnostic Benchmark
Expert Design Guidance
GPU Kernel Generation
LLM Agents
Hierarchical Optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Daifeng Li
HKUST
Huiqiang Jiang
Huiqiang Jiang
Microsoft Research Asia
Efficient AILLMsMLSys
Chengruidong Zhang
Chengruidong Zhang
Research SDE, Microsoft
AI for System & System for AI
W
Wei Wu
USTC
X
Xudong Guo
Alibaba Group
J
Jianhong Tu
Alibaba Group
J
Jianwei Zhang
Alibaba Group
B
Binhang Yuan
HKUST
D
Dayiheng Liu
Alibaba Group