MetaKernelBench: Measuring GPU Kernel Knowledge Transfer Beyond Code

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing GPU kernel optimization benchmarks evaluate only code correctness, failing to quantify the value of cross-language transfer of empirical knowledge. This work constructs the first cross-DSL knowledge transfer benchmark comprising 74 CuTe/TIRx variant pairs and proposes a large language model-based agent system that integrates natural language skill distillation, conditionalized attempts, and end-to-end runtime performance evaluation to isolate and verify skill reuse gains. Experimental results demonstrate that six models exhibit performance variations ranging from -19% to +29% during bidirectional transfer, confirming that relative performance in the source domain significantly determines transfer success.
📝 Abstract
Recent GPU kernel optimization agents retain what they learn in knowledge bases or as distilled skills. Kernel benchmarks score each attempt's implementation for correctness and speed but leave the reuse value of retained experience unmeasured. We introduce MetaKernelBench, which measures whether experience distilled from an attempt in one kernel domain-specific language (DSL) improves a fresh attempt at the same problem in another. Its 74 problems are fused subgraphs in six families, each posed as a pair of CuTe DSL and TIRx variants that differ only in the DSL. The agent first attempts each variant solo and is instructed to distill what it learns into a natural-language skill, which is transferred whether or not the source attempt passes verification. The skill is the only extra input to a skill-conditioned attempt by the same model in the other DSL. We compare each skill-conditioned attempt with the solo attempt on the same variant under matched per-attempt budgets, scoring correctness and end-to-end runtime. Across six models and both directions, paired lift over solo attempts ranges from -19% to +29%. Four models gain in both directions, yet regressions occur on 16% to 45% of problems in every model and direction. Outcomes follow the source attempt's result relative to the target's solo attempt rather than source success alone, improving in 71% of comparisons when the source stands above and regressing in 54% when it stands below. MetaKernelBench complements implementation-quality metrics by measuring same-problem cross-DSL kernel knowledge transfer.
Problem

Research questions and friction points this paper is trying to address.

GPU kernel optimization
knowledge transfer
cross-DSL
benchmark
experience reuse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-DSL Knowledge Transfer
GPU Kernel Optimization
Skill Distillation
Benchmark
LLM Agents
🔎 Similar Papers
No similar papers found.