Score
Designs, generates, and refines kernel functions or parametric kernel architectures for kernel-based models using human expert priors and constrained search; implements automated synthesis (including LLM-assisted generation) and expert-in-the-loop tuning procedures to bias the kernel search space, reduce required tuning iterations, and improve performance with minimal verifications.
This work addresses the critical bottleneck in modern AI systems—operator performance—whose optimization traditionally relies on expert-crafted implementations that are costly and difficult to scale. To advance this field, the paper presents the first systematic survey of large language model (LLM) and agent-driven approaches for automated operator generation. It introduces a unified framework that integrates existing methodologies, datasets, and evaluation benchmarks, complemented by a structured taxonomy to categorize current techniques. Furthermore, the authors maintain an open-source repository of curated resources, offering a comprehensive reference for researchers. This effort aims to foster standardization, reproducibility, and future innovation in the emerging domain of automated high-performance operator synthesis.
This work addresses the challenge of kernel design in high-dimensional Bayesian optimization, where Gaussian process kernels are typically handcrafted and existing automated discovery methods suffer from limited search spaces or reliance on raw observational data, hindering scalability. To overcome these limitations, the authors propose Kernel Discovery, a framework that leverages large language model (LLM)-driven evolutionary strategies to explore a broader space of kernel functions beyond conventional additive and multiplicative compositions—without requiring any observed data. The approach employs a two-stage LLM-based generation process to produce diverse yet functionally valid kernel candidates and incorporates a leave-one-out continuous ranked probability score (LOO-CRPS) criterion to mitigate overfitting. Evaluated on five high-dimensional Bayesian optimization benchmarks, the method achieves an average rank of 1.2 out of 17, significantly outperforming current baselines.
This work addresses the challenge of automatically generating high-performance GPU kernels that simultaneously achieve correctness and efficiency, a task often hindered by the absence of expert guidance. To this end, the authors propose EGG, a novel framework that decomposes kernel synthesis into two hierarchical stages: algorithmic structure design and hardware-specific tuning. EGG employs a stage-aware multi-agent collaboration architecture, where large language models make progressive decisions guided by expert-derived optimization principles. By integrating techniques such as parallel mapping, tensor tiling, and memory optimization, the framework enables co-optimization of algorithms and hardware characteristics. Experimental results demonstrate that EGG achieves an average performance 2.13× that of PyTorch on both KernelBench and real-world workloads, significantly outperforming existing agent-based and reinforcement learning approaches.
This work addresses a critical limitation of large language models (LLMs) in GPU kernel generation: while they know *what* optimizations to apply, they lack awareness of *when* those optimizations are safe and effective. To bridge this gap, the authors propose the first method to reverse-engineer transferable optimization skills—with explicit validity conditions—from expert-written kernel families. Each skill precisely specifies its applicable scenarios, preconditions for effectiveness, expected performance gains, and pitfalls to avoid. By integrating reverse simplification, multidimensional validation gating (covering compilation, correctness, and performance), and formalized skill representation with LLM-guided optimization, the approach significantly outperforms existing memory-based methods across five workloads on two NVIDIA architectures. Under identical computational budgets, it achieves higher kernel quality and optimization efficiency, with strong generalization demonstrated across 22 independent test cases and no evidence of overfitting.
This work addresses the autoformulation problem—automatically translating natural-language problem descriptions into solvable mathematical optimization models. We propose the first LLM-driven Monte Carlo Tree Search (MCTS) framework for this task, enabling dynamic hypothesis generation and formal correctness evaluation. Our method integrates hierarchical optimization modeling representations, LLM-based semantic understanding, and MCTS-based search strategies. A key innovation is an equivalence-aware pruning mechanism that reduces search overhead by over 40%. Empirically, our approach achieves state-of-the-art performance on LP/MIP benchmarks, outperforming all existing baselines. LLM-assisted verification accelerates correctness assessment significantly. Moreover, this work formally defines the autoformulation task for the first time, establishing a scalable, automated paradigm to lower the barrier to optimization modeling and empower domain experts.
This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.
This work proposes a human-in-the-loop framework for reliably generating efficient and correct GPU kernel code. To address the unreliability of large language models (e.g., Codex, Claude Code) in kernel optimization, we introduce a decoupled architecture comprising an evaluation harness and an optimization controller. The harness handles compilation, correctness verification, vendor-aligned benchmarking, and archival, while the controller leverages performance profiling to guide the LLM in generating candidate kernels under human-imposed constraints and high-quality reference implementations. Evaluated on NVIDIA Blackwell B200, our approach achieves significant speedups across multiple operators, outperforming the FlashInfer baseline by average latency reductions of 1.62×, 18.05×, 29.68×, 1.12×, and 13.70× across five operators, respectively. These results demonstrate the efficacy of integrating expert knowledge with LLM capabilities.
This work addresses the challenge of scaling purely data-driven approaches in real-world AI tasks with long feedback information loops (FILs), such as GPU programming, where sparse validation signals hinder learning. Challenging the prevailing paradigm that general-purpose methods will ultimately prevail, this study introduces FIL duration as a critical scalability dimension and proposes integrating human prior knowledge to inject inductive bias. Specifically, it constrains the solution space, designs kernel functions grounded in domain-specific inductive biases, and models GPU kernel performance to effectively compensate for the limitations of data-driven methods. Experimental results demonstrate that the proposed approach significantly outperforms purely data-driven baselines on realistic GPU programming tasks, underscoring the efficacy and necessity of combining expert knowledge with inductive bias in long-FIL scenarios.