SyclKittens: A Tile Programming Model for Programmers and Coding Agents on Intel GPUs

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance bottlenecks in AI accelerator kernel development caused by insufficient architectural expertise by proposing a hardware-aware tile programming model for Intel GPUs. The proposed approach encapsulates best practices—such as matrix engine layout optimization and L1 cache prefetching—into programmable interfaces, enabling efficient encoding of data movement and computation strategies. This significantly lowers the barrier to low-level optimization and empowers programmers and AI agents to collaboratively construct high-performance kernels. In a SYCL-based prototype implementation, the developed GEMM kernel achieves 96% of oneDNN performance, while Llama inference is accelerated by up to 2.91×.
📝 Abstract
New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and optimization. What agents produce depends on the interfaces they are given. These interfaces may expose a machine's raw capabilities or encode the known-good methods for exploiting its hardware features efficiently. We test the performance impact of these two interfaces on Intel GPUs through controlled experiments with three coding models on three kernels. With raw SYCL and execution feedback, Opus 4.8, the strongest of the three models tested, writes a single-GPU GEMM kernel reaching only 57.5% of Intel's tuned oneDNN library. We present SyclKittens, a hardware-aware tile programming model for Intel GPUs that encodes these known-good methods, so its operations place operands in matrix-engine layouts, prefetch through the L1 cache, and communicate over the fabric that links GPU stacks into a node. With SyclKittens under the same agent, task, and feedback budget, the agent reaches 82.1%, showing that encoded methods turn hardware capabilities into performance. Workload-specific policies stay programmable, and engineers and agents jointly refine these schedules in SyclKittens to build a kernel suite that reaches ~96% of oneDNN in geometric mean across GEMM shapes. The suite runs Llama-3.1-8B inference 1.59x faster than torch.compile on one GPU and up to 2.91x faster than a matched multi-GPU decode path on Intel's oneCCL. SyclKittens is open source and available at https://github.com/intel/SyclKittens.
Problem

Research questions and friction points this paper is trying to address.

AI accelerators
kernel optimization
coding agents
programming model
Intel GPUs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tile Programming Model
AI Coding Agents
Intel GPUs
Kernel Optimization
Hardware-aware Abstraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yehong Jiang
Massachusetts Institute of Technology
S
Sheng Chen
Intel Corporation
F
Fangwen Fu
Intel Corporation
Yen-Kuang Chen
Yen-Kuang Chen
Intel Corporation
X
Xinmin Tian
Intel Corporation
S
Stuart H. Sul
Stanford University
Simran Arora
Simran Arora
Computer Science, Stanford University
Computer ScienceAI Systems