Mixture-of-Kittens: MoE Megakernel for NVL72s

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of existing Mixture-of-Experts (MoE) training systems on large-scale architectures such as NVL72, where they can underperform even naive baselines. To overcome this limitation, this work proposes MoK, an NVL72-oriented MoE system optimized through three core insights: adaptively selecting push or pull communication strategies per operator, restructuring computation-communication overlap mechanisms, and completely eliminating CPU-GPU synchronization overhead. Furthermore, by integrating token scheduling with expert feed-forward networks, MoK constructs a deterministic giant training kernel. Experimental results demonstrate that MoK achieves up to a 2.37× improvement in single-layer throughput and a 1.41× increase in end-to-end training throughput across 512 GPUs.
📝 Abstract
AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single-hop fabrics. We find that existing Mixture-of-Experts (MoE) training systems, optimized for conventional scale-out networks, transfer poorly to this setting, often running slower than a naive baseline built with PyTorch and NCCL. With industry roadmaps pointing toward even larger scale-up domains, understanding the performance tradeoffs of this hardware regime is increasingly important. We present Mixture-of-Kittens (MoK), an MoE training system designed for Nvidia NVL72. MoK builds on three insights that unlock performance on scale-up domains: (1) choosing push- or pull-based communication per operator, (2) restructuring the computation-communication overlap, and (3) fully eliminating CPU-GPU synchronization. MoK distills these insights into a single deterministic training megakernel that fuses token dispatch, shared and routed expert FFNs, and token combine. Across MoE layer shapes from four widely used open-weight models, MoK delivers up to $2.37\times$ the throughput of the strongest publicly available baseline. In a production run on 512 GPUs spanning multiple GB300 NVL72 racks, MoK improves end-to-end training throughput by $1.41\times$.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
scale-up architecture
MoE training
NVL72
GPU communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Megakernel
Scale-up Architecture
NVL72
Communication-Computation Overlap
🔎 Similar Papers
No similar papers found.