🤖 AI Summary
This work addresses the performance bottleneck in GPU-based Byte Pair Encoding (BPE) tokenization caused by frequent data movement. We propose an efficient parallel algorithm based on array-linked lists, which represent token sequences such that merge operations are reduced to constant-time pointer updates. Furthermore, kernel fusion techniques are introduced to eliminate redundant memory accesses, thereby overcoming the memory-bound limitations of conventional GPU implementations. Experimental results demonstrate that the proposed approach achieves a 5.2× throughput improvement over the state-of-the-art GPU implementation and a 24.6× speedup compared to optimized CPU baselines.
📝 Abstract
Every LLM inference begins with tokenization, which converts raw input bytes into the discrete token sequence the model consumes. For text, this step is often implemented using Byte Pair Encoding (BPE), an algorithm originally introduced for data compression. BPE has traditionally run on the CPU with extensive optimization, but recent work has moved it to the GPU for higher throughput. We show that these GPU implementations are bottlenecked not by computation but by data movement. We develop LLBPE that represents the token sequence as an array-based linked list so that each merge reduces to a constant- time pointer update. Furthermore, LLBPE fuses rank lookup, minimum selection, and merging into a single kernel to eliminate redundant hash map queries. LLBPE achieves up to 5.2x higher throughput than the best existing GPU implementation and 24.6x over optimized CPU implementations, at the cost of minor discrepancies in tokenized output.