TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection

📅 2024-11-05
🏛️ arXiv.org
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) suffer from both performance degradation and high latency—due to quadratic computational complexity—when processing ultra-long contexts. To address this, we propose a training-free, dynamic token-level KV cache selection mechanism. Our approach introduces the first token-level quantification of KV importance based on discontinuous attention sparsity, coupled with a per-head soft-voting strategy for fine-grained cache pruning. We further design a dedicated Selection Cache structure and a customized high-efficiency dot-product kernel to jointly optimize accuracy and inference speed. Experiments demonstrate that our method maintains state-of-the-art performance on long-context tasks while accelerating attention computation by up to 23.84× and reducing end-to-end latency by 2.28× compared to existing approaches.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
The rapid advancement of Large Language Models (LLMs) has driven growing demand for processing extended context sequences in contemporary applications. However, this progress faces two major challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by the quadratic computational complexity of attention. These issues hinder the application of LLMs in long-context scenarios. In this paper, we propose Dynamic Token-Level KV Cache Selection (TokenSelect), a training-free method for efficient and accurate long-context inference. TokenSelect builds upon the observation of non-contiguous attention sparsity, using Query-Key dot products to measure per-head KV Cache criticality at token-level. By per-head soft voting mechanism, TokenSelect selectively involves a few critical KV cache tokens in attention calculation without sacrificing accuracy. To further accelerate TokenSelect, we design the Selection Cache based on observations of consecutive Query similarity and implemented efficient dot product kernel, significantly reducing the overhead. A comprehensive evaluation of TokenSelect demonstrates up to 23.84x speedup in attention computation and up to 2.28x acceleration in end-to-end latency, while providing superior performance compared to state-of-the-art long-context inference methods.
Problem

Research questions and friction points this paper is trying to address.

Addresses performance degradation in LLMs with long sequences.
Reduces excessive inference times due to quadratic attention complexity.
Enables efficient long-context inference without sacrificing accuracy.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Token-Level KV Cache Selection
Query-Key dot products for KV Cache criticality
Selection Cache with efficient dot product kernel
University of Science and Technology of China | Tsinghua University | Alibaba Cloud Computing | The Hong Kong University of Science and Technology
W
Wei Wu
School of Artificial Intelligence and Data Science, University of Science and Technology of China, Hefei, China
Zhuoshi Pan
Zhuoshi Pan
Tsinghua University
deep learningnatural language processing
C
Chao Wang
School of Artificial Intelligence and Data Science, University of Science and Technology of China, Hefei, China
Liyi Chen
Liyi Chen
PhD at PolyU, HK
Y
Yunchu Bai
School of Management, University of Science and Technology of China, Hefei, China
K
Kun Fu
Alibaba Cloud Computing, Beijing, China
Z
Zheng Wang
Alibaba Cloud Computing, Beijing, China
Hui Xiong
Hui Xiong
Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser