🤖 AI Summary
This study addresses the limitations of existing training-free sparsification methods, which struggle to simultaneously achieve token-level adaptivity and fixed-budget control while neglecting sensitivity variations across Transformer blocks. To this end, we propose TopK-Guided, a training-free activation sparsification method that uniquely integrates token-level dynamic adaptation with block-wise differentiated budget allocation. Specifically, it employs a bounded TopK mechanism for token-level sparsity adaptivity and a sensitivity-aware algorithm for layer-wise budget distribution, overcoming the inability of conventional approaches to precisely control computational cost without sacrificing accuracy. Experiments on Llama models demonstrate that our method significantly improves perplexity and downstream task accuracy, with advantages particularly pronounced at high sparsity ratios, while maintaining computational overhead comparable to WINA.
📝 Abstract
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.