MpFA: Hardware-Efficient Train-Free QK4V8 FlashAttention Kernels on Blackwell GPUs

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of end-to-end acceleration in all-FP4 attention on Blackwell GPUs caused by non-matmul overheads. To this end, it proposes a training-free FlashAttention kernel that introduces a novel QK4PV8 mixed-precision strategy. By integrating Tensor Core-based rank-one smoothing compensation, asynchronous pipelining, and adaptive parallel partitioning, the proposed method effectively optimizes long-context inference efficiency. Experimental results demonstrate that this kernel achieves a 2.81× output throughput improvement over BF16 on B200 GPUs while recovering 62.5% of the accuracy loss with only approximately 2% computational overhead.
📝 Abstract
Long-context LLM inference pushes modern GPU serving stacks into an attention-bound regime, where both compute and memory are dominated by the softmax-GEMM pipeline. On NVIDIA Blackwell GPUs, FP4 Tensor Cores offer high matmul throughput, but we find that fully FP4 attention often fails to translate this throughput into end-to-end speedups due to non-matmul costs: online quantization after softmax, tensor/shared-memory data movement, and contention on the softmax path. We present MpFA, a training-free FlashAttention kernel optimized for Blackwell. Guided by hardware characterization, MpFA uses mixed precision: NVFP4 for QK and FP8 for PV (QK4PV8). This preserves low-bit QK throughput while avoiding the conversion and scaling overheads of FP4 PV. To recover accuracy without further stressing the softmax pipeline, MpFA introduces rank-one smoothing compensation implemented as an additional Tensor Core MMA. MpFA further improves performance with a fine-grained asynchronous pipeline, tensor-memory reuse, and adaptive parallel partitioning across prefill and decode. On an NVIDIA B200 and across 16K-128K contexts, MpFA improves prefill throughput over state-of-the-art BF16/FP8 baselines and increases end-to-end output throughput by 2.81$\times$ over BF16 FA4 across Llama-3.1-8B and Qwen3-14B. Across five benchmark suites and two models, rank-one compensation recovers 62.5% of the accuracy loss with about 2.0% kernel overhead.
Problem

Research questions and friction points this paper is trying to address.

FlashAttention
long-context LLM inference
Blackwell GPUs
mixed precision
attention-bound
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixed-Precision Attention
Rank-One Smoothing Compensation
Train-Free Quantization
FlashAttention
Blackwell GPUs
🔎 Similar Papers
No similar papers found.