LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited acceleration of existing attention kernels caused by their underutilization of low-precision computation. To overcome this, we propose a hardware-algorithm co-designed mixed-precision FlashAttention framework. Specifically, this work introduces a forward-looking pipeline tailored for dedicated accelerators alongside an adaptive precision selection mechanism. Efficiency is enhanced through 8-bit accumulation and the selective recomputation of sensitive sub-blocks at 16-bit precision. Experimental evaluations on Qwen3 and Gemma3 demonstrate that the proposed approach recovers baseline performance by converting only a minimal fraction of sub-blocks to higher precision, thereby enabling highly efficient inference.
📝 Abstract
While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.
Problem

Research questions and friction points this paper is trying to address.

FlashAttention
mixed-precision
dedicated accelerator
numerical stability
hardware-algorithm co-design
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixed-Precision FlashAttention
Hardware-Algorithm Co-design
Dedicated Accelerator
Adaptive Sub-block Recomputation
8-bit Attention
S
Stanislav Budzinskiy
Faculty of Mathematics, University of Vienna, Austria
M
Marian Gloser
Faculty of Mathematics, University of Vienna, Austria
T
Tolunay Yilmaz
Faculty of Mathematics, University of Vienna, Austria
Y
Ying Hong Tham
Huawei Heisenberg Research Center, Munich, Germany
Y
Yuanyi Lin
Huawei Technologies Co. Ltd
W
Wenyi Fang
Huawei Technologies Co. Ltd
F
Fan Wu
Huawei Technologies Co. Ltd
Philipp Petersen
Philipp Petersen
University of Vienna
Applied Harmonic AnalysisDifferential equationsNeural network approximation