Empowering Hybrid Attention Models on NPUs

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inference bottlenecks of hybrid-attention large language models (LLMs) on edge neural processing units (NPUs), which stem from memory inefficiency and architectural mismatch. We propose HA-NPU, the first system enabling efficient deployment of hybrid LLMs without algorithmic modifications. HA-NPU innovatively reconstructs dataflow across core, operator, and tensor levels by leveraging key techniques including head-dimension partitioning, dependent operator fusion, immediate consumption of intermediate tensors, and dataflow-aware layout planning to deeply optimize the hardware execution efficiency of linear attention. Experimental results demonstrate that HA-NPU accelerates linear attention kernels by 35.95×, reduces energy consumption by 36.14×, and improves end-to-end inference latency by 2.03×, thereby achieving efficient on-device deployment.
📝 Abstract
Hybrid attention models have emerged as a crucial architecture for Large Language Models (LLMs) (e.g., the Qwen3.5 and Kimi series). Their memory and computational efficiency make them highly attractive for on-device inference, forming a promising synergy with edge Neural Processing Units (NPUs). However, naive execution of these hybrid models on edge NPUs fails to deliver these benefits, often bottlenecking the prefill stage due to severe memory-system inefficiencies and architectural mismatches within the linear attention (LA) layers. We present HA-NPU, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms. HA-NPU enhances execution efficiency by reorganizing the dataflow of the LA components across three levels: (1) At the core level, it partitions workloads by the head dimension and fuses dependent operators, eliminating cross-core global memory accesses; (2) At the operator level, it reorders execution to consume intermediate tensors immediately, drastically minimizing local-buffer pressure; (3) At the tensor level, it employs dataflow-aware layout planning to minimize transformation overhead between matrix and vector processing units. Compared to competitive baselines, HA-NPU achieves up to 35.95$\times$ LA kernel speedup and 36.14$\times$ energy reduction, delivering up to 2.03$\times$ faster end-to-end request latency. The source code will be made publicly available at https://github.com/yinyuanzhang/HA-NPU
Problem

Research questions and friction points this paper is trying to address.

Hybrid Attention Models
Edge NPUs
Linear Attention
Prefill Bottleneck
On-device Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Attention
Edge NPU
Dataflow Optimization
Linear Attention
On-device Inference