SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the privacy leakage risks arising from the input-dependent execution patterns of sparse attention mechanisms in shared GPU environments. It presents the first identification of the resulting SIMA microarchitectural side channel and proposes a phase-aware attack framework. This framework reconstructs sparsity profiles by extracting inference traces and conducting page-level observations, subsequently executing attacks through profiling-based learning algorithms. The proposed approach enables both query attribute inference and response reconstruction. Evaluated in real-world large language model serving scenarios, the method achieves a 90.9% success rate for attribute inference and an 87.3% success rate for response reconstruction, demonstrating significant practical security implications for shared deployment settings.
📝 Abstract
Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior introduces previously unexplored privacy risks. We identify a new GPU micro-architectural side channel, termed Sparsity-Induced Memory Access (SIMA), which arises from secret-dependent key-value cache access patterns induced by sparse attention. Based on this observation, we present SparLeak, a phase-aware side-channel attack that extracts SIMA traces during LLM inference and enables two practical privacy extractions: query attribute inference from prefill-phase traces and autoregressive response reconstruction from decoding-phase traces. By reconstructing approximate token-level sparsity profiles from page-level observations and applying profiling-based learning, SparLeak accurately recovers sensitive information, including user-query attributes and private LLM response content. Extensive evaluation across three LLM architectures, three sparse attention mechanisms, and three privacy-sensitive datasets shows that SparLeak achieves average attack success rates of 90.9% for attribute inference and 87.3% for response reconstruction under real-world LLM serving settings, highlighting the significance to account for SIMA leakage when deploying sparse-attention-based LLM systems. We provide anonymized SIMA traces, trained attack models, evaluation scripts, and documentation as artifacts at https://anonymous.4open.science/r/Janus_artifacts/.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Side-Channel Attack
Privacy Leakage
Large Language Models
Shared GPUs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Side-Channel Attack
GPU Micro-architecture
Privacy Leakage
Large Language Models