TranScope: What the Software Hides About LLM Training Data, the Hardware Reveals at Scale, and Accelerators Magnify

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of hardware-based out-of-distribution detection mechanisms in black-box large language models and their consequent vulnerability to privacy leakage via membership inference attacks. To mitigate these risks, this work is the first to elucidate how training data influences memory locality and alters microarchitectural states through tokenization. By leveraging cycle-level microarchitectural analysis and translation lookaside buffer (TLB) performance monitoring, we propose a hardware-level robust membership detection tool that operates without requiring surrogate models. Experimental evaluations on the PETAL dataset demonstrate that the proposed method achieves an AUC of 0.9, substantially outperforming the baseline of 0.6. These findings establish hardware side channels as a novel paradigm for copyright verification and privacy auditing in large language models.
📝 Abstract
Membership is the root privacy primitive in machine learning: to date, no hardware-based out-of-distribution detection on black-box models has been demonstrated against constant-time, static neural networks with masked confidence. This paper performs the first cycle-level examination of how large language models and vision transformers interact with various modern microarchitecture components, including integrated accelerators, as LLMs scale in size and answers the question of whether the data that a model was trained on affects its execution footprint even without any input-dependent branch, dynamic optimization, or early exit and in constant-time models. The results confirm that the answer is yes and identify which modern hardware components, such as TLBs or on-core accelerators, reveal or amplify that effect. The results also answer whether the signal is informative enough to reliably classify the in-/vs/out-of-distribution property of membership. To understand why, we perform a systematic root cause analysis and find that the transformer's tokenization steps, which happen during training, alter the locality of the accesses the model makes to fetch the vocabulary token later during inference and, as a result, change the page table access patterns and TLB in a previously unknown data-dependent way, causing microarchitectural state to vary significantly based on whether or not the input was in the distribution of the transformer training data. Building on the above observation, we introduce TranScope: the first microarchitecture tool for detecting membership information with low cost, no need for a surrogate model, and significantly higher robustness, e.g., 0.6 AUC for PETAL (best previously reported) vs 0.9 AUC (ours). This reintroduces hardware as both an opportunity, e.g., a tool for checking copyright violation for the first time, and a new channel for inferring membership (MIA).
Problem

Research questions and friction points this paper is trying to address.

Membership Inference
Microarchitectural Side-channel
Large Language Models
Out-of-distribution Detection
Hardware Security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Membership Inference Attack
Microarchitectural Side-Channel
Large Language Models
TLB Locality
TranScope
💼 Related Jobs
No related jobs found.
Joshua Kalyanapu
Joshua Kalyanapu
Ph.D Candidate, North Carolina State University
microarchitectural side-channel attacksformal verificationmicroarchitectureaccelerator
Darsh Asher
Darsh Asher
Ph.D Student, NC State University
Microarchitectural SeculrityMachine Learning Security
K
Kaushal Mhapsekar
North Carolina State University
B
Bita Aslrousta
North Carolina State University
R
Rushiraj Chaitanyakumar Sheth
North Carolina State University
M
Maharshi Mukeshkumar Oza
North Carolina State University
A
Achyuta Kannan
North Carolina State University
Samira Mirbagher Ajorpaz
Samira Mirbagher Ajorpaz
North Carolina State University
Computer ArchitectureSecurityMachine Learning