SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment

📅 2025-10-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Deploying large language models (LLMs) on edge devices faces a dual challenge—vulnerability to model extraction attacks and difficulty in leveraging untrusted GPU accelerators without compromising model confidentiality. Method: This paper proposes a heterogeneous secure inference framework that introduces the first threat-aware, operator-level tensor partitioning strategy: sensitive nonlinear layers execute inside Intel SGX enclaves, while linear layers are encrypted and outsourced to GPUs. The framework integrates encrypted tensor outsourcing, secure reconstruction mechanisms, and information-theoretic security guarantees. Contribution/Results: Evaluated on LLaMA-2, the prototype achieves strong model confidentiality with only ~1.8× inference latency overhead—significantly outperforming enclave-only baselines. It enables efficient, privacy-preserving LLM deployment on resource-constrained edge devices, offering a practical solution for privacy-sensitive applications.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

Security and Privacy: Large-scale security measurementsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
With the increasing deployment of Large Language Models (LLMs) on mobile and edge platforms, securing them against model extraction attacks has become a pressing concern. However, protecting model privacy without sacrificing the performance benefits of untrusted AI accelerators, such as GPUs, presents a challenging trade-off. In this paper, we initiate the study of high-performance execution on LLMs and present SecureInfer, a hybrid framework that leverages a heterogeneous Trusted Execution Environments (TEEs)-GPU architecture to isolate privacy-critical components while offloading compute-intensive operations to untrusted accelerators. Building upon an outsourcing scheme, SecureInfer adopts an information-theoretic and threat-informed partitioning strategy: security-sensitive components, including non-linear layers, projection of attention head, FNN transformations, and LoRA adapters, are executed inside an SGX enclave, while other linear operations (matrix multiplication) are performed on the GPU after encryption and are securely restored within the enclave. We implement a prototype of SecureInfer using the LLaMA-2 model and evaluate it across performance and security metrics. Our results show that SecureInfer offers strong security guarantees with reasonable performance, offering a practical solution for secure on-device model inference.
Problem

Research questions and friction points this paper is trying to address.

Securing LLMs against model extraction attacks
Protecting model privacy while using untrusted accelerators
Isolating privacy-critical components in TEE-GPU architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid TEE-GPU architecture isolates privacy-critical components
Information-theoretic partitioning secures sensitive model operations
Encrypted linear operations offloaded to untrusted GPU accelerators
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tushar Nayan
Florida International University, Miami, FL, USA
Z
Ziqi Zhang
University of Illinois Urbana-Champaign, Urbana, IL, USA
Ruimin Sun
Ruimin Sun
Florida International University, Miami, FL, USA