secure llm serving

Designs, implements, and evaluates runtime architectures, protocols, and tooling to perform large language model inference inside secure execution contexts (e.g., TEEs) or on-device secure environments so that model weights and user data are protected from OS compromise. This work covers secure enclave integration with host OS services (memory management, attestation), device-side secure inference flows, and resource/Isolation strategies to minimize latency while preserving confidentiality and integrity.

securellmserving

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical security vulnerability in deploying large language models (LLMs) on third-party devices, where existing approaches often rely on partially trusted execution environments (TEEs) and precomputed noise to protect intellectual property. We demonstrate for the first time that the common practice of statically reusing secret bases—introduced to accelerate encrypted inference in mainstream LLMs—leaks key information. Leveraging cryptanalysis, linear algebraic reconstruction, and adversarial queries, we develop two novel attacks that reverse-engineer TEE protocols and encrypted inference pipelines. Our methods recover secret parameters from a single layer of the LLaMA-3 8B model within six minutes and scale effectively to models as large as 405B, entirely bypassing integrity checks employed by systems such as Soter and TSQP.

integrityLLM inferencemodel confidentiality

SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment

Oct 22, 2025
TN
Tushar Nayan
🏛️ Florida International University | University of Illinois Urbana-Champaign

Deploying large language models (LLMs) on edge devices faces a dual challenge—vulnerability to model extraction attacks and difficulty in leveraging untrusted GPU accelerators without compromising model confidentiality. Method: This paper proposes a heterogeneous secure inference framework that introduces the first threat-aware, operator-level tensor partitioning strategy: sensitive nonlinear layers execute inside Intel SGX enclaves, while linear layers are encrypted and outsourced to GPUs. The framework integrates encrypted tensor outsourcing, secure reconstruction mechanisms, and information-theoretic security guarantees. Contribution/Results: Evaluated on LLaMA-2, the prototype achieves strong model confidentiality with only ~1.8× inference latency overhead—significantly outperforming enclave-only baselines. It enables efficient, privacy-preserving LLM deployment on resource-constrained edge devices, offering a practical solution for privacy-sensitive applications.

Isolating privacy-critical components in TEE-GPU architectureProtecting model privacy while using untrusted acceleratorsSecuring LLMs against model extraction attacks

TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone

Nov 17, 2025
XW
Xunjie Wang
🏛️ Shanghai Jiao Tong University

To address the risks of model parameter leakage and the security-efficiency trade-off in deploying large language models (LLMs) on mobile devices, this paper proposes a lightweight secure inference framework leveraging Arm TrustZone. Methodologically, it introduces a TEE-REE co-driven architecture enabling time-division sharing of the NPU to minimize the trusted computing base; designs a pipelined parameter prefetching mechanism integrating deterministic memory access prediction, on-demand decryption, and encrypted memory management; and develops a lightweight NPU data-plane driver. Evaluated on a Rockchip platform, the framework reduces first-token latency by 90.9% and improves decoding throughput by up to 23.2%, achieving a significant balance between robust model intellectual property protection and high-performance local inference.

Enabling secure NPU time-sharing between Rich and Trusted Execution EnvironmentsResolving memory efficiency versus fast inference dilemma in TEESecuring proprietary LLMs on mobile devices against user leakage risks

Latest Papers

What's happening recently
View more

This work proposes a black-box reverse engineering method that infers architectural parameters and inference optimization strategies of closed-source large language models solely from the timing of remote API responses. By constructing fine-grained models of per-token generation latency on NVIDIA GPUs and integrating architectural space search with speculative decoding detection algorithms, the approach achieves architecture-level probing without requiring access to model weights or internal logs. Experimental results demonstrate that, on Llama-family models, the true architecture is ranked among the top ten hypotheses in over 90% of cases. Furthermore, the method successfully uncovers that Gemini Flash 2.5 employs speculative decoding with a 128K context window, substantially expanding the frontier of black-box model analysis.

inference optimizationlanguage modelsmodel architecture

This work demonstrates that subtle numerical discrepancies introduced by different components—such as inference engines, attention backends, and hardware platforms—in large language model (LLM) inference systems can serve as distinctive fingerprints, inadvertently revealing system configurations and posing security risks. The paper presents the first fingerprinting methodology based on prompt-response behavior and numerical deviation analysis, achieving high-accuracy identification of these components through empirical evaluation. The study shows that this approach reliably infers system configurations even under non-zero temperature sampling, highlighting inherent limitations in existing defense mechanisms. Furthermore, it explores potential mitigation strategies and analyzes their practical implications, underscoring the tension between deployability and robustness in securing LLM inference pipelines.

fingerprintinginference systemslarge language models

This study addresses the security–efficiency paradox arising when deploying large language models (LLMs) at the edge under resource constraints, where efficiency-oriented optimizations—such as quantization and pruning—introduce significant security and privacy vulnerabilities. The work presents the first systematic characterization of this inherent trade-off, introducing a deployment taxonomy grounded in memory walls, compute walls, and secondary walls. It formulates a unified constraint model to quantify the critical thresholds beyond which optimizations become unsafe and proposes a Security-Operational Efficiency Score (SOES) to guide deployment configurations. Through integrated analyses of model compression, threat modeling, privacy attack simulations, and multi-objective optimization, the framework jointly evaluates model accuracy, jailbreak resistance, and privacy preservation, delivering actionable guidelines and mitigation strategies for secure edge deployment of LLMs.

Edge ComputingLarge Language ModelsModel Compression

Centralized AI deployments entail exposing sensitive data and code to cloud providers, posing significant privacy and security risks. This work proposes the first end-to-end confidential AI workflow that systematically integrates CPU-based trusted execution environments (TEEs)—such as Intel TDX and AMD SEV-SNP—with GPU TEEs on NVIDIA H100/H200 platforms, ensuring data confidentiality and integrity across both virtual machine and application layers. The study identifies and mitigates novel threats, including unauthorized access to confidential VM contents by Kubernetes administrators, by leveraging remote attestation and end-to-end encrypted execution to provide strong security guarantees. The proposed framework is evaluated on an integrated Intel TDX and NVIDIA H200 platform using industry-standard benchmarks, quantifying performance overhead and demonstrating the feasibility and robustness of the approach.

Confidential AICPU/GPU TEEsLarge Language Models

Hot Scholars

CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
QZ

Qiao Zhang

Shandong University
Privacy in Machine Learning
HW

Hongyi Wu

IEEE Fellow, Professor and Department Head, ECE, The University of Arizona
Intelligent and Secure Computing and Communication Systems
SY

Shuzhou Yang

Peking University
Computer VisionMachine LearningGenerative Model
JS

Jiaqi Sun

Carnegie Mellon University
Causalitygraph representation learning