Score
Designs, implements, and evaluates runtime architectures, protocols, and tooling to perform large language model inference inside secure execution contexts (e.g., TEEs) or on-device secure environments so that model weights and user data are protected from OS compromise. This work covers secure enclave integration with host OS services (memory management, attestation), device-side secure inference flows, and resource/Isolation strategies to minimize latency while preserving confidentiality and integrity.
This work addresses a critical security vulnerability in deploying large language models (LLMs) on third-party devices, where existing approaches often rely on partially trusted execution environments (TEEs) and precomputed noise to protect intellectual property. We demonstrate for the first time that the common practice of statically reusing secret bases—introduced to accelerate encrypted inference in mainstream LLMs—leaks key information. Leveraging cryptanalysis, linear algebraic reconstruction, and adversarial queries, we develop two novel attacks that reverse-engineer TEE protocols and encrypted inference pipelines. Our methods recover secret parameters from a single layer of the LLaMA-3 8B model within six minutes and scale effectively to models as large as 405B, entirely bypassing integrity checks employed by systems such as Soter and TSQP.
Deploying large language models (LLMs) on edge devices faces a dual challenge—vulnerability to model extraction attacks and difficulty in leveraging untrusted GPU accelerators without compromising model confidentiality. Method: This paper proposes a heterogeneous secure inference framework that introduces the first threat-aware, operator-level tensor partitioning strategy: sensitive nonlinear layers execute inside Intel SGX enclaves, while linear layers are encrypted and outsourced to GPUs. The framework integrates encrypted tensor outsourcing, secure reconstruction mechanisms, and information-theoretic security guarantees. Contribution/Results: Evaluated on LLaMA-2, the prototype achieves strong model confidentiality with only ~1.8× inference latency overhead—significantly outperforming enclave-only baselines. It enables efficient, privacy-preserving LLM deployment on resource-constrained edge devices, offering a practical solution for privacy-sensitive applications.
To address the risks of model parameter leakage and the security-efficiency trade-off in deploying large language models (LLMs) on mobile devices, this paper proposes a lightweight secure inference framework leveraging Arm TrustZone. Methodologically, it introduces a TEE-REE co-driven architecture enabling time-division sharing of the NPU to minimize the trusted computing base; designs a pipelined parameter prefetching mechanism integrating deterministic memory access prediction, on-demand decryption, and encrypted memory management; and develops a lightweight NPU data-plane driver. Evaluated on a Rockchip platform, the framework reduces first-token latency by 90.9% and improves decoding throughput by up to 23.2%, achieving a significant balance between robust model intellectual property protection and high-performance local inference.
This work proposes a black-box reverse engineering method that infers architectural parameters and inference optimization strategies of closed-source large language models solely from the timing of remote API responses. By constructing fine-grained models of per-token generation latency on NVIDIA GPUs and integrating architectural space search with speculative decoding detection algorithms, the approach achieves architecture-level probing without requiring access to model weights or internal logs. Experimental results demonstrate that, on Llama-family models, the true architecture is ranked among the top ten hypotheses in over 90% of cases. Furthermore, the method successfully uncovers that Gemini Flash 2.5 employs speculative decoding with a 128K context window, substantially expanding the frontier of black-box model analysis.
This work demonstrates that subtle numerical discrepancies introduced by different components—such as inference engines, attention backends, and hardware platforms—in large language model (LLM) inference systems can serve as distinctive fingerprints, inadvertently revealing system configurations and posing security risks. The paper presents the first fingerprinting methodology based on prompt-response behavior and numerical deviation analysis, achieving high-accuracy identification of these components through empirical evaluation. The study shows that this approach reliably infers system configurations even under non-zero temperature sampling, highlighting inherent limitations in existing defense mechanisms. Furthermore, it explores potential mitigation strategies and analyzes their practical implications, underscoring the tension between deployability and robustness in securing LLM inference pipelines.
This study addresses the security–efficiency paradox arising when deploying large language models (LLMs) at the edge under resource constraints, where efficiency-oriented optimizations—such as quantization and pruning—introduce significant security and privacy vulnerabilities. The work presents the first systematic characterization of this inherent trade-off, introducing a deployment taxonomy grounded in memory walls, compute walls, and secondary walls. It formulates a unified constraint model to quantify the critical thresholds beyond which optimizations become unsafe and proposes a Security-Operational Efficiency Score (SOES) to guide deployment configurations. Through integrated analyses of model compression, threat modeling, privacy attack simulations, and multi-objective optimization, the framework jointly evaluates model accuracy, jailbreak resistance, and privacy preservation, delivering actionable guidelines and mitigation strategies for secure edge deployment of LLMs.
Centralized AI deployments entail exposing sensitive data and code to cloud providers, posing significant privacy and security risks. This work proposes the first end-to-end confidential AI workflow that systematically integrates CPU-based trusted execution environments (TEEs)—such as Intel TDX and AMD SEV-SNP—with GPU TEEs on NVIDIA H100/H200 platforms, ensuring data confidentiality and integrity across both virtual machine and application layers. The study identifies and mitigates novel threats, including unauthorized access to confidential VM contents by Kubernetes administrators, by leveraging remote attestation and end-to-end encrypted execution to provide strong security guarantees. The proposed framework is evaluated on an integrated Intel TDX and NVIDIA H200 platform using industry-standard benchmarks, quantifying performance overhead and demonstrating the feasibility and robustness of the approach.