inference on encrypted data

Design, implement, and evaluate systems that perform machine learning model inference directly on encrypted data (ciphertexts) using homomorphic encryption; this includes implementing and optimizing homomorphic arithmetic and ciphertext representations, managing noise and computation-depth tradeoffs, minimizing inference-time ciphertext overhead, and measuring inference accuracy, latency, and resource use.

inferenceonencrypteddata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a unified privacy-preserving framework based on the CKKS homomorphic encryption scheme to mitigate the risk of privacy leakage associated with processing sensitive data in plaintext during machine learning. For the first time, it enables encrypted training of both k-nearest neighbors (KNN) and linear regression, as well as encrypted inference for multilayer perceptrons, within a single system. By integrating approximation techniques to handle non-polynomial operations and effectively managing ciphertext noise to enhance computational efficiency, the framework maintains end-to-end data encryption while achieving model accuracy comparable to that of plaintext training. The results demonstrate the practical feasibility of privacy-preserving machine learning and highlight key challenges remaining in computational overhead and functional expressiveness.

data confidentialityencrypted data traininghomomorphic encryption

To address the high computational overhead, orchestration complexity, and poor model compatibility of homomorphic encryption (HE) in privacy-preserving machine learning (ML) inference within cloud-native environments, this paper proposes the first cloud-native HE inference framework tailored for ML-as-a-Service (MLaaS). The framework containerizes HE components and leverages Kubernetes for elastic scaling and cross-cluster parallel encrypted computation. It introduces three key optimizations—ciphertext packing, adaptive polynomial modulus adjustment, and operator fusion—to enable efficient end-to-end model inference directly over encrypted data. Experimental evaluation demonstrates that, compared to conventional HE pipelines, the framework achieves up to a 3.2× speedup in inference latency and reduces memory consumption by 40%. These improvements significantly enhance secure, scalable, and deployable privacy-preserving computation capabilities in zero-trust cloud environments.

Integrating encrypted computation with Kubernetes orchestration systemsOptimizing homomorphic encryption workflows for cloud ML inferenceReducing computational overhead in privacy-preserving ML services

This work addresses the dual privacy risks in cloud-based AI inference—exposure of both user inputs and model weights—and tackles the impractical computational overhead of fully homomorphic encryption (FHE). To bridge this gap, the authors propose a co-design paradigm that integrates FHE with AI inference through a novel “meet-in-the-middle” optimization framework. This approach jointly tailors cryptographic primitives and neural network architectures: on one side, it customizes an FHE scheme and compiler to align with the static structure of the inference circuit; on the other, it imposes architectural constraints on the AI model to minimize dominant homomorphic operations. The resulting synergy substantially reduces FHE inference costs, offering a practical and efficient pathway toward privacy-preserving AI inference.

AI inferencecloud computingfully homomorphic encryption

This work addresses the scalability limitations of existing fully homomorphic encryption (FHE)-based approaches for privacy-preserving large language model inference, particularly their inefficiency with long sequences and excessive computational overhead in nonlinear layers due to outlier values. To overcome these challenges, the authors propose an unbalanced chunked prefilling framework that encrypts only the sensitive portion of the input—specifically the last 128 tokens—under the CKKS scheme, while leveraging a hybrid plaintext-ciphertext computation paradigm. The approach integrates a novel homomorphic matrix multiplication algorithm, an efficient polynomial evaluation method, and training-free mitigation strategies such as token shifting and rotation to suppress outliers. Evaluated on an 8×RTX 4090 GPU cluster, the system achieves the first end-to-end private inference for Llama-2-7B on 4096-token inputs, requiring only 85 seconds and 33 seconds per output token for summarization and generation tasks, respectively.

fully homomorphic encryptioninput token lengthlarge language models

This work systematically compares fully homomorphic encryption (FHE) and garbled circuits (GC) for privacy-preserving machine learning inference, focusing on secure neural network evaluation under joint data and model confidentiality. Under a unified threat model, we conduct the first quantitative, five-dimensional comparison—covering inference error, end-to-end latency, memory footprint, communication rounds, and bandwidth—using CKKS (Microsoft SEAL) for FHE and TinyGarble2.0 (Intel Labs) for GC, with a two-layer neural network as the benchmark. Results show that GC achieves lower latency and memory overhead, whereas FHE enables single-round, non-interactive inference, offering superior communication efficiency and stronger privacy guarantees. The study uncovers a fundamental trade-off among interactivity, computational efficiency, and security strength, providing empirical guidance for scenario-driven selection of privacy-enhancing technologies in practical deployments.

Assessing computational efficiency under semi-honest threat modelComparing FHE and GC for secure neural network inferenceEvaluating performance trade-offs in privacy-preserving machine learning

Latest Papers

What's happening recently
View more

This work addresses the challenge of deploying machine learning in the cloud when legal constraints prohibit sharing sensitive data by proposing a privacy-preserving image classification method based on fully homomorphic encryption (FHE). The authors redesign convolutional neural network architectures to operate efficiently in the encrypted domain, extending the TenSEAL framework to support multi-channel color images, multi-layer convolutions, and average pooling operations, with an efficient inference implementation built upon Microsoft SEAL. Experimental results demonstrate that the proposed approach achieves classification accuracy on MNIST, Kuzushiji-MNIST, Fashion-MNIST, and CIFAR-10 datasets nearly matching that of plaintext models, while maintaining relatively low computational overhead, thereby significantly advancing the practical applicability of FHE to real-world computer vision tasks.

CNNComputer VisionEncrypted Inference

As security demands increase, the importance of secure computation technologies grows, yet these technologies can often seem overwhelming to practitioners. Furthermore, many approaches focus only on a single technology, potentially overlooking superior alternatives. This work aims to address the issue of selecting the right technology for secure computation by presenting a comparative analysis of two highly relevant cryptographic methods and their software implementations, with a particular focus on machine learning. Firstly, we provide a theoretical summary and comparison of the secure computation paradigms of secure multi-party computation (SMPC) and fully homomorphic encryption (FHE). We outline the advantages and limitations of the protocols, as well as the relevant open-source software implementations. Secondly, we present the results of extensive benchmarking of the main software frameworks identified for machine learning operations and models. Regarding the current state of the art in FHE, we observe that it outperforms SMPC for regressions. Additionally it may be faster for simple dense networks using GPUs or Hybrid Models. Conversely, SMPC showed superior performance for complex models such as CNNs. Our results should pave the way for more technology-agnostic benchmarking of secure computation technologies for machine learning, providing guidance for practitioners looking to adopt these technologies.

FHEmachine learningsecure computation

This work addresses the privacy risks in federated learning, where model updates may inadvertently leak sensitive user information. To mitigate this, the authors propose a novel federated learning framework that integrates homomorphic encryption with differential privacy. Specifically, homomorphic encryption enables secure aggregation of model updates in encrypted form, while differential privacy introduces calibrated noise to these updates, thereby providing dual-layer privacy protection without requiring clients to upload raw local data. The efficacy of the approach is empirically validated on real-world datasets—including Framingham, Pima Indians Diabetes, and Bank Marketing—demonstrating its ability to maintain high model accuracy while ensuring strong privacy guarantees in sensitive domains such as healthcare and finance. The study also systematically investigates the impact of data heterogeneity on performance and presents corresponding optimization strategies.

Data PrivacyFederated LearningModel Updates

Hot Scholars

QL

Qian Lou

Assistant Professor of Computer Science, University of Central Florida,
Secure & Private ComputingAI InfrastructureMachine Learning Systems
AK

Afsana Khan

PhD Candidate, Maastricht University
Federated LearningArtificial IntelligenceData FusionInternet of Things
BL

Bo Luo

Professor, The University of Kansas
SecurityPrivacyAI Security
YS

Young-Sik Kim

Professor, Depart. Electrical Engineering and Computer Science, DGIST
Post-Quantum CryptographyFully Homomorphic EncryptionPrivacy-Preserving Machine Learning