🤖 AI Summary
To address the high computational overhead, orchestration complexity, and poor model compatibility of homomorphic encryption (HE) in privacy-preserving machine learning (ML) inference within cloud-native environments, this paper proposes the first cloud-native HE inference framework tailored for ML-as-a-Service (MLaaS). The framework containerizes HE components and leverages Kubernetes for elastic scaling and cross-cluster parallel encrypted computation. It introduces three key optimizations—ciphertext packing, adaptive polynomial modulus adjustment, and operator fusion—to enable efficient end-to-end model inference directly over encrypted data. Experimental evaluation demonstrates that, compared to conventional HE pipelines, the framework achieves up to a 3.2× speedup in inference latency and reduces memory consumption by 40%. These improvements significantly enhance secure, scalable, and deployable privacy-preserving computation capabilities in zero-trust cloud environments.
📝 Abstract
As machine learning (ML) models become increasingly deployed through cloud infrastructures, the confidentiality of user data during inference poses a significant security challenge. Homomorphic Encryption (HE) has emerged as a compelling cryptographic technique that enables computation on encrypted data, allowing predictions to be generated without decrypting sensitive inputs. However, the integration of HE within large scale cloud native pipelines remains constrained by high computational overhead, orchestration complexity, and model compatibility issues.
This paper presents a systematic framework for the design and optimization of cloud native homomorphic encryption workflows that support privacy-preserving ML inference. The proposed architecture integrates containerized HE modules with Kubernetes-based orchestration, enabling elastic scaling and parallel encrypted computation across distributed environments. Furthermore, optimization strategies including ciphertext packing, polynomial modulus adjustment, and operator fusion are employed to minimize latency and resource consumption while preserving cryptographic integrity. Experimental results demonstrate that the proposed system achieves up to 3.2times inference acceleration and 40% reduction in memory utilization compared to conventional HE pipelines. These findings illustrate a practical pathway for deploying secure ML-as-a-Service (MLaaS) systems that guarantee data confidentiality under zero-trust cloud conditions.