Score
Designs and builds systems that split large images into multiple encrypted sub-images and represent each sub-image as a separate ciphertext within a multi-ciphertext (FHE) processing framework, enabling parallel encrypted-image computation. Analyzes and optimizes sub-image partitioning, ciphertext-level parallelism, bootstrapping placement, and FHE parameter and key-size tradeoffs to minimize computational, memory, and communication overhead.
为解决高效处理加密细粒度数据的问题,本文提出PixCrypt,一种基于缓存的加速机制,通过减少密文生成和使用系数级操作来加快全同态加密速度。
This work addresses the substantial computational overhead of existing homomorphic encryption schemes when processing high-resolution images. To mitigate this, the authors propose a multi-ciphertext privacy-preserving framework that enables parallel computation through image tiling and encrypted-domain convolution optimization via repeated packing. An efficient Sobel operator is specifically designed to support gradient computation on encrypted data. Key innovations include a tiling strategy that reduces ciphertext parameter size, a novel bootstrapping placement mechanism to minimize computational cost, and a sign-function-based polynomial approximation for reciprocal computation that enhances gradient direction accuracy. Experimental results demonstrate that the proposed approach significantly reduces the complexity of encrypting high-resolution images and computing their gradients, while simultaneously improving both client-side efficiency and server-side processing performance.
Multi-bit TFHE suffers from narrow numerical representation ranges and low computational efficiency, hindering large-scale privacy-preserving computation in cloud environments. This paper proposes Taurus, a hardware-accelerated architecture enabling, for the first time, private inference of large language models (e.g., GPT-2) using multi-bit TFHE. Taurus jointly enhances ciphertext computation throughput and numerical precision through a customized FFT unit, key-value reuse mechanism, memory bandwidth optimization, and a compiler supporting operation deduplication. Experimental results demonstrate that Taurus achieves up to 2600× and 1200× speedup over CPU and GPU baselines, respectively, and outperforms the state-of-the-art TFHE accelerator by 7×. These advances significantly advance the practical deployment of high-precision, wide-dynamic-range privacy-preserving computation.
To address the high computational overhead and poor scalability of fully homomorphic encryption (FHE) for private inference, this paper introduces the first frequency-domain private inference paradigm. Our method integrates the discrete cosine transform (DCT) into the FHE inference pipeline, enabling lightweight activation functions and optimized bootstrapping directly on JPEG-compatible frequency-domain representations. Leveraging the energy concentration property of DCT coefficients in low-frequency bands, we design a low-frequency-aware training strategy and a dedicated frequency-domain neural network architecture, coupled with dynamic bootstrapping scheduling. Experiments on ImageNet demonstrate that inference time is reduced from 12.5 to 2.5 hours (5.3× speedup), exhibiting superlinear scalability. Moreover, ciphertext noise is significantly suppressed, yielding improved prediction robustness. This work establishes a novel pathway toward high-accuracy, high-efficiency privacy-preserving inference.
Fully homomorphic encryption (FHE) faces critical challenges in deep learning inference—including prohibitive computational overhead, inefficient vector packing, uncontrolled noise growth, and lack of high-level programming abstractions. Method: This paper introduces the first end-to-end FHE inference framework for PyTorch. It proposes a novel single-shot multi-channel convolutional packing strategy; designs a noise-aware, fully automated bootstrap placement and dynamic scaling mechanism; and implements automatic compilation and optimized scheduling—from PyTorch models to CKKS circuits—including relinearization and rescaling. Contribution/Results: It achieves the first complete FHE inference for ResNet-50 (ImageNet) and YOLO-v1 (139M parameters, high-resolution input); ResNet-20 inference is 2.38× faster than the state of the art; and the implementation is open-sourced.
为解决全同态加密(FHE)内存瓶颈问题,提出BXT框架,通过压缩、序列化、延迟生成种子及剪枝等方法优化内存使用,提高计算效率。
Existing fully homomorphic encryption (FHE) compilers perform optimizations only at the ciphertext level, which is insufficient to eliminate polynomial-level redundant computations across ciphertexts, thereby limiting performance gains. This work proposes Recifhe, a multi-level FHE compiler that introduces, for the first time, polynomial-level optimization. Recifhe transforms conventional programs into FHE-compatible ones while integrating the RNS-CKKS scheme, achieving finer-grained computation reduction through ciphertext management, global program transformation, and cross-ciphertext polynomial redundancy elimination. Compared to approaches restricted to ciphertext-level optimization, Recifhe delivers an average speedup of 1.25×.
Fully homomorphic encryption (FHE) remains challenging to execute efficiently due to its substantial computational and memory overhead, exacerbated by a longstanding disconnect between cryptographic optimizations and hardware design. This work proposes a memory-centric, architecture-aware hardware-software co-optimization approach that tightly integrates the CKKS scheme with accelerator design, drastically reducing off-chip memory accesses and temporary data storage. Key innovations include an accelerator-oriented fine-grained coefficient-to-slot transformation, plaintext compression, intermediate modulus upping, dedicated on-chip buffers, and extended functional units. These synergistic enhancements enable, for the first time, sub-millisecond CKKS bootstrapping. Compared to the state-of-the-art FHE accelerators, the proposed design achieves 1.38× to 8.74× higher performance per unit area.
为解决云LLM服务中的隐私风险,本文提出Odin系统,通过共同设计密文打包和模型执行优化CKKS-based LLM推理,显著减少计算时间和内存消耗。
This work addresses the vulnerability of CKKS homomorphic encryption to transient hardware faults on CPUs, which can cause silent data corruption, while existing fault-tolerance mechanisms incur prohibitive overhead. To mitigate this, the authors propose a three-tiered, low-overhead fault-tolerance scheme comprising modulus-aware bucket checking, intra-operator fused verification, and inter-operator check fusion. This approach ensures end-to-end error detection while substantially reducing overheads associated with modular arithmetic, memory access, and execution. Implemented atop OpenFHE, the solution achieves 100% error detection across 150,000 non-crash fault injections, with runtime overhead ranging from 6.0% to 8.4% (averaging 6.8%)—a 4.9× reduction in protection cost compared to conventional methods.