๐ค AI Summary
This study addresses the prohibitive overhead of nonlinear operations in secure Transformer inference and the accuracy degradation caused by compounding independent compression strategies. We propose a unified optimization framework centered on polynomial degree, which formulates secure compression as a structured sparsity problem to jointly optimize approximation, pruning, and model architecture. By employing computation-aligned null-degree removal and low-degree Softmax/GeLU approximation-aware training, the method effectively prevents error accumulation. Experimental results demonstrate that this approach achieves 2.29ร to 6.63ร speedups across vision and language tasks. Notably, it attains 92.68% accuracy on BERT/SST-2 with an inference latency of 110.55 seconds, yielding an accuracyโlatency trade-off that significantly surpasses existing hybrid protocols.
๐ Abstract
Secure Transformer inference protects sensitive inputs but incurs substantial cryptographic overhead, with nonlinear operations such as Softmax and GeLU becoming major bottlenecks. Existing compression methods reduce nonlinear complexity, sequence-dependent computation, or model structure through separately defined compression variables. Under aggressive compression, however, these independently optimized perturbations can accumulate: at a matched compression level, stacking representative approximation, token-pruning, and model-pruning methods reduces ViT-S accuracy from 80.20% to 76.41%. We introduce DegreeSpar, which formulates secure Transformer compression as structured sparsification over nonlinear polynomial degrees. Polynomial degree directly controls the cost of secure nonlinear evaluation, while computation-aligned zero-degree structures expose token-level and model-dimension computation as removable within the same optimization space. DegreeSpar further incorporates approximation-aware training for low-degree Softmax and GeLU, enabling aggressive degree reduction and creating the optimization headroom required for structured computation removal. Across vision and language Transformers, DegreeSpar consistently improves the accuracy-latency trade-off across model scales, tasks, and sequence lengths, achieving speedups from 2.29x to 6.63x over the corresponding baselines. Under the same network setting, DegreeSpar achieves 92.68% accuracy on BERT/SST-2 in 110.55 s, compared with 92.66% in 167.26 s for CipherPrune, the closest prior hybrid secure-inference approach. These results establish structured polynomial degree as an effective shared optimization space for secure Transformer compression.