Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic research on large language model quantization under zero-knowledge proofs (ZKPs), where balancing privacy and efficiency remains challenging. We present the first systematic evaluation of ZK-friendly quantization strategies by leveraging finite-field arithmetic modeling and mixed-precision quantization techniques. Our analysis reveals that activation precision is more sensitive than weight precision, with RMSNorm identified as a critical bottleneck. Accordingly, we propose an operator-aware precision selection method that overcomes the limitations of conventional low-bit heuristics. The findings demonstrate that indiscriminately reducing bit-width does not linearly decrease proving overhead, whereas selectively increasing the precision of bottleneck layers can restore near-baseline utility. This work provides both theoretical insights and practical guidance for efficient ZK inference deployment.
📝 Abstract
Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.
Problem

Research questions and friction points this paper is trying to address.

Zero-Knowledge Proofs
Large Language Models
Quantization
Verifiable Inference
Proving Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Knowledge Proofs
ZK-Friendly Quantization
Verifiable LLM Inference
Lookup Table Approximation
Operator-Aware Precision
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.