Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference

arXiv cs.AI Papers

Summary

This paper presents the first systematic study of zero-knowledge-friendly quantization for LLMs, formalizing the concept and evaluating nine models (including Qwen2.5-14B and Qwen3-30B-A3B) across weight, activation, and nonlinear lookup table precision. It finds that activation precision and RMSNorm inverse-square-root lookups dominate ZK proving costs and that conventional low-bit quantization heuristics do not translate to proportional proving savings.

arXiv:2609.36437v1 Announce Type: new Abstract: Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:45 AM

# Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
Source: [https://arxiv.org/html/2609.36437](https://arxiv.org/html/2609.36437)
Yupeng ZhangXiaojing LiaoAffiliation:Siebel School of Computing and Data ScienceAffiliation:University of Illinois Urbana\-ChampaignEmail:[\{tuyoon2,zhangyp,xjliao\}@illinois\.edu](mailto:)

###### Abstract

Zero\-knowledge \(ZK\) proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training\-data usage or LLM inference\-time behavior, while model providers must protect proprietary model parameters\. However, despite the growing interest in ZK\-LLMs, the understanding of ZK\-friendly quantization remains limited\. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference\. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints\. Understanding ZK\-friendly quantization is therefore essential for making ZK\-LLMs practical\. In this work, we present the first systematic study of ZK\-friendly quantization for LLMs\. We first formalize the definition of ZK\-friendly quantization, capturing the properties required for ZK proof generation\. We then evaluate nine language models, including Qwen2\.5\-14B and the mixture\-of\-experts model Qwen3\-30B\-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision\. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation\. Also, we identify RMSNorm inverse\-square\-root lookups as a recurring bottleneck in several large models and recover near\-baseline utility by selectively increasing precision only at the bottleneck\. Finally, we show that reducing bit\-width or lookup\-table size does not necessarily yield proportional end\-to\-end proving savings, showing that conventional low\-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator\-aware precision selection\.

## 1Introduction

Large language models \(LLMs\) are increasingly deployed through opaque cloud services, yet users, auditors, and regulators often lack practical mechanisms to verify what computation was performed, which model was used, or whether specific governance\-relevant claims hold\. Zero\-knowledge \(ZK\) proofs offer a promising path toward verifiable private LLM inference\([Sun et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib32);[Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34);[Xie et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib16);[Chen et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib33);[Liao et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib22);[Wang et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib40);[Gailly et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib42)\): a provider can prove that an inference\-time computation or auditing statistic was correctly computed without revealing proprietary model weights or sensitive operational details\. However, proving LLM inference directly is costly because ZK proof systems operate over finite fields, while modern transformers rely on floating\-point arithmetic, nonlinear functions, normalization, and softmax operations that are expensive to arithmetize\. Compared with conventional machine learning \(ML\) quantization, ZK\-friendly quantization is therefore not only a compression technique; it determines the arithmetic structure, constraint complexity, and proving cost of the resulting ZK\-LLM system\.

Despite rapid progress in ZK\-ML and LLM quantization, there is still no formal definition of ZK\-friendly quantization or a systematic understanding of which quantization schemes are actually suitable for ZK\-LLM\. Prior ZK\-LLM systems typically adopt a specific integer representation or lookup approximation as an implementation choice, often validating it on a limited model family or scale\([Sun et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib32);[Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34)\), while conventional LLM quantization benchmarks optimize for deployment efficiency rather than proof compatibility\([Yang et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib6);[Huang et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib29)\)\. This leaves open a fundamental research gap: we do not yet know how ZK\-friendly quantization mechanisms affect ZK proof generation across models, tasks, and operator types\. This gap matters because, in the ZK setting, quantization directly determines the arithmetic structure of the proved computation, including its constraint complexity, lookup requirements, and prover cost\. In addition, these trade\-offs directly affect governance applications that must preserve model behavior while remaining practical to prove\. We demonstrate this in a ZK training\-data membership case study based on Min\-K% inference \(Appendix[H](https://arxiv.org/html/2609.36437#A8)\)\. Note that although this study focuses on ZK\-LLMs, these considerations may also extend to other cryptographic approaches for private or verifiable inference, such as secure multiparty computation\([Yao, 1982](https://arxiv.org/html/2609.36437#bib.bib17);[Hou et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib19)\)and fully homomorphic encryption\([Gentry, 2009](https://arxiv.org/html/2609.36437#bib.bib18)\), as they often represent computations over finite fields or polynomial rings and therefore face similar challenges\. Understanding ZK\-friendly quantization is essential for making proof systems practical\.

Contributions\. This study makes three major contributions toward understanding and designing ZK\-friendly quantization for LLMs\. First, we formalize*ZK\-friendly quantization*as a joint utility\-proof\-efficiency problem, capturing the constraints imposed by bounded integer representations, fixed rescaling rules, static calibration, and lookup\-based nonlinear computation\. Building on this formulation, we develop a scalable methodology that enables systematic study across model scales and architectures\. Second, we show that the utility bottlenecks of ZK\-friendly quantization differ substantially from those of traditional LLM quantization\. Across nine language models, we find that model utility is markedly more sensitive to activation precision than weight precision, while fixed\-precision nonlinear lookup table \(LUT\) can dominate utility degradation even when linear quantization remains accurate\. Also importantly, we identify the inverse\-square\-root LUT in RMSNorm as a recurring bottleneck in several large models and show that selectively increasing its LUT precision can recover near\-baseline utility without uniformly increasing precision across all nonlinearities\. Third, we characterize how these quantization choices translate into realized ZK proving cost\. Using direct proof\-system measurements, we find that reducing weight/activation bit\-width or LUT size does not necessarily yield proportional ZK prover savings\. These results motivate heterogeneous, operator\-aware precision choices rather than uniformly minimizing bit\-width, and expose the utility\-proving\-cost trade\-offs that determine practical ZK\-LLM inference\.

## 2Background and Related Work

##### Zero\-Knowledge Proofs\.

Zero\-knowledge proofs \(ZKPs\) are cryptographic protocols that allow a prover to convince a verifier that a computation was executed correctly without revealing the secret witness\. In ML, a model provider can use a ZKP to prove that an inference or other properties, such as fairness, robustness, and explainability\([Yadav et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib37);[Zhang et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib38);[Yadav et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib39)\), are honestly computed on the secret model\([Feng et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib36);[Lee et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib2);[Chen et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib33)\)\. ZKP protocols represent computations as arithmetic operations defined over a large prime field𝔽p\\mathbb\{F\}\_\{p\}\. The natively supported operations include modular additions,a\+bmodpa\+b\\mod p, and modular multiplications,a×bmodpa\\times b\\mod pfora,b∈𝔽pa,b\\in\\mathbb\{F\}\_\{p\}\. Modern proof systems also support lookup arguments\([Setty et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib43), e\.g\.,\)that check whether an input\-output pair belongs to a fixed table, denoted asy←𝖫𝗈𝗈𝗄𝗎𝗉⁡\(T,x\)y\\leftarrow\\mathsf\{Lookup\}\(T,x\)\. We use the term “BB\-bit LUT” to denote a table with2B2^\{B\}entries\. The proving cost depends on the table size and the total number of lookup operations\. Meanwhile, divisions are not native field operations\. Also, representing floating\-point arithmetic or continuous functions inside a proof requires extra constraints, such as range checks, bit decomposition, or lookup arguments\. These overheads make ZK\-LLM inference expensive\.

##### Post\-Training Quantization\.

Post\-training quantization \(PTQ\) is an approach to reduce the storage and memory cost of LLMs by lowering the bit\-width of weights and/or activations\([Dettmers et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib5);[Yang et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib6);[Dettmers et al\., 2022](https://arxiv.org/html/2609.36437#bib.bib9);[Xiao et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib7);[Jacob et al\., 2018](https://arxiv.org/html/2609.36437#bib.bib8);[Frantar et al\., 2022](https://arxiv.org/html/2609.36437#bib.bib23)\)\. In a basic form of quantization, a tensorXXis assigned a scales=max⁡\|X\|2b−1−1,s=\\frac\{\\max\|X\|\}\{2^\{b\-1\}\-1\},wherebbis the target bit width\. Each tensor valuexxis then mapped to a signedbb\-bit integer by roundingx/sx/sand clipping to the representable range\. Static quantization computes the scale with calibration samples, whereas dynamic quantization computes it from runtime values\([Xie et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib16);[Xiao et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib7)\)\. A line of work\([Kurtic et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib15);[Yang et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib6);[Yao et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib31);[Huang et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib29);[Gong et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib30)\)has explored how different quantization schemes affect performance and cost efficiency\. LLMCBench\([Yang et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib6)\)presents a benchmark for LLM compression algorithms, including seven modern quantization methods\([Ma et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib24);[Sun et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib25);[Frantar and Alistarh, 2023](https://arxiv.org/html/2609.36437#bib.bib26);[Frantar et al\., 2022](https://arxiv.org/html/2609.36437#bib.bib23);[Xiao et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib7);[Lin et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib27);[Shao et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib28)\), across multiple metrics such as compression performance\.[Kurtic et al\. \(2025\)](https://arxiv.org/html/2609.36437#bib.bib15)investigate the accuracy and cost trade\-offs of specific quantization formats, such as W4A16, on both academic and real\-world benchmarks, and show that W4A16\-INT is one of the most cost\-efficient settings\. However, these benchmarks are designed for general ML environments, rather than the arithmetic constraints and proving costs of ZK protocols\. In ZK protocols, rescaling and nonlinear operations must also be encoded using field arithmetic, lookup arguments, or other proof\-compatible constructions\.

##### Quantization in ZK\-ML\.

Early ZK\-ML frameworks targeted relatively simple neural networks\([Chen et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib33);[Liu et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib13);[Lee et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib2);[Feng et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib36)\)and recent systems such as zkGPT\([Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34)\)and zkLLM\([Sun et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib32)\)extend proof generation to LLM inference with protocol\-level optimizations\. In many existing ZK machine learning schemes, quantization is an essential building block, as otherwise the overhead of simulating floating\-point arithmetic would be too high in ZK protocols\. Many of the systems\([Liu et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib13);[Feng et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib36);[Sun et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib32);[Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34)\)use integer representations by adopting an affine quantization scheme\([Jacob et al\., 2018](https://arxiv.org/html/2609.36437#bib.bib8)\)\. Rescaling is performed after multiplication to avoid overflow\([Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34)\)in𝔽p\\mathbb\{F\}\_\{p\}\. To avoid expensive operations, they also use approximation techniques such as LUTs or polynomial approximations\. Specifically, zkGPT\([Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34)\)uses high\-precision quantization and integer\-ratio rescaling for GPT\-2 Small\([Radford et al\., 2019](https://arxiv.org/html/2609.36437#bib.bib14)\), while zkLLM\([Sun et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib32)\)uses specialized attention arithmetization and lookup arguments for nonlinear operations\. However, the quantization was only tested on the small model evaluated in each paper, and the techniques may not generalize to other models\. Moreover, existing implementations of ZK machine learning are optimized for proof generation but not for fast plaintext inference\. They are mostly written in C\+\+ and Rust, and lack support for LLM acceleration\. Therefore, it is inefficient to test different quantization methods on different models\. Our paper addresses this gap by developing a framework to implement different ZK\-friendly quantization schemes and benchmarking their performance\. This framework is important to develop new ZK\-ML schemes in the future for larger models and different architectures \(e\.g\., MoE\) that are currently not supported by existing ZK\-ML systems\.

## 3ZK\-friendly Quantization

In this section, we first introduce the terminology and syntax of ZK\-friendly quantization in general\. We then describe the properties that make such a scheme useful in practice\.

Letℳ\\mathcal\{M\}be a large language model with real\-valued weights and forward pass computations \(e\.g\., floating\-point numbers and arithmetic\)\. For an input promptxx,ℳ⁡\(x\)→y\\mathcal\{M\}\(x\)\\to ydenotes the process of running the model and producing an outputyy, such as next\-token logits or a sequence of tokens\. The main object of a ZK\-friendly quantization scheme is a quantization procedure

𝖰𝗎𝖺𝗇𝗍𝗂𝗓𝖾⁡\(ℳ,𝔽p\)→ℳQ,\\mathsf\{Quantize\}\(\\mathcal\{M\},\\mathbb\{F\}\_\{p\}\)\\to\\mathcal\{M\}\_\{Q\},which takes a real\-valued modelℳ\\mathcal\{M\}and outputs a quantized modelℳQ\\mathcal\{M\}\_\{Q\}whose inference computation is represented over𝔽p\\mathbb\{F\}\_\{p\}of the ZK protocol\. Any numerical formats, scaling rules, approximation ranges, LUTs and public constants are treated as part of the quantization scheme\.

Specifically, we consider two basic quantization algorithms:

- •𝖫𝗂𝗇𝖰𝗎𝖺𝗇𝗍⁡\(𝖫,𝔽p\)→𝖫~\\mathsf\{LinQuant\}\(\\mathsf\{L\},\\mathbb\{F\}\_\{p\}\)\\to\\widetilde\{\\mathsf\{L\}\}takes a real\-valued linear tensor operation𝖫\\mathsf\{L\}, such as matrix multiplication, convolution andQ,K,VQ,K,Vprojections, and replaces it with a quantized operation𝖫~\\widetilde\{\\mathsf\{L\}\}\.
- •𝖭𝗈𝗇𝗅𝗂𝗇𝖰𝗎𝖺𝗇𝗍⁡\(g,𝔽p\)→g~\\mathsf\{NonlinQuant\}\(g,\\mathbb\{F\}\_\{p\}\)\\to\\widetilde\{g\}takes a real\-valued nonlinear operationgg, such as softmax, GeLU and normalization with inverse square root, and replaces it with a quantized arithmetizationg~\\widetilde\{g\}\.

With these operations, the real\-valued computation ofℳ\\mathcal\{M\}is replaced by a computation represented over𝔽p\\mathbb\{F\}\_\{p\}\. Then the inference under the quantized model can be represented as:ℳQ​\(x\)→yQ\.\\mathcal\{M\}\_\{Q\}\(x\)\\to y\_\{Q\}\.

##### ZK efficiency\.

In principle, any function that is computable by a computer can be realized in a ZK protocol\. However, the overhead of modeling the function could be high\. Therefore, a useful quantization scheme should improve the efficiency of the ZK protocol on the quantized model, compared to the original model and computation\. Efficiency measures include prover time, proof size and verifier time\. All three measures depend on the size of the relationℛ\\mathcal\{R\}to be proven in ZK, which we use to formalize this property\.

We defineℛ=\{\(x,y\);ℳ\|ℳ\(x\)=y\}\\mathcal\{R\}=\\\{\(x,y\);\\mathcal\{M\}\|\\mathcal\{M\}\(x\)=y\\\}as the original relation of LLM inference with real\-valued modelℳ\\mathcal\{M\}andℛQ,p=\{\(x,yQ\);ℳQ\|ℳQ\(x\)=yQ\}\\mathcal\{R\}\_\{Q,p\}=\\\{\(x,y\_\{Q\}\);\\mathcal\{M\}\_\{Q\}\|\\mathcal\{M\}\_\{Q\}\(x\)=y\_\{Q\}\\\}as the quantized relation with the quantized model and computations\. For a ZK protocolΠ\\Pi, we use\|ℛ\|Π\|\\mathcal\{R\}\|\_\{\\Pi\}to denote the size of the relation underΠ\\Pi, which is usually measured by the number of arithmetic operations over𝔽p\\mathbb\{F\}\_\{p\}and lookup operations\.

###### Definition 1\.

A quantization scheme hasΔ\\Delta\-ZK efficiency with respect to a ZK protocolΠ\\Pi, whereΔ=\|ℛQ,p\|Π\|ℛ\|Π\\Delta=\\frac\{\|\\mathcal\{R\}\_\{Q,p\}\|\_\{\\Pi\}\}\{\|\\mathcal\{R\}\|\_\{\\Pi\}\}\.

A quantization scheme is interesting only ifΔ<1\\Delta<1, which is often significantly smaller than 1 in practice\. The speedup is also not smooth, as we will demonstrate in §[5\.2](https://arxiv.org/html/2609.36437#S5.SS2)\.

##### Fidelity\.

Quantization should not substantially degrade the quality of the model; otherwise the ZK proof would be meaningless in practice\.

###### Definition 2\.

Letddbe a discrepancy measure between the two outputs, and let𝒟\\mathcal\{D\}be an input distribution of the prompts or token sequences\. The fidelity error ofQQonℳ\\mathcal\{M\}is defined as:

δd​\(ℳ,Q,𝒟\)=𝔼x∼𝒟​\[d⁡\(ℳ⁡\(x\),ℳQ​\(x\)\)\]\.\\delta\_\{d\}\(\\mathcal\{M\},Q;\\mathcal\{D\}\)=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}\\left\[d\\big\(\\mathcal\{M\}\(x\),\\mathcal\{M\}\_\{Q\}\(x\)\\big\)\\right\]\.

Different choices ofddcapture different levels of fidelity\. At the numerical level,ddmay measure layer\-wise error or logit distance\. At the distributional level, it may measure KL divergence between next\-token distributions or agreement among top\-ranked tokens\. At the task level, it may measure the gap in perplexity, accuracy, or another downstream metric\. We measure three levels of fidelity in §[5](https://arxiv.org/html/2609.36437#S5)\.

Static compilability\.Almost all ZK schemes cannot handle dynamic relations efficiently\. Therefore, when modeled as an arithmetic circuit/constraint system, the relation of the quantized model inference should not depend on the input known at runtime\. It does not mean that all intermediate values are known in advance\. Activations, attention scores, and other hidden states still depend on the input\. Rather, it means that the rules for representing and checking these values are fixed before the proof is generated\. This distinction is important because input\-dependent choices, such as computing new quantization parameters inside the forward pass, can substantially increase the cost of the ZK proof\.

###### Definition 3\.

A quantization scheme is statically compilable if in the quantized relationℛQ,p=\{\(x,yQ\);ℳQ\|ℳQ\(x\)=yQ\}\\mathcal\{R\}\_\{Q,p\}=\\\{\(x,y\_\{Q\}\);\\mathcal\{M\}\_\{Q\}\|\\mathcal\{M\}\_\{Q\}\(x\)=y\_\{Q\}\\\}, all arithmetic operations, public constants and lookup tables are fixed and independent of the runtime inputxx\.

With the definitions above, the goal of a ZK\-friendly quantization scheme is to minimizeΔ\\Deltawhile ensuringδd​\(ℳ,Q,𝒟\)≤ε\\delta\_\{d\}\(\\mathcal\{M\},Q;\\mathcal\{D\}\)\\leq\\varepsilonfor an acceptable threshold of fidelityε\\varepsilon\. The static compilability is implicitly captured byΔ\\Deltaas it directly affects the size of\|ℛQ,p\|\|\\mathcal\{R\}\_\{Q,p\}\|\.

Why ZK\-friendly quantization?ZK\-friendly quantization is a central systems primitive for practical zkML: by mapping floating\-point inference to bounded integer arithmetic and lookup\-compatible nonlinearities, it can substantially reduce proof complexity and proving cost\. Without ZK\-friendly quantization, the overhead of simulating floating\-point arithmetic could make the proving time more than 1200×\\timesslower, as we demonstrate in §[5\.2](https://arxiv.org/html/2609.36437#S5.SS2)\. However, prior ZK\-LLM systems largely tailor quantization to a specific model and proof stack, leaving little guidance on how precision and LUT design should transfer across architectures and scales\. We address this gap by systematically investigating the utility\-proving\-cost trade\-offs in ZK\-friendly quantizations\.

## 4Methods

##### Research Questions\.

The central research question of our study is:*How should ZK\-friendly quantization be designed to make LLM inference practical to prove while preserving the utility of the original model?*We decompose this question into three sub\-questions\. \(1\) How much utility degradation does ZK\-friendly quantization introduce when enabling LUT nonlinearities? \(2\) What are the root causes of utility degradation, and what kind of ZK\-friendly quantization best preserves model utility? \(3\) How should ZK\-friendly quantization be designed to achieve a favorable trade\-off between model utility and ZK proving cost?

##### ZK\-aware Quantization Framework\.

We develop a ZK\-aware quantization framework for a suite of Transformer models, including commonly studied model scales and architectures in recent ZK\-LLM work\([Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34);[Sun et al\., 2024](https://arxiv.org/html/2609.36437#bib.bib32)\), and MoE architectures that are widely deployed, but have not been studied in the ZK\-LLM literature\. Grounded in the criteria defined in §[3](https://arxiv.org/html/2609.36437#S3): a ZK\-friendly quantization must use bounded integer representations, fixed rescaling rules, static calibration, and fixed LUTs so that the resulting inference relation can be compiled before proof generation\. Hence, given a pretrained model and a quantization configuration, the framework is designed to instrument the Transformer forward pass at the operation level to emulate ZK\-compatible computation\. More specifically, linear tensor operations are replaced by bounded\-integer quantize\-dequantize modules, while selected nonlinear functions are replaced by fixed LUT modules\.

∙\\bulletZK\-aware quantization design space\. Our design space includes five quantization axes: weight and activation bit\-width, scale type, symmetry, activation calibration, and LUT precision\. We write WnnAmmfornn\-bit weight andmm\-bit activation quantization\. For scale type, the framework supports standard floating\-point scales, power\-of\-two scales,c1/c2c\_\{1\}/c\_\{2\}integer\-ratio rescaling, and ZK\-dedicated signed symmetric scales\([Gailly et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib42)\)\. It also supports symmetric and asymmetric quantization, as well as static\-max, static\-percentile, and dynamic activation calibration \(see Appendix[G](https://arxiv.org/html/2609.36437#A7)\)\.

Nonlinear operators introduce an additional design dimension because they are realized through LUTs rather than ordinary arithmetic\. Therefore, the framework supports operator\-specific LUT construction and substitution, allowing the input\-address precisionBBto be configured independently for each nonlinear operator type\. Throughout the experiments, we use the notation B16, for example, to denote a LUT with2162^\{16\}entries\. This is important because different nonlinearities exhibit distinct input ranges and approximation sensitivities, and therefore need not admit the same LUT configuration\. This operator\-level construction enables us to systematically explore heterogeneous LUT designs and identify which nonlinear approximations dominate the utility\-proving\-cost trade\-off \(see Appendix[C](https://arxiv.org/html/2609.36437#A3)\)\.

∙\\bulletOperation coverage for ZK\-LLMs\. We instrument the Transformer forward pass at the level of the operations that affect ZK\-compatible execution\. Learned parameters, including attention and MLP projections, the language\-modeling head, embeddings, and normalization affine parameters, are quantized according to the selected configuration\. Intermediate activations are subjected to quantize\-dequantize \(QDQ\) simulation at designated boundaries, including projection, normalization, activation\-function, and residual\-addition outputs\. Specifically, QDQ maps each tensor to the bounded integer domain implied by the target precision and scaling policy and then dequantizes it for continued execution with standard PyTorch operators\. For MoE architectures, we additionally capture expert and router projections, including functional linear calls that are not exposed asnn\.Linearmodules\. Note that for QDQ to faithfully evaluate the utility of a ZK\-compatible configuration, it must reproduce the numerical effect of the integer computation encoded by the prover\. We establish this correspondence algebraically and validate it empirically \(see details in Appendix[A](https://arxiv.org/html/2609.36437#A1)\)\.

We handle nonlinear computation separately: GeLU/SiLU, the inverse square root in LayerNorm/RMSNorm, and the exponential and reciprocal components of softmax are replaced by precomputed LUTs under the chosen operator\-specific precision\. Operations that are not directly approximated by the quantization configuration, such as attention\-score computation, context aggregation, RoPE, masking, reductions, and top\-\(k\) selection, remain in floating\-point arithmetic in the utility evaluator\. This separation makes the approximation boundary explicit: it isolates the effects of the ZK\-relevant numerical transformations while leaving unrelated computation unchanged\.

##### Scalable Estimation of ZK Proving Cost\.

Directly evaluating the ZK cost of every quantization configuration is impractical: each candidate must be realized as an integer computation graph and executed by a prover, making exhaustive exploration prohibitively expensive for large models\. We address this challenge with a two\-stage methodology that decouples large\-scale quantization exploration from proof\-system execution\. We first use the aforementioned QDQ simulation to instantiate arbitrary weight, activation, and nonlinear\-LUT precisions with standard PyTorch operators and GPU acceleration\. This allows us to explore configurations, including unconventional precisions such as INT12, without implementing a dedicated integer kernel or ZK circuit for each candidate\. Then, we map selected configurations to actual proof\-system execution using DeepProve\([Gailly et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib42)\)\. We extend DeepProve to independently control weight, activation, and nonlinear\-LUT precision and obtain direct measurements of prover time, verifier time, and proof size\. This two\-stage methodology makes it possible to study the utility\-proving\-cost design space at model scales and quantization granularities that would be impractical to explore through proof generation alone\.

##### Implementation and experiment setup\.

We implement our approach in PyTorch and Hugging Face and evaluate nine language models spanning different scales and architectures, including a Mixture\-of\-Experts model\. Models are loaded in BF16 and modified according to a specified quantization configuration\. We measure language\-modeling quality with perplexity on WikiText\-2\([Merity et al\., 2016](https://arxiv.org/html/2609.36437#bib.bib21)\)and C4\([Raffel et al\., 2020](https://arxiv.org/html/2609.36437#bib.bib20)\), downstream accuracy on ARC\-Easy and MMLU using thelm\-evaluation\-harnessmultiple\-choice protocol\([Gao et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib12)\), and token\-distribution fidelity to the BF16 model using KL divergence, Top\-5 overlap, and bottom\-20% overlap \(see Appendix[F](https://arxiv.org/html/2609.36437#A6)\)\. For proof evaluation, we report prover time, verifier time, and proof size\. Weights are quantized post hoc, while activation scales are obtained from calibration data and then fixed for evaluation\. Unless stated otherwise, we use signed symmetric per\-tensor quantization with standard scales and static\-max activation calibration over eight fixed WikiText\-2 training windows\. Auxiliary domains, including embeddings, positional representations, and residual paths, are kept at 24 bits while weight and activation precision are varied\. QDQ modules are inserted at the activation boundaries, and nonlinear operators are replaced by calibrated LUTs at the selected LUT precision\. Further implementation details and QDQ validation are provided in Appendices[A](https://arxiv.org/html/2609.36437#A1)and[B](https://arxiv.org/html/2609.36437#A2)\.

## 5Results

### 5\.1Broad Utility Study

##### R1\. Activations are more sensitive to low precision than weights

We first isolate linear quantization, which accounts for most bounded\-integer arithmetic in the ZK computation\. For activation sensitivity, we fix weight bit precision at 16 bits, and scale the bit precision of activations from 8 to 16, incrementing by 2\. For weight sensitivity, we fix activations at 16 bits, and increase the bit precision of weights from 8 to 16 by 2\.

Figure[1](https://arxiv.org/html/2609.36437#S5.F1)shows whether weight or activation precision limits utility\. In Figure[1](https://arxiv.org/html/2609.36437#S5.F1)\(a\), even at the lowest precision of 8 bits, the PPL ratio compared to the BF16 baseline remained below about 1\.2\. Figure[1](https://arxiv.org/html/2609.36437#S5.F1)\(b\) shows that most models exhibited PPL ratios above 10 at 8 bits\. Most models required at least 12 or 14 bits for the ratio to approach 1\. This sensitivity does not vary monotonically with model size, and the same trend holds across WikiText\-2 and C4\. We provide the complete results in Appendix[D](https://arxiv.org/html/2609.36437#A4)\.

These results show that activation quantization has greater impact on model utility than weight quantization\. Since LUT approximations can only introduce additional error, 8\-bit activations already leave too little utility margin for the full ZK\-friendly pipeline\. Depending on the model, activation precisions of at least 12 or 14 bits provide a practical starting point for preserving utility\.

Figure 1:Perplexity ratio under linear quantization to BF16 baseline\.
##### R2\. A uniform B16 LUT configuration does not reliably preserve utility

After identifying the linear quantization regimes, we evaluate the full ZK\-friendly pipeline by replacing nonlinear transformer operations with fixed LUTs\. This experiment tests whether utility loss comes mainly from linear quantization, nonlinear approximation, or their interaction\.

∙\\bulletFull\-pipeline evaluation\. Table[1](https://arxiv.org/html/2609.36437#S5.T1)\(a\) reports perplexity ratios for the full pipeline\. GeLU or SiLU activations, normalization inverse square roots, softmax exponentials and reciprocals, including the additional routing\-related functions in Qwen3 are replaced by aBB\-bit LUT over a calibrated input range\. We evaluateB∈\{8,12,16\}B\\in\\\{8,12,16\\\}on top of the W16A16 linear baseline\.

AtB=16B=16, LUT replacement introduces little additional degradation for GPT\-2 Small, Medium, and XL, and Qwen2\.5\-3B\. However, the ratios reach about 32 for Qwen2\.5\-14B, 130 for Qwen3\-30B\-A3B, and10310^\{3\}for Qwen2\.5\-7B\. Thus, even B16 LUTs can fail catastrophically on larger models, despite preserving utility on smaller ones\.

∙\\bulletModel and operator dependent precision requirements\. The same 16\-bit LUT preserves utility for some models and leaves significant degradation in several large models\. A fixed\-precision LUT divides its calibrated input range into uniform steps\. For narrow ranges, even a smaller table can approximate the operator well\. For wide ranges, the sameBBproduces coarse steps whose error can compound through the residual stream\. Therefore, it requires sufficient LUT resolution to approximate the original nonlinear operator over the inputs during inference\. The requirement depends on the model and operator, which motivates fine\-grained diagnosis of utility bottlenecks\. With our framework, we are able to isolate each LUT bit precision and evaluate whether targeting the bottleneck can help restore the lost utility\.

Table 1:PPL ratios relative to the W16A16 linear\-only baseline: \(a\) full\-LUT precision and \(b\) operator\-wise diagnosis on WikiText\-2\. Complete results for all models are provided in Appendix[D](https://arxiv.org/html/2609.36437#A4)\.WikiText\-2C4ModelB8B12B16B8B12B16GPT\-2 S1866\.21\.0111\.0001439\.41\.0181\.002GPT\-2 M3213\.61\.0021\.0011465\.41\.0000\.999Llama 8B\>104\>10^\{4\}1247\.91126\.8\>104\>10^\{4\}422\.2380\.5Qwen 14B\>104\>10^\{4\}\>104\>10^\{4\}32\.5\>104\>10^\{4\}\>104\>10^\{4\}18\.3Qwen3 30B\>104\>10^\{4\}148\.2131\.3\>104\>10^\{4\}78\.7120\.7LUT configurationModelFull B16RMS onlyRMS exactRMS B24Llama 8B1126\.84531126\.79720\.99931\.0002Qwen 7B1496\.90871528\.49880\.99990\.9987Qwen 14B32\.460932\.61391\.00020\.9997Qwen3 30B131\.33521\.03981\.00341\.0014\(a\) LUT precision sensitivity\(b\) Targeted precision adjustment

In \(b\), RMS only uses B16 RMS inverse\-square\-root LUTs with other nonlinear functions exact\. RMS exact and RMS B24 use B16 for the other LUTs\.

##### R3\. RMSNorm inverse\-square\-root LUTs are a recurring utility bottleneck

To identify the root cause of utility loss, we evaluate each model with a B16 LUT for one nonlinear operator\. The rest operators are present in their original functions, which isolates the effect of a specific operator\.

Table[1](https://arxiv.org/html/2609.36437#S5.T1)\(b\) shows that RMS inverse square root is a shared bottleneck of four models\. When we only use LUT for RMSNorm in Llama\-3\.1, Qwen2\.5\-7B, and 14B, the PPLs were very close to those of the full B16 pipeline\. However, we observed that the single RMSNorm LUT setting does not impact Qwen3\-30B like other models\. Nevertheless, substituting the RMSNorm LUTs with their exact functions can bring the PPL back to near baseline, suggesting that interactions between operators lead to the utility loss\.

∙\\bulletRoot cause and increasing bit precision\. We investigate the root cause of the bottleneck\. At one RMSNorm site in Qwen2\.5\-7B, the calibrated input range extends from 0 to 65,461\.31\. A B16 LUT contains 65,536 entries, so the interval between table grid points is about 0\.9989\. According to our experiments, 66% of the inputs are mapped to the grid point at zero\. This implies that the grid of a B16 table is too coarse to represent most inputs faithfully\. It leads to incorrect normalization and potentially disrupts the successive operations such as projection and residual computations\. Increasing the LUT input precision to 24 bits improves the resolution\. It reduces the grid spacing to 0\.0039 over the same calibrated input range\. None of the sampled inputs then maps to the grid point at zero\. Instead, each input is represented by a nearby point on the finer grid\.

∙\\bulletWhy not use higher bit precision for all LUTs?One could use high bit precision for all LUTs and it will likely lead to good utility\. However, this will incur high overhead on both proving time and memory usage\. Increasing the input precision from 16 to 24 bits expands the number of table entries by a factor of 256, which increases the memory usage from 256 KiB to 64 MiB per table withint32entries\. The proving time is also proportional to the size of the table; see §[5\.2](https://arxiv.org/html/2609.36437#S5.SS2)\. We provide fine\-grained analysis for different non\-linear operations and our results show that near\-baseline utility can be achieved by introducing high\-precision LUTs only for the RMSNorm layer\.

### 5\.2End\-to\-End Proving Costs

#### 5\.2\.1Floating\-point vs\. Quantized proving at Qwen\-14B scale

In this section, we investigate the proving cost of ZK\-friendly quantization\. We begin with its cost advantages over floating\-point proving\. Specifically, we target the projection dimensions of the largest dense model in our evaluation settings, Qwen2\.5\-14B\. For floating\-point arithmetic in ZK, we use the implementation of[Ernstberger et al\. \(2025\)](https://arxiv.org/html/2609.36437#bib.bib41), which was later used in[Riasi et al\. \(2025\)](https://arxiv.org/html/2609.36437#bib.bib44)to prove ML computations\. The protocol supports basic operations such as floating\-point addition, multiplication, and division\. Therefore, our comparison focuses on linear projections in the LLM as a lower\-bound for the proving cost\. For quantized integer ZK proving, we measure the cost based on the zkML implementation of DeepProve\([Gailly et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib42)\)\.

Table 2:Full\-width projection proving times \(seconds\)\.Each Qwen\-14B transformer block has seven projections, including Q, K, V, O, and the FFN gate, up, and down projections\. Each projection is a matrix multiplication of the formy=x​Wy=xW, with an inner dimension of either 5,120 or 13,824\. Since directly proving all output columns with the floating\-point circuit is prohibitively expensive, we measure the proving time for eight output columns of each projection and extrapolate to the full projection width by scaling in proportion to the number of output columns\. We then sum the estimated costs of the seven projections to obtain the estimated floating\-point linear\-projection proving cost\.

On the other hand, DeepProve allows us to directly prove the full\-width projections with the W16A16 configuration\. Table[2](https://arxiv.org/html/2609.36437#S5.T2)compares these measured integer costs with the floating\-point estimates\. We measure each distinct full\-width projection shape once\. The integer total is the sum of these measured costs across the seven projections\. We put more detailed results with direct measurement of floating\-point proving and RAM usage in Table[13](https://arxiv.org/html/2609.36437#A9.T13), Appendix[I](https://arxiv.org/html/2609.36437#A9)\.

For one input token, the estimated floating\-point proving cost of a single block is 33\.87 hours, whereas the quantized integer proving time is 95\.46 seconds\. If we extend this to all 48 blocks, about 67 days will be taken for proving with floating\-point numbers\. This big gap occurs not only in proving time but also in peak RAM usage\. Floating\-point proving uses about 73–139 times as much peak RAM\. This is because floating\-point arithmetic requires additional constraints for exponent handling, normalization, and rounding\. Instead, ZK\-friendly quantization uses integer arithmetic with fixed scales, avoiding these per\-operation checks and handling rescaling through output requantization\. This experimental result justifies the importance of ZK\-friendly quantization\.

#### 5\.2\.2Proving cost does not scale smoothly with quantization

##### R4\. End\-to\-end proving time with different precisions for small models

In the previous section, we compared the proving costs of quantized integer and floating\-point computation\. Due to the high memory requirements of the current ZK protocol, full\-model proof generation is not yet feasible at the largest model scale\. We therefore evaluate end\-to\-end proving performance on smaller models \(i\.e\., GPT\-2 Small, Medium, and Large\) across different integer precisions\.

Figure 2:Proving time ratio relative to the W16A16 baseline for each model\.Figure[2](https://arxiv.org/html/2609.36437#S5.F2)shows the proving time for GPT\-2 Small, Medium, and Large, normalized by each model’s W16A16 mean\. The left panel examines the weight precision sensitivity and the right panel is for the activation sensitivity\. We run each configuration three times\.

We observe that proving cost does not strongly depend on the bit precision in this range\. This is because as long as the computations do not lead to overflows in𝔽p\\mathbb\{F\}\_\{p\}, different precisions result in the same number of arithmetic operations and lookups in the ZK protocol\. That is,\|ℛQ,p\|\|\\mathcal\{R\}\_\{Q,p\}\|is roughly the same for these quantization schemes\. Still, bit precision affects the integer representation and range of values, leading to different requantizations and range decompositions\. However, in the experiments, we observe that these precision\-dependent components are too small to materially change total prover time\. We put more detailed results in Appendix[E](https://arxiv.org/html/2609.36437#A5)\. Our result suggests that for small models that have already been measured in prior work\([Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34);[Gailly et al\., 2026](https://arxiv.org/html/2609.36437#bib.bib42)\), one could use higher precisions such as W16A16 for better utilities, with minimal overhead on the proving cost\.

However, the result above does not mean we can always increase the precision with no cost\. At higher precision, products and accumulated sums may no longer fit within the range supported by the underlying finite field\. For example, multiplying two integers wider than 16 bits can exceed the capacity of a small field, such as the 31\-bit BabyBear fieldp=15⋅227\+1p=15\\cdot 2^\{27\}\+1or the Mersenne primep=231−1p=2^\{31\}\-1, which are commonly used in many ZK protocols due to fast arithmetic over these fields\([Polygon, 2024](https://arxiv.org/html/2609.36437#bib.bib45);[RISC Zero, 2026](https://arxiv.org/html/2609.36437#bib.bib46)\)\. Once the overflow occurs, each integer value needs to be represented by multiple field elements, which will significantly increase\|ℛQ,p\|\|\\mathcal\{R\}\_\{Q,p\}\|, and thus the proving cost\. Therefore, the proving cost does not scale smoothly with the bit precision, and our work provides the first framework to select the minimum bit precision with good utility for different models for future zkML designs\.

##### R5\. End\-to\-end proving cost with different LUT sizes

We measure the end\-to\-end proving cost with different LUT sizes as well\. We vary both the GeLU and softmax exponential LUT bit precision between 8 and 24 bits in GPT\-2 Small\. We observe similar phenomena of non\-smooth scaling as well\. Reducing both LUTs from B16 to B8 shrinks each table by256×256\\times, yet the mean proving time remains essentially unchanged \(102\.65 vs\. 104\.75 s\)\. However, increasing the LUTs to B24 makes the proving time 2\.6×\\timesslower \(265\.54 s\)\. More configurations are reported in Table[12](https://arxiv.org/html/2609.36437#A9.T12)in Appendix[I](https://arxiv.org/html/2609.36437#A9)\.

The reason is different from the bit precision\. Recall that the proving cost of lookup arguments is proportional to both the size of the LUT and the number of lookup queries\. In this experiment, for the same model the number of direct lookup queries remains fixed at 0\.823 million; reducing table precision changes table size, but not the number of lookups, which is the same as the number of GeLU and exponential evaluations\. Reducing the LUT size from B16 to B8 decreases part of the proving time related to the LUT from 1\.325 to 0\.483 seconds, but this is not the dominating cost in end\-to\-end proving\. However, increasing the LUT size from B16 to B24 makes the table size the new bottleneck \(16 million entries\), thus significantly increasing the prover time\. These results justify the necessity of our fine\-grained utility analysis in §[5\.1](https://arxiv.org/html/2609.36437#S5.SS1), where we only need a B24 LUT for the RMSNorm while keeping other LUTs at B16\.

## 6Conclusion

We introduce a formalization of ZK\-friendly quantization and a scalable framework for studying its effects on LLM utility and ZK proving cost\. We evaluate nine language models across weight, activation, and nonlinear lookup precisions\. Our experiments identify RMSNorm inverse\-square\-root approximation as a recurring utility bottleneck in several models at 7B scale and above, and show that targeted increases in LUT precision can recover near\-baseline utility without uniformly increasing the precision of all nonlinear operators\. We further show that end\-to\-end proving cost does not necessarily scale linearly with quantization bit\-width or LUT size\. These findings suggest three design principles for ZK\-LLM quantization: preserve sufficient activation precision, allocate nonlinear precision on an operator\-specific basis, and reduce numerical precision only when it meaningfully simplifies the underlying proof relation\. Together, these principles motivate heterogeneous, operator\-aware quantization for balancing model fidelity with proving efficiency\.

#### Acknowledgments

This work is partially supported by the National Science Foundation \(NSF\) under Grant No\. 2453149, 2613388 and Amazon research award\. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author\(s\) and do not necessarily reflect the views of these institutes\.

## References

- Chenet al\.\(2024\)B\. Chen, S\. Waiwitlikhit, I\. Stoica, and D\. KangZkml: an optimizing system for ml inference in zero\-knowledge proofs\.InProceedings of the Nineteenth European Conference on Computer Systems,pp\. 560–574\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[Appendix F](https://arxiv.org/html/2609.36437#A6.p6.1)\.
- Dettmerset al\.\(2022\)T\. Dettmers, M\. Lewis, Y\. Belkada, and L\. ZettlemoyerGpt3\. int8 \(\): 8\-bit matrix multiplication for transformers at scale\.Advances in neural information processing systems35,pp\. 30318–30332\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Dettmerset al\.\(2023\)T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. ZettlemoyerQlora: efficient finetuning of quantized llms\.Advances in neural information processing systems36,pp\. 10088–10115\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Ernstbergeret al\.\(2025\)J\. Ernstberger, C\. Zhang, L\. Ciprian, P\. Jovanovic, and S\. SteinhorstZero\-knowledge location privacy via accurate floating\-point snarks\.In2025 IEEE Symposium on Security and Privacy \(SP\),pp\. 3440–3459\.Cited by:[§5\.2\.1](https://arxiv.org/html/2609.36437#S5.SS2.SSS1.p1.1)\.
- European Parliament and Council of the European Union \(2016\)European Parliament and Council of the European UnionRegulation \(EU\) 2016/679 of the European Parliament and of the Council of 27 April 2016\.Note:[https://eur\-lex\.europa\.eu/eli/reg/2016/679/oj](https://eur-lex.europa.eu/eli/reg/2016/679/oj)Cited by:[Appendix H](https://arxiv.org/html/2609.36437#A8.p1.1)\.
- European Parliament and Council of the European Union \(2024\)European Parliament and Council of the European UnionRegulation \(EU\) 2024/1689 of the European Parliament and of the Council of 13 June 2024\.Note:[https://eur\-lex\.europa\.eu/eli/reg/2024/1689/oj](https://eur-lex.europa.eu/eli/reg/2024/1689/oj)Cited by:[Appendix H](https://arxiv.org/html/2609.36437#A8.p1.1)\.
- Fenget al\.\(2021\)B\. Feng, L\. Qin, Z\. Zhang, Y\. Ding, and S\. ChuZen: an optimizing compiler for verifiable, zero\-knowledge neural network inferences\.Cryptology ePrint Archive\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1)\.
- Frantar and Alistarh \(2023\)E\. Frantar and D\. AlistarhSparsegpt: massive language models can be accurately pruned in one\-shot\.InInternational conference on machine learning,pp\. 10323–10337\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Frantaret al\.\(2022\)E\. Frantar, S\. Ashkboos, T\. Hoefler, and D\. AlistarhGptq: accurate post\-training quantization for generative pre\-trained transformers\.arXiv preprint arXiv:2210\.17323\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Gaillyet al\.\(2026\)N\. Gailly, I\. Hishon\-Rezaizadeh, T\. Liu, N\. Mainardi, D\. Papadopoulos, C\. Papamanthou, C\. Pappas, S\. Srinivasan, Z\. Youell, and Y\. ZhangDeepProve: verifiable end\-to\-end large language model inference\.Cryptology ePrint Archive\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p1.1),[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px2.p2.1),[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px3.p1.1),[§5\.2\.1](https://arxiv.org/html/2609.36437#S5.SS2.SSS1.p1.1),[§5\.2\.2](https://arxiv.org/html/2609.36437#S5.SS2.SSS2.Px1.p3.1)\.
- Gaoet al\.\(2021\)L\. Gao, J\. Tow, S\. Biderman, S\. Black, A\. DiPofi, C\. Foster, L\. Golding, J\. Hsu, K\. McDonell, N\. Muennighoff,et al\.A framework for few\-shot language model evaluation\.Zenodo\.Cited by:[Appendix F](https://arxiv.org/html/2609.36437#A6.p6.1),[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px4.p1.1)\.
- Gentry \(2009\)C\. GentryFully homomorphic encryption using ideal lattices\.InProceedings of the forty\-first annual ACM symposium on Theory of computing,pp\. 169–178\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p2.1)\.
- Gonget al\.\(2024\)R\. Gong, Y\. Yong, S\. Gu, Y\. Huang, C\. Lv, Y\. Zhang, X\. Liu, and D\. TaoLLMC: benchmarking large language model quantization with a versatile compression toolkit\.External Links:2405\.06001,[Link](https://arxiv.org/abs/2405.06001)Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Hendryckset al\.\(2020\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[Appendix F](https://arxiv.org/html/2609.36437#A6.p6.1)\.
- Houet al\.\(2026\)X\. Hou, J\. Liu, J\. Li, Y\. Li, J\. Zhang, W\. Lu, C\. Hong, and K\. RenCiphergpt: secure two\-party gpt inference\.IEEE Transactions on Dependable and Secure Computing\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p2.1)\.
- Huanget al\.\(2024\)W\. Huang, X\. Ma, H\. Qin, X\. Zheng, C\. Lv, H\. Chen, J\. Luo, X\. Qi, X\. Liu, and M\. MagnoHow good are low\-bit quantized llama3 models? an empirical study\.Vol\.22,Apr\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p2.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Jacobet al\.\(2018\)B\. Jacob, S\. Kligys, B\. Chen, M\. Zhu, M\. Tang, A\. Howard, H\. Adam, and D\. KalenichenkoQuantization and training of neural networks for efficient integer\-arithmetic\-only inference\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 2704–2713\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1)\.
- Kurticet al\.\(2025\)E\. Kurtic, A\. N\. Marques, S\. Pandit, M\. Kurtz, and D\. Alistarh“Give me bf16 or give me death”? accuracy\-performance trade\-offs in llm quantization\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 26872–26886\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Leeet al\.\(2024\)S\. Lee, H\. Ko, J\. Kim, and H\. OhVcnn: verifiable convolutional neural network based on zk\-snarks\.IEEE Transactions on Dependable and Secure Computing21\(4\),pp\. 4254–4270\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1)\.
- Liaoet al\.\(2025\)G\. Liao, T\. Wang, S\. Zhang, J\. Zhang, S\. Long, and D\. TaoVeriLoRA: fine\-tuning large language models with verifiable security via zero\-knowledge proofs\.arXiv preprint arXiv:2508\.21393\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p1.1)\.
- Linet al\.\(2024\)J\. Lin, J\. Tang, H\. Tang, S\. Yang, W\. Chen, W\. Wang, G\. Xiao, X\. Dang, C\. Gan, and S\. HanAwq: activation\-aware weight quantization for on\-device llm compression and acceleration\.Proceedings of machine learning and systems6,pp\. 87–100\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2021\)T\. Liu, X\. Xie, and Y\. ZhangZkcnn: zero knowledge proofs for convolutional neural network predictions and accuracy\.InProceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security,pp\. 2968–2985\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1)\.
- Maet al\.\(2023\)X\. Ma, G\. Fang, and X\. WangLlm\-pruner: on the structural pruning of large language models\.Advances in neural information processing systems36,pp\. 21702–21720\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Merityet al\.\(2016\)S\. Merity, C\. Xiong, J\. Bradbury, and R\. SocherPointer sentinel mixture models\.arXiv preprint arXiv:1609\.07843\.Cited by:[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px4.p1.1)\.
- Polygon \(2024\)PolygonPlonky3: a toolkit for polynomial IOPs \(PIOPs\)\.Note:[https://github\.com/Plonky3/Plonky3](https://github.com/Plonky3/Plonky3)Accessed: September 25, 2026Cited by:[Appendix H](https://arxiv.org/html/2609.36437#A8.SS0.SSS0.Px2.p1.1),[§5\.2\.2](https://arxiv.org/html/2609.36437#S5.SS2.SSS2.Px1.p4.1)\.
- Quet al\.\(2025\)W\. Qu, Y\. Sun, X\. Liu, T\. Lu, Y\. Guo, K\. Chen, and J\. Zhang\{\\\{zkgpt\}\\\}: An efficient non\-interactive zero\-knowledge proof framework for\{\\\{llm\}\\\}inference\.In34th USENIX Security Symposium \(USENIX Security 25\),pp\. 2045–2063\.Cited by:[Appendix C](https://arxiv.org/html/2609.36437#A3.SS0.SSS0.Px2.p1.1),[Appendix H](https://arxiv.org/html/2609.36437#A8.p3.1),[§1](https://arxiv.org/html/2609.36437#S1.p1.1),[§1](https://arxiv.org/html/2609.36437#S1.p2.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px2.p1.1),[§5\.2\.2](https://arxiv.org/html/2609.36437#S5.SS2.SSS2.Px1.p3.1)\.
- Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,et al\.Language models are unsupervised multitask learners\.OpenAI blog1\(8\),pp\. 9\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1)\.
- Raffelet al\.\(2020\)C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. LiuExploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of machine learning research21\(140\),pp\. 1–67\.Cited by:[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px4.p1.1)\.
- Riasiet al\.\(2025\)A\. Riasi, H\. Wang, R\. Behnia, V\. Vo, and T\. HoangZero\-knowledge ai inference with high precision\.InProceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security,pp\. 1053–1067\.Cited by:[§5\.2\.1](https://arxiv.org/html/2609.36437#S5.SS2.SSS1.p1.1)\.
- RISC Zero \(2026\)RISC ZeroRISC Zero: verifiable computation\.Note:[https://risczero\.com/](https://risczero.com/)Accessed: September 25, 2026Cited by:[§5\.2\.2](https://arxiv.org/html/2609.36437#S5.SS2.SSS2.Px1.p4.1)\.
- Settyet al\.\(2024\)S\. Setty, J\. Thaler, and R\. WahbyUnlocking the lookup singularity with lasso\.InAnnual International Conference on the Theory and Applications of Cryptographic Techniques,pp\. 180–209\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1)\.
- Shaoet al\.\(2023\)W\. Shao, M\. Chen, Z\. Zhang, P\. Xu, L\. Zhao, Z\. Li, K\. Zhang, P\. Gao, Y\. Qiao, and P\. LuoOmniquant: omnidirectionally calibrated quantization for large language models\.arXiv preprint arXiv:2308\.13137\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Shiet al\.\(2023\)W\. Shi, A\. Ajith, M\. Xia, Y\. Huang, D\. Liu, T\. Blevins, D\. Chen, and L\. ZettlemoyerDetecting pretraining data from large language models\.arXiv preprint arXiv:2310\.16789\.Cited by:[Appendix H](https://arxiv.org/html/2609.36437#A8.p2.1)\.
- Succinct Labs \(2026\)Succinct LabsSP1: a performant, open\-source zkvm\.Note:[https://github\.com/succinctlabs/sp1](https://github.com/succinctlabs/sp1)Accessed: 2026\-05\-01Cited by:[Appendix H](https://arxiv.org/html/2609.36437#A8.SS0.SSS0.Px2.p1.1)\.
- Sunet al\.\(2024\)H\. Sun, J\. Li, and H\. ZhangZkllm: zero knowledge proofs for large language models\.InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security,pp\. 4405–4419\.Cited by:[Appendix H](https://arxiv.org/html/2609.36437#A8.p3.1),[§1](https://arxiv.org/html/2609.36437#S1.p1.1),[§1](https://arxiv.org/html/2609.36437#S1.p2.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.36437#S4.SS0.SSS0.Px2.p1.1)\.
- Sunet al\.\(2023\)M\. Sun, Z\. Liu, A\. Bair, and J\. Z\. KolterA simple and effective pruning approach for large language models\.arXiv preprint arXiv:2306\.11695\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2026\)L\. Wang, H\. Lou, C\. Li, Y\. Yu, and Y\. HuZkAgent: verifiable agent execution via one\-shot complete llm inference proof\.Cryptology ePrint Archive\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p1.1)\.
- Xiaoet al\.\(2023\)G\. Xiao, J\. Lin, M\. Seznec, H\. Wu, J\. Demouth, and S\. HanSmoothquant: accurate and efficient post\-training quantization for large language models\.InInternational conference on machine learning,pp\. 38087–38099\.Cited by:[Appendix C](https://arxiv.org/html/2609.36437#A3.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Xieet al\.\(2025\)T\. Xie, T\. Lu, Z\. Fang, S\. Wang, Z\. Zhang, Y\. Jia, D\. Song, and J\. ZhangZkpytorch: a hierarchical optimized compiler for zero\-knowledge machine learning\.Cryptology ePrint Archive\.Cited by:[Appendix C](https://arxiv.org/html/2609.36437#A3.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.36437#S1.p1.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Yadavet al\.\(2024\)C\. Yadav, A\. R\. Chowdhury, D\. Boneh, and K\. ChaudhuriFairProof: confidential and certifiable fairness for neural networks\.InProceedings of the 41st International Conference on Machine Learning,ICML ’24\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1)\.
- Yadavet al\.\(2026\)C\. Yadav, E\. M\. Laufer, D\. Boneh, and K\. ChaudhuriExpproof: operationalizing explanations for confidential models with zkps\.InProceedings of the 41st International Conference on Machine Learning,ICML ’26\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2024\)G\. Yang, C\. He, J\. Guo, J\. Wu, Y\. Ding, A\. Liu, H\. Qin, P\. Ji, and X\. LiuLlmcbench: benchmarking large language model compression for efficient deployment\.Advances in Neural Information Processing Systems37,pp\. 87532–87544\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p2.1),[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Yao \(1982\)A\. C\. YaoProtocols for secure computations\.In23rd annual symposium on foundations of computer science \(sfcs 1982\),pp\. 160–164\.Cited by:[§1](https://arxiv.org/html/2609.36437#S1.p2.1)\.
- Yaoet al\.\(2023\)Z\. Yao, X\. Wu, C\. Li, S\. Youn, and Y\. HeZeroQuant\-v2: exploring post\-training quantization in llms from comprehensive study to low rank compensation\.External Links:2303\.08302,[Link](https://arxiv.org/abs/2303.08302)Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025\)T\. Zhang, S\. Dong, O\. D\. Kose, Y\. Shen, and Y\. ZhangFairzk: a scalable system to prove machine learning fairness in zero\-knowledge\.In2025 IEEE Symposium on Security and Privacy \(SP\),pp\. 3460–3478\.Cited by:[§2](https://arxiv.org/html/2609.36437#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix ASimulated Integer Execution

Our framework uses quantize\-dequantize \(QDQ\) simulation to evaluate quantization configurations with existing PyTorch operations and GPU acceleration\. This avoids implementing dedicated integer kernels for every bit\-width, such as INT12, and makes large\-model evaluation practical\. It raises one question: how closely does their output reproduce the corresponding integer computation? We answer this question first algebraically and then experimentally\. We separate the error introduced by quantization from the additional error introduced by its QDQ simulation\.

##### Algebraic comparison

QDQ first rounds and clips each tensor to its bounded\-integer representation, then dequantizes it before the PyTorch operation\. Dequantization restores the scale of the values, but not the information lost through rounding and clipping\. LetX∈ℝM×KX\\in\\mathbb\{R\}^\{M\\times K\}andW∈ℝK×NW\\in\\mathbb\{R\}^\{K\\times N\}have fixed per\-tensor scalessxs\_\{x\}andsws\_\{w\}\. For either tensorZZwith scaleszs\_\{z\}, we define signedqq\-bit quantization as

QZ=clip\[−2q−1,2q−1−1\]⁡\(round⁡\(Zsz\)\)\.Q\_\{Z\}=\\operatorname\{clip\}\_\{\[\-2^\{q\-1\},\\,2^\{q\-1\}\-1\]\}\\\!\\left\(\\operatorname\{round\}\\\!\\left\(\\frac\{Z\}\{s\_\{z\}\}\\right\)\\right\)\.\(1\)The QDQ and integer paths use the same quantized tensors but apply their scales at different stages:

Yint\\displaystyle Y\_\{\\mathrm\{int\}\}=\(QX​QW\)​sx​sw\\displaystyle=\(Q\_\{X\}Q\_\{W\}\)s\_\{x\}s\_\{w\}\(integer multiplication, then rescale\),\\displaystyle\\text\{\(integer multiplication, then rescale\)\},\(2\)YQDQ\\displaystyle Y\_\{\\mathrm\{QDQ\}\}=\(sx​QX\)​\(sw​QW\)\\displaystyle=\(s\_\{x\}Q\_\{X\}\)\(s\_\{w\}Q\_\{W\}\)\(dequantize, then multiply\)\.\\displaystyle\\text\{\(dequantize, then multiply\)\}\.\(3\)
The above expressions are identical under exact arithmetic\. On a computer, however, dequantization and matrix multiplication round the number, so algebraic equivalence does not promise numerical agreement\.

##### Experimental validation

We examine the numerical errors introduced by quantization and QDQ in actual computation\. Using nine projections in total from GPT\-2 Small, we compute three outputs: the unquantized productYFP=X​WY\_\{\\mathrm\{FP\}\}=XW, the integer\-accumulation referenceYintY\_\{\\mathrm\{int\}\}, and the QDQ outputYQDQY\_\{\\mathrm\{QDQ\}\}\. From the configurations used in our experiments, we select W8A8, W12A12, and W16A16, testing both the standard scale used in the main experiments and the power\-of\-two \(PoT\) scale\. PoT provides a useful comparison because binary floating\-point arithmetic can apply power\-of\-two scaling exactly when the scaled values are representable\. We execute the QDQ projections in both FP32 and BF16\.

We define two errors to distinguish the effect of quantization from the additional error introduced by its simulation\. The quantization error,EquantE\_\{\\mathrm\{quant\}\}, measures how much quantization changes the original floating\-point computation\. Both the integer and QDQ paths use quantized operands, so this quantization effect is present in both paths\. We isolate it by comparingYintY\_\{\\mathrm\{int\}\}withYFPY\_\{\\mathrm\{FP\}\}\. The simulation error,EsimE\_\{\\mathrm\{sim\}\}, instead measures how closely QDQ reproduces the integer reference\. It captures the additional numerical error introduced by floating\-point dequantization, dtype conversion, and matrix multiplication\. Thus, even whenEquantE\_\{\\mathrm\{quant\}\}is large, a sufficiently smallEsimE\_\{\\mathrm\{sim\}\}implies that QDQ faithfully reproduces the quantized computation\. To summarize the errors across all nine projections, we flatten and concatenate their outputs into one vector for each computation path\.

Equant=‖Yint−YFP‖2‖YFP‖2,Esim=‖YQDQ−Yint‖2‖Yint‖2\.E\_\{\\mathrm\{quant\}\}=\\frac\{\\left\\\|Y\_\{\\mathrm\{int\}\}\-Y\_\{\\mathrm\{FP\}\}\\right\\\|\_\{2\}\}\{\\left\\\|Y\_\{\\mathrm\{FP\}\}\\right\\\|\_\{2\}\},\\qquad E\_\{\\mathrm\{sim\}\}=\\frac\{\\left\\\|Y\_\{\\mathrm\{QDQ\}\}\-Y\_\{\\mathrm\{int\}\}\\right\\\|\_\{2\}\}\{\\left\\\|Y\_\{\\mathrm\{int\}\}\\right\\\|\_\{2\}\}\.\(4\)
Table 3:Numerical validation of QDQ on actual GPT\-2 Small projections\. All values are relative errors, not percentages\.
##### Results\.

Table[3](https://arxiv.org/html/2609.36437#A1.T3)presents the numerical validation results\. As expected, increasing quantization precision reduces the error from the unquantized floating\-point result\. Our main interest is the simulation error\. In particular, the PoT configurations produce identical outputs at 8 and 12 bits in this experiment\. This is because power\-of\-two scaling can be applied exactly to representable binary floating\-point values, removing a source of rounding error\. With BF16 execution, the simulation error is generally larger than with FP32, but still remains small\. BF16 inherently has fewer significant bits than FP32, so storing the dequantized operands and outputs in BF16 introduces additional rounding\. For some cases that require more numerical equivalence, FP32 execution with PoT scaling can be a useful option\. This choice, however, comes with the higher memory requirements of FP32 execution\.

## Appendix BLookup Tables

In our framework, the key mechanism for replacing non\-linear operations is a LUT, a fixed table of precomputed input\-output pairs for a target function\. For a non\-linear operatorgg, we first choose an input range by calibration or an operator\-specific predefined bound\. Once we fix the input range, we discretize it into a uniform grid with2B2^\{B\}entries, whereBBdenotes the LUT input precision\. We then evaluateggat each grid point and quantize the outputs to 24\-bit integers stored inint32\. During evaluation, each runtime input is rounded to the nearest grid and replaced by the corresponding output in the table\. The LUT for softmax exponential, for example, has an input range of\[−20,0\]\[\-20,0\]\. If the input is out of the range, then we clip it to the nearest endpoint\. Our simulation collects aggregate clipping rates and lookup counts per model\.

The PyTorch implementation follows the same rule in general, while the exact implementation for each non\-linear operator is different\. This is because the place where the operator appears and the way it is called are different\. For example, we replace the forward computation of GeLU and SiLU modules with LUT evaluation while preserving the existing modules\. For normalization, we modify the forward computation to replace only the inverse\-square\-root function with a LUT, preserving the mean and variance computation in LayerNorm and the mean\-square computation in RMSNorm\. For attention softmax, we temporarily patch the softmax evaluation path so that the exponential and reciprocal are computed through LUTs\. In particular, for Qwen3, LUT coverage also includes Q/K RMSNorm and the exponential and reciprocal operations used in MoE routing\. In terms of memory, anint32B16 table requires 256 KiB, whereas a B24 table requires 64 MiB, both much smaller than the model weights in our experiments\. Additionally, lookup is implemented as batched tensor indexing rather than element\-wise Python lookup, avoiding per\-element Python overhead\.

## Appendix CDetailed Framework Implementation

Our framework implements ZK\-friendly quantization as a config\-driven PyTorch and Hugging Face pipeline\. Given a model and a quantization configuration, it patches the transformer forward pass at the operation level: linear tensor operations use quantized weights and activations with quantize\-dequantize modules, while selected nonlinear functions are replaced by LUT modules \(See Appendix[B](https://arxiv.org/html/2609.36437#A2)for a more detailed explanation of non\-linear operations\)\. The configuration specifies the design choices that affect both ZK compatibility and model utility, including weight and activation bit\-widths, scale type, quantization symmetry, activation calibration, and LUT bit\-width\.

##### Linear operations\.

All matrix multiplications in the transformer produce unbounded floating\-point intermediate results that must be converted to bounded integers for ZK circuit compatibility\. This category includes everynn\.Linear\(orConv1Din GPT\-2\) layer in the architecture, such as the QKV projections, the output projection, the MLP layers, and the final language\-modeling head\. The language\-modeling head uses quantized weights but does not apply an additional A\-bit requantization to its output logits\. Attention score computationQ​K⊤QK^\{\\top\}and context aggregationsoftmax⁡\(S\)​V\\mathrm\{softmax\}\(S\)Vremain native tensor operations; the evaluated policy does not insert a separate QDQ operation after each of these matrix products\. The evaluation policy also covers embedding weights and residual outputs, whose precision is fixed at 24 bits in the main utility experiments, as well as the expert and router projections in Qwen3\. Given this operation\-level classification, the framework implements each configuration by patching the corresponding PyTorch modules in place\.

##### Scales\.

For each quantized tensor, the configuration determines how its scale is chosen\. We use one scale per weight tensor in all experiments\. A standard scale uses the floating\-point range of the tensor directly\. The main utility experiments use the scales=2​a/\(2b−1\)s=2a/\(2^\{b\}\-1\), whereas the standard\-scale ablation usess=a/\(2b−1−1\)s=a/\(2^\{b\-1\}\-1\), withaameaning the maximum absolute value andbbthe target bit\-width\. A power\-of\-two scale is computed ass=2⌈log2⁡\(a/2b−1\)⌉s=2^\{\\lceil\\log\_\{2\}\(a/2^\{b\-1\}\)\\rceil\}\. The reason a power\-of\-two scale is ZK\-friendly is that it can be compiled as a fixed shift\-like integer operation or a multiplication by a public constant\. Thec1/c2c\_\{1\}/c\_\{2\}scale uses an integer\-ratio rescaling rule, wherec1c\_\{1\}andc2c\_\{2\}are fixed integers as in zkGPT\([Qu et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib34)\)\. It can avoid floating\-point arithmetic by using integer multiplication and division by a constant\.

##### Activation calibration\.

Activation calibration is needed because activations are runtime values\. We calibrate activation ranges on eight fixed WikiText\-2 training windows with context length 512 and stride 256, before weight quantization\. For each static configuration, scales are determined from these ranges and reused across evaluations, including C4\. We use one static scale per designated activation tensor, rather than one scale for an entire transformer layer\. For a linear layery=x​W\+by=xW\+b, the weightsWWand biasbbare fixed and can be quantized offline, but the output activationyydepends on the input prompt\. The outputyyalso becomes the input of the next layer\. Static\-max calibration fixes activation scales before inference\([Xie et al\., 2025](https://arxiv.org/html/2609.36437#bib.bib16);[Xiao et al\., 2023](https://arxiv.org/html/2609.36437#bib.bib7)\)\. We run calibration prompts once, record a maximum activation range at each designated site, and reuse that fixed range during evaluation\. Dynamic activation quantization instead computes the scale from the current input, which incurs a lot of overhead\.

##### Compute environment\.

The main model utility experiments were run with two Intel Xeon Gold 6526Y CPUs, 251 GiB RAM, and two NVIDIA H100 NVL GPUs\. The PyTorch environment used version 2\.10\.0 with CUDA 12\.6\.

## Appendix DAdditional Utility Results

Figure 3:Perplexity ratio under linear quantization to BF16 baseline\.Table 4:Full\-LUT PPL ratios relative to the W16A16 linear\-only baseline\.Table 5:PPL ratios relative to the W16A16 linear\-only baseline with targeted precision adjustment\.Figure[3](https://arxiv.org/html/2609.36437#A4.F3)extends the linear\-quantization comparison to all nine models on WikiText\-2 and C4\. Activation quantization generally causes greater utility loss than weight quantization, although the sensitivity varies across models\. Table[4](https://arxiv.org/html/2609.36437#A4.T4)reports the additional effect of LUT approximation relative to the W16A16 linear\-only baseline\. B8 LUTs cause severe degradation across all models, while increasing precision to B16 preserves utility for some models but remains insufficient for others\.

Table[5](https://arxiv.org/html/2609.36437#A4.T5)examines four affected models on WikiText\-2\. RMS inverse\-square\-root LUTs alone reproduce much of the full\-B16 degradation in Llama\-3\.1\-8B and Qwen2\.5\-7B/14B\. For Qwen3\-30B\-A3B, the degradation instead emerges when these LUTs are combined with the other approximations\. Replacing only the RMS LUTs with exact functions brings all four models close to their linear\-only baselines\. Increasing their input precision to B24 achieves a similar result\.

## Appendix EAdditional Proving Cost Results

##### Weight and activation precision

Table[6](https://arxiv.org/html/2609.36437#A5.T6)reports the absolute proving times underlying Figure[2](https://arxiv.org/html/2609.36437#S5.F2)\. We evaluate GPT\-2 Small, Medium, and Large at a context length of 16 tokens, with batch size one and 32 CPU threads\. We report the mean, minimum, and maximum proving times\. The reported time excludes graph preparation, inference, and verification\.

Table 6:Proving time under different W/A precisions \(seconds\)\.
##### Context length

We extend the W/A precision experiment by measuring GPT\-2 Small and Medium at context lengths of 4, 8, 16, and 32 tokens, using W16A16 and W12A16\. We use the same quantization policies and DeepProve backend as in the main experiment \(§[5\.2](https://arxiv.org/html/2609.36437#S5.SS2)\), with batch size one and 32 CPU threads\. For each setting, we prove the full prefill computation and report one formal measurement\. The reported time excludes graph preparation, inference, and verification\.

Table 7:Proving time across context lengths \(seconds\)\.Table[7](https://arxiv.org/html/2609.36437#A5.T7)shows that proving time increases consistently with context length in both models\. Increasing the context from 4 to 32 tokens raises proving time by about 1\.8 times across the four settings\. The increase with context follows a change in the amount of computation: longer inputs produce more activation elements and larger attention matrices whose relations must be proved\. The recorded number of intermediate elements increases by approximately 8\.4 times from context 4 to 32, whereas reducing weight precision leaves both the graph node count and intermediate tensor sizes unchanged at each context\. In conclusion, we observe that context length has a much larger effect on proving time than the change of weight precision\.

## Appendix FBehavioral Fidelity Experimentation

Perplexity is a useful scalar measure of language\-modeling utility, but it is not sufficient for ZK\-LLM governance, where the proved computation may depend on fine\-grained output\-distribution behavior\. Auditing model behavior, verifying token\-level claims, and proving membership\-inference statistics require quantization to preserve token rankings, low\-probability regions, and downstream decisions, not only average likelihood\. We therefore evaluate W16A16 and W12A12 with the model\-specific LUT configurations selected in our utility study \(§[5\.1](https://arxiv.org/html/2609.36437#S5.SS1)\), alongside a W16A16 linear\-only baseline\.

∙\\bulletLogit\-level fidelity\.

Table[8](https://arxiv.org/html/2609.36437#A6.T8)compares BF16 and quantized next\-token distributions on WikiText\-2 and C4 for the largest five models in our suite\. Each configuration is evaluated once on the same 8,192 scored tokens per model and dataset\. For each token position, we compute KL divergence from the BF16 distribution to the quantized distribution and the fraction of shared tokens in their Top\-5 predictions\. We also measure bottom\-20% overlap: each model selects the 20% of evaluated positions with the lowest probabilities assigned to the actual target tokens, and we report the intersection divided by the selected set size\.

The full\-pipeline configurations use B16 LUTs throughout Qwen2\.5\-3B\. For the other four models, block and final RMSNorm inverse\-square\-root LUTs use B24, while the remaining LUTs use B16\. Qwen3’s Q/K RMSNorm and router LUTs also remain at B16\. The same LUT settings and calibrated ranges are used for W16A16 and W12A12\.

W16A16 with these LUT configurations consistently preserves token\-level behavior\. It maintains low KL divergence and high overlap in both the top\-ranked and low\-probability token regions across datasets and models\. This matters for ZK\-LLM governance applications because the proof may certify a computation over token probabilities, not only a final generated token\. By contrast, W12A12 with the same LUT settings has lower fidelity across all five models\. This confirms that perplexity alone is not enough to certify behavioral fidelity: a configuration can appear acceptable under scalar likelihood while still changing distributional structure\.

Table 8:Logit\-level fidelity relative to the BF16 reference\.Linear: nonlinear functions remain exact\. Full: the model\-specific B16/B24 LUT configuration described in the text, held fixed across W/A settings\.

∙\\bulletDownstream task accuracy\. We evaluate two multiple\-choice benchmarks: ARC\-easy\([Clark et al\., 2018](https://arxiv.org/html/2609.36437#bib.bib1)\)for basic reasoning and MMLU\([Hendrycks et al\., 2020](https://arxiv.org/html/2609.36437#bib.bib10)\)for broad knowledge, using the lm\-evaluation\-harness\([Gao et al\., 2021](https://arxiv.org/html/2609.36437#bib.bib12)\)MCQ scoring protocol\. These tasks test whether the quantized model preserves end\-task decisions, rather than only pretraining likelihood or token\-distribution similarity\. We evaluate all 2,376 ARC\-easy test questions with 0\-shot prompts and all 14,042 MMLU test questions across 57 subjects with 5\-shot prompts\.

Figure[4](https://arxiv.org/html/2609.36437#A6.F4)describes the accuracy results under different quantization configurations\. Both Full16 and Full12 settings are from our utility study: B24 for block and final RMSNorm lookups with B16 for the rest\. The result suggests that the strongest ZK\-friendly configuration \(Full 16\) preserves downstream behavior as well as perplexity and logit\-level fidelity\. While Full 12 is more model\-dependent\. In particular, it shows larger loss in Llama\-3\.1\-8B\. Hence, the results present that our framework can evaluate quantization configurations not only through perplexity and logit\-level fidelity, but also through downstream tasks that can reflect practical use of LLMs\.

Figure 4:Downstream task accuracy on ARC\-easy and MMLU under different configurations\.
## Appendix GQuantization Design\-Choice Ablations

We examine how quantization design choices affect utility at W12A12 and W16A16 across nine models on WikiText\-2 and C4\. The baseline uses static per\-tensor signed\-symmetric quantization with non\-power\-of\-two scales\. Starting from this baseline, we separately substitute power\-of\-two scales, conventional symmetric scales, bounded integer\-ratio scales, asymmetric quantization, or dynamic per\-token activation quantization\. The dynamic setting uses the baseline weight quantization\. Across all settings, embedding and residual quantizers remain fixed at 24 bits, and nonlinear functions are evaluated without LUT approximation\. Each configuration is evaluated once on 32 windows per dataset, with a context length of 512 and 8,192 scored tokens in total\.

Table 9:Design\-choice ablations: perplexity with exact nonlinear functions\.Table[9](https://arxiv.org/html/2609.36437#A7.T9)shows that design choices have little effect at W16A16: all evaluated settings remain within approximately 1\.2% of the BF16 reference PPL\. At W12A12, however, activation scale selection becomes substantially more important\. Dynamic per\-token quantization improves PPL over the static baseline in 17 of the 18 model–dataset pairs\. For example, Llama\-3\.1\-8B’s C4 PPL decreases from 17\.30 to 9\.13, close to the BF16 reference of 9\.12\. Changing only the scale representation does not recover this loss, while asymmetric quantization provides model\-dependent improvements\. These results suggest that the degradation at W12A12 is not solely determined by bit width: adapting activation scales can use the available precision more effectively\. However, the dynamic setting changes both scale granularity and runtime adaptation, so this experiment does not isolate their individual effects or establish their proving\-cost implications\.

## Appendix HUse Case: Proving Training\-Data \(Non\-\)Membership in Zero\-Knowledge

In this section, we demonstrate a ZK\-LLM governance use case with our ZK\-friendly quantization for mandated training\-data disclosure \(EU AI Act[European Parliament and Council of the European Union, 2024](https://arxiv.org/html/2609.36437#bib.bib4), GDPR Article 15[European Parliament and Council of the European Union, 2016](https://arxiv.org/html/2609.36437#bib.bib3)\): proving whether a given text was likely included in a model’s training data without revealing the model itself\.

In this setting, a model provider can use a ZKP to certify the result of a membership\-inference computation, based on the Min\-K% probability score[Shi et al\. \(2023\)](https://arxiv.org/html/2609.36437#bib.bib35), while keeping the model parameters and intermediate inference values private\.

Specifically, given a textXX, Min\-K% Prob computeslog⁡P⁡\(xi∣x<i\)\\log P\(x\_\{i\}\\mid x\_\{<i\}\)for every token ofXX, takes the K% tokens with the smallest log\-probability, and computes their mean\. A score above a calibrated thresholdε\\varepsilonsuggests thatXXwas seen during training\. To prove this score in ZK, the prover needs to prove the LLM inference onXXto obtain the probability of each token, followed by the algorithm above\. As there are already multiple prior works on LLM inference[Sun et al\. \(2024\)](https://arxiv.org/html/2609.36437#bib.bib32);[Qu et al\. \(2025\)](https://arxiv.org/html/2609.36437#bib.bib34), we focus on the Min\-K% Prob score part\.

##### Protocol\.

Our protocol is described in Figure[5](https://arxiv.org/html/2609.36437#A8.F5)\. We present the protocol in terms of lookup operations and modular arithmetic, so that the protocol is agnostic to the underlying ZK protocol\. The protocol takes the quantized per\-token probability vector computed by the ZK inference of the LLM as a secret witness\. It computes the fixed\-point negative log\-probabilities using a precomputed lookup table, validates the sorted vectorℓ⋆\\ell^\{\\star\}, computes the sum of thekkentries of the least likely tokens, and compares the sum with the threshold\. Only the membership decision is revealed, while the score stays private\. All computations can be realized in ZK protocols efficiently\.

Public input:sequence lengthNN, Min\-K parameterkkwith0<k<N0<k<N, thresholdTT, and preprocessed lookup tablesTlogT\_\{\\log\}andTrangeT\_\{\\mathrm\{range\}\}\. Forq∈\[1,216\)q\\in\[1,2^\{16\}\), the log table mapsqqto the fixed\-point negative log\-probability⌊−log\(q/216\)⋅215⌉\\lfloor\-\\log\(q/2^\{16\}\)\\cdot 2^\{15\}\\rceil, which is below2192^\{19\}\. The range table isTrange=\[0,216\)T\_\{\\mathrm\{range\}\}=\[0,2^\{16\}\)denoting non\-negative numbers\.Secret Witness:secret per\-token quantized probability vectorq=\(q0,…,qN−1\)∈\[1,216\)Nq=\(q\_\{0\},\\ldots,q\_\{N\-1\}\)\\in\[1,2^\{16\}\)^\{N\}from the quantized model’s forward pass on textXX\.Auxiliary input:negative log\-probabilitiesℓ∈𝔽N\\ell\\in\\mathbb\{F\}^\{N\}, sorted copyℓ⋆∈𝔽N\\ell^\{\\star\}\\in\\mathbb\{F\}^\{N\}, table multiplicities forTlogT\_\{\\log\}andTrangeT\_\{\\mathrm\{range\}\}used in lookup arguments\. The accumulator’s final valuesfinals\_\{\\mathrm\{final\}\}stays private inside the committed trace\.Output:𝖺𝖼𝖼𝖾𝗉𝗍/𝗋𝖾𝗃𝖾𝖼𝗍\\mathsf\{accept\}/\\mathsf\{reject\}\.1\.The protocol computes the negative log\-probability of each token via lookup operations:ℓi=Lookup​\(Tlog,qi\)\.\\ell\_\{i\}=\\texttt\{Lookup\}\(T\_\{\\log\},q\_\{i\}\)\.2\.The protocol checks thatℓ⋆\\ell^\{\\star\}andℓ\\ellare permutations of each other via a permutation argument\.3\.The protocol checks thatℓ⋆\\ell^\{\\star\}is sorted in descending order\. That is, definedi=ℓi⋆−ℓi\+1⋆d\_\{i\}=\\ell^\{\\star\}\_\{i\}\-\\ell^\{\\star\}\_\{i\+1\}fori=0,…,N−2i=0,\\ldots,N\-2and checkdi∈\[0,224\)d\_\{i\}\\in\[0,2^\{24\}\)via lookups of its limbs inTrangeT\_\{\\mathrm\{range\}\}\. Sinceℓi⋆<219\\ell^\{\\star\}\_\{i\}<2^\{19\}, an unsorted pair would givedi≥p−219d\_\{i\}\\geq p\-2^\{19\}in𝔽p\\mathbb\{F\}\_\{p\}, outside this range\.4\.The protocol sums thekkentries of the least likely tokens:sfinal=∑j=0k−1ℓj⋆\.s\_\{\\mathrm\{final\}\}=\\sum\_\{j=0\}^\{k\-1\}\\ell^\{\\star\}\_\{j\}\.5\.Ifsfinal<Ts\_\{\\mathrm\{final\}\}<T, outputaccept; otherwise, outputreject\. WithT=⌈−kε⋅215⌉T=\\lceil\-k\\varepsilon\\cdot 2^\{15\}\\rceil, this is the Min\-K% test of mean log\-probability aboveε\\varepsilon\.

Figure 5:Protocol for the Min\-K% membership inference zero\-knowledge proof\.
##### Implementation\.

We implement our protocol using the Plonky3\([Polygon, 2024](https://arxiv.org/html/2609.36437#bib.bib45)\)library\. The performance is measured using itsp3\-batch\-starkprover over the BabyBear field with a Poseidon2\-based FRI commitment scheme\. As baselines, we implement the same algorithm as two programs for the SP1 zkVM\([Succinct Labs, 2026](https://arxiv.org/html/2609.36437#bib.bib11)\): one computes the Min\-K% Prob score directly in FP32 without ZK\-friendly quantization, and the other uses our fixed\-point encoding\. The two programs differ only in number representation\. In all three implementations, the prover supplies the sorted order, the proof checks it in linear time, and only\(N,k,T\)\(N,k,T\)and the decision are public\. All implementations use the same FRI parameters as the SP1 core prover and are measured on a single socket of an Intel Xeon Gold 6526Y \(16 cores, 256GB DDR5\)\.

##### Results\.

Table[10](https://arxiv.org/html/2609.36437#A8.T10)summarizes the proving cost\. Within the same zkVM, ZK\-friendly quantization reduces the number of executed RISC\-V cycles compared with FP32, because each logarithm becomes a single table lookup instead of a software floating\-point routine\. For short texts, both programs are dominated by the fixed cost of the zkVM \(about 25 seconds\), but the gap in proving time grows with the text length\. Our protocol proves the decision in about 3 seconds for everyNNup to 4096\. This is faster than the fixed\-point zkVM program and up to 30 times faster than the FP32 program\. The cost is dominated by committing the two2162^\{16\}\-entry lookup tables, so it barely depends onNN\. Moreover, as shown in Table[8](https://arxiv.org/html/2609.36437#A6.T8), the W16A16 pipeline preserves near 98% of the bottom\-20% token positions selected by the BF16 model, which are exactly the tokens that Min\-K% aggregates\. Hence, the proved decision is computed over nearly the same tokens as the original model would select\.

In conclusion, this use case again confirms that ZK\-friendly quantization lowers the cost of proving the Min\-K% statistic\. It makes the computation expressible as lookups and field arithmetic, allowing our protocol to be proved in a few seconds regardless of the text length\.

Table 10:Cost of proving the Min\-K% decision with and without ZK\-friendly quantization \(mean of three proofs\)\.

## Appendix IDetailed Experimental Results

Table 11:Complete operator\-wise LUT ablations\. All values are perplexity ratios relative to the corresponding W16A16 LUT\-off baseline\. O: only the indicated operator family uses a B16 LUT; all other nonlinear functions are exact\. E: the full B16 pipeline uses the exact function only for the indicated operator family\.Table 12:LUT cost sensitivity on GPT\-2 Small\.Times are means \(standard deviation\) over three runs\. AtB=16B=16, all three columns use the same configuration\.

Table 13:Linear proving costs for one Qwen2\.5\-14B block\.

Similar Articles

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

arXiv cs.CL

CAT-Q introduces a post-training ternary quantization method for LLMs that uses learnable modulation and softened ternarization, achieving superior performance over BitNet 1.58-bit while using only 512 calibration samples and scaling to 235B parameters.