Tag
This paper presents the first systematic study of zero-knowledge-friendly quantization for LLMs, formalizing the concept and evaluating nine models (including Qwen2.5-14B and Qwen3-30B-A3B) across weight, activation, and nonlinear lookup table precision. It finds that activation precision and RMSNorm inverse-square-root lookups dominate ZK proving costs and that conventional low-bit quantization heuristics do not translate to proportional proving savings.
The article discusses the concept of verifiable AI inference, exploring methods like trusted attestation and cryptographic proofs to ensure the authenticity and provenance of AI-generated outputs without rerunning the model.