Tag
A 1.6 MB model using ternary 3:4 weights runs on WebGPU in the browser, playing Connect Four as well as a 7.8 MB int8 model, demonstrating the effectiveness of ternary formats for lightweight AI deployment.
This paper revisits Kashin-decomposition-based weight quantization for large language models and proposes an improved algorithm using structured orthogonal transforms, reducing computational cost and ensuring numerical stability compared to methods like OPTQ and QuIP.
CubicQuant proposes a parametric non-uniform scalar quantization format for LLM weights, using a monotonic cubic curve to adapt reconstruction levels at 1-8 bit widths while retaining dense integer code streams for GPU efficiency. Experiments show RMSE reductions over uniform and floating-point baselines, with preliminary H200 kernel measurements.
Introduces QAM-W, a joint 2D codebook quantization method for LLM weights using Hadamard rotation and activation-aware scaling, achieving near BF16 perplexity at 5–6 bits per weight and matching SmoothQuant W8A8 quality with 32% fewer weight bits.