Tag
PRQuant is a training-free and low-overhead framework for quantizing linear layers in large language models, using permutation and residual compensation to reduce inference latency while improving accuracy over baselines like MXFP4.
Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.
HyQuant improves low-bit LLM attention quantization by preserving critical vertical-line tokens and local windows in high precision while quantizing the rest, maintaining accuracy with low overhead.
ExTernD introduces an expanded-rank ternary decomposition for post-training LLM quantization, enabling accuracy approaching bf16 by using a factored representation with free inner rank. It matches Q4_K accuracy at 5.2-5.5 effective bits per weight on models like Gemma-4 and Qwen3.5.
InfoQuant introduces a train-free method, Peak Suppression Orthogonal Transformation (PSOT), to reshape activation distributions for low-bit LLM quantization, preserving 97% floating-point accuracy under W4A4KV4 and outperforming prior PTQ methods.
OSCAR is an offline spectral covariance-aware rotation method for 2-bit KV cache quantization that aligns quantization with attention covariance structures, achieving high accuracy and efficiency for long-context LLM serving.