Tag
This paper introduces GPTQ-2D, a method for two-sided adaptive rounding that produces identical results to applying GPTQ on vectorized matrices but runs in cubic time instead of quartic time.
A Twitter thread explains five key LLM quantization techniques (RTN, GPTQ, AWQ, LLM.int8(), QAT) for fitting large models on limited hardware, and references a comprehensive study paper.
Community release of Qwen3.6 35B A3B uncensored variant with full 19 MTP tensors preserved, available in multiple formats including Safetensors, GGUF, NVFP4 and GPTQ-Int4.