gptq

Tag

Cards List
#gptq

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

Hugging Face Daily Papers · 2026-07-29 Cached

This paper introduces GPTQ-2D, a method for two-sided adaptive rounding that produces identical results to applying GPTQ on vectorized matrices but runs in cubic time instead of quartic time.

0 favorites 0 likes
#gptq

@akshay_pachaar: LLM quantization techniques I'd learn if I had to fit a 70B model on a single GPU: (bookmark this) A 70B model in FP16 …

X AI KOLs Following · 2026-07-22 Cached

A Twitter thread explains five key LLM quantization techniques (RTN, GPTQ, AWQ, LLM.int8(), QAT) for fitting large models on limited hardware, and references a comprehensive study paper.

0 favorites 0 likes
#gptq

Qwen3.6 35B A3B uncensored heretic Native MTP Preserved is Out Now With KLD 0.0015, 10/100 Refusals and the Full 19 MTPs Preserved and Retained, Available in Safetensors, GGUFs. NVFP4, NVFP4 GGUFs and GPTQ-Int4 Formats

Reddit r/LocalLLaMA · 2026-05-09

Community release of Qwen3.6 35B A3B uncensored variant with full 19 MTP tensors preserved, available in multiple formats including Safetensors, GGUF, NVFP4 and GPTQ-Int4.

0 favorites 0 likes
← Back to home

Submit Feedback