awq

Tag

Cards List
#awq

@akshay_pachaar: LLM quantization techniques I'd learn if I had to fit a 70B model on a single GPU: (bookmark this) A 70B model in FP16 …

X AI KOLs Following · 2026-07-22 Cached

A Twitter thread explains five key LLM quantization techniques (RTN, GPTQ, AWQ, LLM.int8(), QAT) for fitting large models on limited hardware, and references a comprehensive study paper.

0 favorites 0 likes
#awq

Why is AutoRound being slept on so hard?

Reddit r/LocalLLaMA · 2026-06-21

A user questions why AutoRound, a quantization tool offering superior accuracy retention at low bits and direct GGUF export, is overlooked despite outperforming standard AWQ and RTN, especially on complex models like Qwen3.6 27B.

0 favorites 0 likes
#awq

DiffusionGemma 26b on a 4090 at up to 475t/s... and some thoughts...

Reddit r/LocalLLaMA · 2026-06-18

A user shares their experience running DiffusionGemma 26B on a 4090 GPU via vLLM, achieving up to 475t/s but noting drawbacks like single-user limitation, lower accuracy, and short context, concluding it's not worth using over the regular 26B model.

0 favorites 0 likes
#awq

Introducing cyankiwi AWQ 4-bit Quantization — 26.05 update

Reddit r/LocalLLaMA · 2026-05-14

Cyankiwi introduced an updated version of their AWQ 4-bit quantization method that jointly optimizes scales and quantization ranges, achieving lower KL divergence than existing methods on Llama-3 models.

0 favorites 0 likes
← Back to home

Submit Feedback