q2-k

Tag

Cards List
#q2-k

There's a new PR for llamacpp claiming to boost prompt processing with rocm by around 15%, also fixes a bug which makes Q2_K 28x faster

Reddit r/LocalLLaMA · 2026-07-21 Cached

A new PR for llama.cpp boosts prompt processing on ROCm by ~15% and fixes a bug making Q2_K quantization 28x faster.

0 favorites 0 likes
#q2-k

@antirez: GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

X AI KOLs Following · 2026-06-28 Cached

GLM 5.2 model runs with Q2_K quantized routed experts (effective 2.6 bits) using SSD streaming on an M5 Max 128GB computer.

0 favorites 0 likes
← Back to home

Submit Feedback