Tag
A new PR for llama.cpp boosts prompt processing on ROCm by ~15% and fixes a bug making Q2_K quantization 28x faster.
GLM 5.2 model runs with Q2_K quantized routed experts (effective 2.6 bits) using SSD streaming on an M5 Max 128GB computer.