A caveman qwen3.6 27B

Reddit r/LocalLLaMA Models

Summary

A new model called grug-27b claims to outperform qwen3.6 27B while reducing token usage by over 90%, potentially making it much faster on consumer hardware.

Just saw this on huggingface: https://huggingface.co/ProCreations/grug-27b The benchmarks claim that it's quite a bit better than qwen3.6 27B original and that they reduced the amount of necessary tokens by more than 90%. It would make 27B running on my old laptop at 3tps feel more like 30tps for the thinking part, if true. Couldn't test it yet.
Original Article

Similar Articles

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

Hacker News Top

Cerebras has announced the availability of the Qwen 3.8 27B model on its inference platform with a speed of 1500 tokens per second, detailing its model compression techniques such as quantization and pruning.

Qwen3.6 27B Pure Quant: 40 tok/s on 16 GB VRAM

Reddit r/LocalLLaMA

A quantized version of Qwen3.6 27B using a pure Q4_K_M method fits entirely in 16 GB VRAM, achieving up to 40 tok/s token generation speed with MTP, and significantly reducing model size compared to other GGUF variants.

Qwen 3.8 27B is faster than expected

Reddit r/LocalLLaMA

A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.