A caveman qwen3.6 27B
Summary
A new model called grug-27b claims to outperform qwen3.6 27B while reducing token usage by over 90%, potentially making it much faster on consumer hardware.
Similar Articles
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Cerebras has announced the availability of the Qwen 3.8 27B model on its inference platform with a speed of 1500 tokens per second, detailing its model compression techniques such as quantization and pruning.
Qwen3.8-27B: slower tokens, faster and better results
Qwen3.8-27B is a new AI model that emphasizes wall-clock time over token speed, offering superior intelligence for local deployment on consumer hardware with 32 GB VRAM.
Qwen3.6 27B Pure Quant: 40 tok/s on 16 GB VRAM
A quantized version of Qwen3.6 27B using a pure Q4_K_M method fits entirely in 16 GB VRAM, achieving up to 40 tok/s token generation speed with MTP, and significantly reducing model size compared to other GGUF variants.
yall are sleeping on qwen 3.8 27b q2 + q2 dflash + q5 kv
A user shares their experience running a quantized Qwen 3.8 27B model using QAT Q2 and Q5 KV, achieving high performance on a 12GB GPU with up to 200K token context, surpassing models like Sonnet 4.6.
Qwen 3.8 27B is faster than expected
A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.