1.3B token used recently, qwen3.8 27b

Reddit r/LocalLLaMA Models

Summary

A Reddit post highlights the strong performance of the Qwen 3.8 27B model, with mention of 1.3B tokens used recently.

https://preview.redd.it/ako9tw504yph1.png?width=1257&format=png&auto=webp&s=8ab3cc2263f24022faddfdc4b8736d366a985aa8 Very strong, very strong model.
Original Article

Similar Articles

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

Hacker News Top

Cerebras has announced the availability of the Qwen 3.8 27B model on its inference platform with a speed of 1500 tokens per second, detailing its model compression techniques such as quantization and pruning.

A caveman qwen3.6 27B

Reddit r/LocalLLaMA

A new model called grug-27b claims to outperform qwen3.6 27B while reducing token usage by over 90%, potentially making it much faster on consumer hardware.

Qwen 3.8 27B is faster than expected

Reddit r/LocalLLaMA

A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.