1.3B token used recently, qwen3.8 27b
Summary
A Reddit post highlights the strong performance of the Qwen 3.8 27B model, with mention of 1.3B tokens used recently.
Similar Articles
Qwen Developers' responses from their recent Twitter/X AMA
Qwen developers held an AMA on Twitter, confirming an upcoming 27B model, sharing details about Qwen 3.8's 2.4T parameters and 95B active, and noting the community's influence on releases.
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Cerebras has announced the availability of the Qwen 3.8 27B model on its inference platform with a speed of 1500 tokens per second, detailing its model compression techniques such as quantization and pruning.
Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72
NVIDIA showcases the high-throughput performance of serving the Qwen3-8B 2.4T parameter model on GB300 NVL72 hardware, achieving over 4k tokens per second per GPU.
A caveman qwen3.6 27B
A new model called grug-27b claims to outperform qwen3.6 27B while reducing token usage by over 90%, potentially making it much faster on consumer hardware.
Qwen 3.8 27B is faster than expected
A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.