Tag
A user's evaluation demonstrates that Qwen 3.8 27b performs excellently on professional certification tests in real estate and finance, achieving scores up to 98.44% with tools and guided search.
Qwen Lab has released Qwen 3.8 27B, which shows significant improvement over previous versions like Qwen 3.6 27B and other open-source models. The author hopes that Qwen publishes papers to help other labs develop similar high-quality small models.
A developer excitedly queues tests for Qwen's new 27B model, which Qwen promises brings a whole new level of capability.
Daniel Han of Unsloth validates that Qwen3.8-27B will run in only 17GB VRAM, making it accessible for local inference.
Qwen announced Qwen3.8, including a new 27B model, generating excitement for local deployment.
A developer built a testing harness that measures KL divergence per weight group during quantization, leading to three custom quantized builds of Qwen3.6-27B (Bedrock, Tightrope, Gambit) with optimized compression. Tool calling is identified as the first capability to degrade under quantization.
DavidAU releases Qwen3.6-27B-Fable-Fusion-711, a multi-stage fine-tune of Qwen 3.6 27B that claims to exceed 700 ARC-C benchmark, surpassing base models and matching closed-source models, available in GGUF format for consumer hardware.
DFlash is a method that accelerates Qwen3.6 27B model inference by 2.2x without quality degradation.
Prism-ML published benchmarks for their Bonsai-27B model.
User tests ternary quantized Qwen3.6 27B on an RTX 3090, achieving 60 tk/s with two slots and 100k KV cache using 21GB VRAM, with good quality and stable tool calls.
Running Qwen3.6 27B on an RTX 5090, achieving 6.4k tokens per second after tuning MTP and cache settings, demonstrating optimization techniques for inference.
Prism ML releases Ternary-Bonsai-27B-mlx-2bit, a ternary-quantized 27B-parameter language model that achieves ~95% of FP16 performance while fitting in ~7.2 GB, enabling full reasoning on laptops.
Mention of Qwen 3.6 27b model in context of Dspark.
A tweet promoting the Qwen 3.6 27b model and recommending UnslothAI for running it on any GPU.
A Hugging Face repository (kaitchup/Qwen3.6-27B-GGUF-MoQ) provides GGUF quantized weights for the Qwen3.6-27B MoQ model, enabling local inference with tools like llama.cpp and Ollama.
A GGUF quantized version of the Qwopus3.6-27B-Coder-MTP model is released on Hugging Face, optimized for local inference and compatible with Transformers, vLLM, SGLang, and Unsloth Studio.
MooreThreads releases MusaCoder-27B, a 27-billion-parameter code generation model, accompanied by a paper on arXiv.
User shares experience with Qwen3.6 27B model, which successfully generated a complete HTML5 breakout game in one shot, showing impressive coherence and attention to detail beyond typical LLM outputs.
Qwen 3.6 27B runs fast on 16 GB VRAM thanks to 'Pure Quant' technology, achieving 40 tokens/s with MTP and supporting 64k contexts, enabling local AI on consumer GPUs like RTX 4060 Ti.
Qwen is highly likely to release a 27B parameter model, though the exact roadmap is still pending.