@sgl_project: The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: -…
Summary
Qwen3.8-27B, a 27B-parameter multimodal AI model from Alibaba, is now open source with day-0 support in SGLang, offering high inference speeds and superior performance in coding and office tasks.
View Cached Full Text
Cached at: 08/16/26, 08:05 PM
The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:
- 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark
- 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks.
Long live the (small model) king! Run it locally with SGLang
Qwen (@Alibaba_Qwen): We promised open weights for Qwen3.8. Now, time to meet them! 🎉
⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M
Similar Articles
@rohanpaul_ai: Alibaba dropped the weights for Qwen3.8-27B as a 27B open-weight multimodal model built for local deployment. - Apache …
Alibaba released Qwen3.8-27B, a 27B open-weight multimodal model for local deployment, which shows frontier-class performance on coding benchmarks like SWE-bench Pro and OSWorld, outperforming some larger models.
Qwen 3.8 27b is out. Big news for local AI
Qwen 3.8 27b, a sub-30 billion parameter AI model, has been released and is suitable for local inference on consumer hardware like RTX 3090 or M4 Pro, potentially replacing cloud-based AI subscriptions and shifting workflows locally.
Qwen3.6-27B
Alibaba's Qwen team released Qwen3.6-27B, a new 27-billion-parameter language model, accompanied by benchmark results.
Qwen3.7 Preview lands on Arena (1 minute read)
Alibaba Qwen announces two major model releases: Qwen3-Omni, the first natively end-to-end omni-modal AI unifying text, image, audio and video, and Qwen3-Next-80B-A3B, an ultra-efficient MoE model with 3B activated parameters per token, achieving SOTA performance and 10x faster inference than Qwen3-32B.
Qwen/Qwen3.6-27B-FP8
Alibaba releases Qwen3.6-27B-FP8, a 27B FP8-quantized model with strong agentic coding and reasoning benchmarks, now available on Hugging Face.