@sgl_project: The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: -…

X AI KOLs Timeline Models

Summary

Qwen3.8-27B, a 27B-parameter multimodal AI model from Alibaba, is now open source with day-0 support in SGLang, offering high inference speeds and superior performance in coding and office tasks.

The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. Long live the (small model) king! Run it locally with SGLang
Original Article
View Cached Full Text

Cached at: 08/16/26, 08:05 PM

The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:

  • 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark
  • 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks.

Long live the (small model) king! Run it locally with SGLang

Qwen (@Alibaba_Qwen): We promised open weights for Qwen3.8. Now, time to meet them! 🎉

⚡ Qwen3.8-27B:

  • A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
  • 262K native context, easily extendable to 1M

Similar Articles

Qwen 3.8 27b is out. Big news for local AI

Reddit r/ArtificialInteligence

Qwen 3.8 27b, a sub-30 billion parameter AI model, has been released and is suitable for local inference on consumer hardware like RTX 3090 or M4 Pro, potentially replacing cloud-based AI subscriptions and shifting workflows locally.

Qwen3.6-27B

Reddit r/LocalLLaMA

Alibaba's Qwen team released Qwen3.6-27B, a new 27-billion-parameter language model, accompanied by benchmark results.

Qwen3.7 Preview lands on Arena (1 minute read)

TLDR AI

Alibaba Qwen announces two major model releases: Qwen3-Omni, the first natively end-to-end omni-modal AI unifying text, image, audio and video, and Qwen3-Next-80B-A3B, an ultra-efficient MoE model with 3B activated parameters per token, achieving SOTA performance and 10x faster inference than Qwen3-32B.

Qwen/Qwen3.6-27B-FP8

Hugging Face Models Trending

Alibaba releases Qwen3.6-27B-FP8, a 27B FP8-quantized model with strong agentic coding and reasoning benchmarks, now available on Hugging Face.