Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM
Summary
Daniel Han of Unsloth validates that Qwen3.8-27B will run in only 17GB VRAM, making it accessible for local inference.
Similar Articles
My Qwen 3.8 27B tests on limited VRAM (16-20GB)
This article tests various quantized versions of the Qwen 3.8 27B AI model on limited VRAM setups, comparing their performance on tasks like animation generation, app development, and word generation.
@UnslothAI: Qwen3.8-27B is coming! Will run locally on 17GB RAM/VRAM setups.
Alibaba announces Qwen3.8-27B open-weights release, capable of running locally on 17GB RAM/VRAM, alongside the larger Qwen3.8-Max.
High VRAM local coding model — still Qwen 3.6 27B?
The user discusses their experience with Qwen 3.6 27B for local coding tasks and asks for recommendations for larger models (100B+) suitable for systems with 224GB of VRAM.
Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context
The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.
For those with 12GB GPUs, you can now run QWEN 3.6 27B wth little loss via the new Ternary version.
A new ternary quantized version of Qwen3.6 27B, called Bonsai 27B, allows running the model on 12GB GPUs with 10x less memory and 95% of original performance, making it accessible for local deployment.