TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram

Reddit r/LocalLLaMA News

Summary

A user shares their experience running the Qwen 3.8 flash next model on a system with 6GB VRAM and 16GB RAM using llama.cpp, achieving 6-7 tokens per second with 1-bit quantization, and asks for recommendations on quantization variants.

After getting tired refreshing and searching sub reddits for qwen 3.8 35 a3b i decided to give qwen 3.8 flash next a try . So I dual booted and built llama cpp I am able to hit 6-7 tps constantly inside ubuntu just with 16 gb vram and 6 gb gpu . Honestly it is fire for 1 bit quant . Will increase the quant variant till I get somewhat decent speed and acceptable results . What quant will be better to try . I don't wanna download delete and redownload the whole day . https://preview.redd.it/vreqm0btn4mh1.png?width=868&format=png&auto=webp&s=5ccd5d6152bd4c7f52782b79e03af5720c219975
Original Article

Similar Articles

Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context

Reddit r/LocalLLaMA

The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.