Qwen 30b MoE - 30tps - 6GB vram - Done!

Reddit r/LocalLLaMA News

Summary

A user shares success running Qwen 30B MoE on an RTX 3050 6GB with 90k context at 30-35 tokens per second, celebrating low-end hardware performance.

So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of those experts and give me room for context. What did I get 10 or less tokens per second. πŸ˜„ Not today!! Today I could run it with 90k context with Hermes I had 20-25 tps. And when I changed harness I got even 30-35tps πŸ₯³πŸ₯³πŸ₯³ NOT benchmarks- but actual session generation with context and actual work being done πŸ˜„πŸ˜„πŸ˜„ I will come to edit the post and add details. Just wanted to share the joy with anyone out there with a peasant rig like mine πŸ˜… May be someone who does better can also share the positive vibe. Cheers for now πŸ™‹πŸΎβ€β™‚οΈ
Original Article

Similar Articles

Qwen3.6-35B-A3B Q4 262k context on 8GB 3070 Ti = +30tps

Reddit r/LocalLLaMA

The author shares detailed tuning tips for running the Qwen3.6-35B-A3B MoE model on an 8GB RTX 3070 Ti with up to 262k context using llama.cpp, achieving 30+ tps, and notes a 25% speed boost when switching from Windows to Ubuntu Server.

Qwen 35B-A3B is very usable with 12GB of VRAM

Reddit r/LocalLLaMA

A user benchmarks Qwen 35B-A3B (a 35B MoE model) on a 12GB RTX 3060, finding that 12GB VRAM is a practical sweet spot for running the model with 32k context, achieving ~47 t/s generation.

Wow! Qwen 3.6:35b-a3b on a 3090... pretty amazing.

Reddit r/artificial

A user shares impressive results running a quantized Qwen 3.6:35b-a3b model on a used RTX 3090, achieving 160 tokens per second output after fitting the model into VRAM, and demonstrates vision capabilities with a 75-second video processing time.