@cniongolo: I’m not sure people realize yet that you can actually run Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated-MTP-GGUF on a dua…
Summary
Demonstrates running a custom Qwen model (Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated-MTP-GGUF) on dual Nvidia RTX PRO 6000 Blackwell GPUs at 195 tokens per second using Hugging Face Inference.
View Cached Full Text
Cached at: 06/09/26, 10:46 AM
I’m not sure people realize yet that you can actually run Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated-MTP-GGUF on a dual‑GPU setup with Nvidia RTX PRO 6000 Blackwell cards and still hit around 195 tokens per second on @huggingface with huggingface Inference ! Already tried it with the pi agent!
@julien_c
Similar Articles
@RoundtableSpace: Qwen 3.8 27B running locally on an RTX 5090 beats Opus 4.8 in personal benchmarks at up to 200 tokens per second with n…
Qwen 3.8 27B, a 27-billion parameter AI model, outperforms Opus 4.8 in personal benchmarks when running locally on an RTX 5090 GPU, achieving up to 200 tokens per second without internet or API access.
Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable
User reports surprisingly usable coding performance from Qwen3-27B-UD-Q6_K_XL.gguf running locally on RTX 5090 at ~50 tok/s with 200K context, marking a significant leap in local model quality.
Wow! Qwen 3.6:35b-a3b on a 3090... pretty amazing.
A user shares impressive results running a quantized Qwen 3.6:35b-a3b model on a used RTX 3090, achieving 160 tokens per second output after fitting the model into VRAM, and demonstrates vision capabilities with a 75-second video processing time.
Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72
NVIDIA showcases the high-throughput performance of serving the Qwen3-8B 2.4T parameter model on GB300 NVL72 hardware, achieving over 4k tokens per second per GPU.
Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti
A user successfully runs the Qwen 3.8 27B AI model on a mixed setup of RTX 3060 and 5060 Ti GPUs using tensor parallelism with exllamav3, achieving around 50 tokens per second with MTP enabled.