40+tok/s - optimized recipe for Qwen 3.5 122B Int4 on a single DGX Spark with vLLM

Reddit r/LocalLLaMA Tools

Summary

User shares an optimized recipe for running Qwen 3.5 122B Int4 on a single DGX Spark with vLLM, achieving over 40 tokens per second. They invite others to try and further optimize it.

Hello guys, two days ago i ran the spark-arena for my Qwen 3.5 122B Recipe on a single DGX Spark and I got the highest score on speed for any context length and concurrency across all 3.5 122B Int4 Recipes. Just wanted to share if somebody wants to try, play around with it and optimize it further. [https://spark-arena.com/benchmark/sub1779146508448](https://spark-arena.com/benchmark/sub1779146508448) https://preview.redd.it/pz2dr3n4fb2h1.png?width=1099&format=png&auto=webp&s=40f078ae3df597545d08ed3df008f84873acca6a
Original Article

Similar Articles

Qwen3.6:35b UD Q4_K_M 80 tok/s on Nvidia P40

Reddit r/LocalLLaMA

A user shares achieving 80 tok/s on a Qwen3.6 35B model with Q4_K_M quantization and 100k context on a single Nvidia P40 using TheTom's TurboQuant fork of llama.cpp, highlighting various optimizations.

Qwen 3.8 27B is faster than expected

Reddit r/LocalLLaMA

A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.