Here's your recipes for models for the most popular hardware

Reddit r/LocalLLaMA Tools

Summary

A curated collection of LLM inference recipes and performance benchmarks for popular hardware configurations, including NVIDIA RTX Pro 6000 Blackwell, H100, AMD Strix Halo, and RTX 5090, with detailed batch size, quantization, and throughput data.

No content available
Original Article
View Cached Full Text

Cached at: 07/11/26, 11:40 PM

# Recipes — LLMRequirements.com Source: [https://llmrequirements.com/recipes](https://llmrequirements.com/recipes) [8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)kimi\-k2\-5\-1t\-moe@ INT4 \(FP8 KV, DCP=8\) on vLLMbatch~900@ 40K\(100\-conc\. aggregate\)—[local\-inference\-lab/rtx6kpro wiki \(Kimi K2\.5 high concurrency, Festr\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ FP8 \(fp8 KV, MTP spec=3\) on vLLMbatch~894@ 175K\(32\-conc\. aggregate\)—[theogravity/dual\-rtx\-6000\-blackwell\-qwen3\.6\-27b\-fp8 \(coding sweep, seqs=32\)↗](https://github.com/theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)deepseek\-v3\-671b\-moe@ FP8 \(TP=8\) on vLLMbatch~620@ 1\.024K\(100\-conc\. aggregate\)—[dzhsurf/deepseek\-v3\-r1\-deploy\-and\-benchmarks \(8xH100 vLLM TP=8, ~100 concurrency\)↗](https://github.com/dzhsurf/deepseek-v3-r1-deploy-and-benchmarks)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)gemma4\-26b\-moe@ native on vLLM cluster TP=2 \(triton, ROCm\)batch~411@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(triton cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/triton/google_gemma-4-26B-A4B-it_cluster_tp2_throughput.json)[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)qwen3\-5\-397b\-a17b\-moe@ NVFP4 on SGLang\+MTP~350@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Qwen3\.5\-397B 8x single\-batch\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~320@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-6\-35b\-a3b\-moe@ AWQ\-4bit / native on vLLM cluster TP=2 \(aiter, ROCm\)batch~287@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)gpt\-oss\-120b@ MXFP4 on vLLM cluster TP=2 \(triton, ROCm\)batch~229@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(triton cluster tp2 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[Single RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-32)qwen3\-coder\-30b@ Q4\_K on llama\.cpp \(CUDA\)~226@ 4K~7093@ 4K[hardware\-corner\.net RTX 5090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-5090/)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)mistral\-medium\-3\-5\-128b@ FP8 on vLLM~220@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/mistralai/Mistral-Medium-3.5-128B)[Single RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-32)gemma4\-26b\-moe@ Q4\_K on llama\.cpp \(CUDA\)~180@ 4K~8799@ 4K[hardware\-corner\.net RTX 5090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-5090/)[Single H100 80 GB workstation](https://llmrequirements.com/hardware/h100-80)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~180@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 \(FP8 KV, MTP n=2\) on vLLMbatch~179\.9@ 393\.216K\(8\-conc\. aggregate\)—[NVIDIA Developer Forum 373808 \(jasl vLLM TP=4, n=8 aggregate\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-on-4x-dgx-spark-via-vllm-jasl-fork-tp-4-rdma-mtp-49-54-tok-s-single-stream-full-recipe-the-traps/373808)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ AWQ\-4bit / native on vLLM \(aiter, ROCm\)batch~178@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter tp1 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~170@ 262K—[GitHub lastloop\-ai↗](https://github.com/lastloop-ai/vllm-blackwell-guide)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ NVFP4 on vLLM\+MTP~156@ 262K~831@ 262K[loFT LLC↗](https://loftllc.dev/en/docs/tech/llm-research/qwen3-6-27b-nvfp4-mtp-vllm-benchmark/)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-coder\-30b@ Q4\_K on llama\.cpp \(CUDA\)~153@ 4K~2988@ 4K[hardware\-corner\.net RTX 3090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/)[Quad RTX Pro 6000 Blackwell build \(384 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)qwen3\-5\-397b\-a17b\-moe@ AWQ\-INT4 \(QuantTrio\) on SGLang\+MTP~152@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Qwen3\.5\-397B single\-batch decode\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[Dual RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x2)qwen3\-6\-35b\-a3b\-moe@ AWQ on vLLM~149@ 4K—[GitHub \- tfriedel \(RTX 3090 lab\)↗](https://github.com/tfriedel/qwen3.6-rtx3090-lab)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ ROCmFP4 \(CHADROCK\) on llama\-server\+ROCmFPX~140@ 4K—[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ FP8 \(native, fp8 KV, MTP spec=3\) on vLLM~137@ 175K—[theogravity/dual\-rtx\-6000\-blackwell\-qwen3\.6\-27b\-fp8 \(benchmark sweep\)↗](https://github.com/theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)mistral\-small\-4\-119b\-moe@ NVFP4 on vLLMbatch~131@ 262\.144K\(20\-conc\. aggregate\)—[Sebastien67 Medium \(DGX Spark vLLM NVFP4, n=20 aggregate\)↗](https://medium.com/@Sebastien67/running-mistral-small-4-119b-nvfp4-locally-on-a-dgx-spark-81cc2fdc4f6f)[Quad RTX Pro 6000 Blackwell build \(384 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)qwen3\-5\-397b\-a17b\-moe@ NVFP4 \(nvidia checkpoint\) on vLLM\+MTP~130@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Qwen3\.5\-397B MTP scaling table, concurrency=1\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)gemma\-4\-31b@ native \(bf16/fp16\) on vLLM cluster TP=2 \(triton, ROCm\)batch~128@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(triton cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/triton/google_gemma-4-31B-it_cluster_tp2_throughput.json)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 \(modelopt\_fp4, fp8 KV\) on SGLang\+MTPbatch~124@ 196\.608K\(8\-conc\. aggregate\)—[NVIDIA Developer Forum 373676 \(SGLang TP=4 EP=4, n=8 aggregate\)↗](https://forums.developer.nvidia.com/t/minimax-m2-5-nvfp4-on-4x-dgx-spark-via-sglang-tp-4-ep-4-124-tok-s-aggregate-n-8-fixing-the-cutlass-moe-compile-oom-with-max-jobs-1/373676)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-next\-80b\-moe@ Q4\_K\_M on ollama / llama\.cpp \(CUDA\)~124@ 4K~3274@ 4K[vaditaslim\.com RTX PRO 6000 Blackwell 8\-model benchmarks↗](https://www.vaditaslim.com/blog/ai/local-llm-benchmarks-rtx-pro-6000)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)gemma4\-26b\-moe@ Q4\_K on llama\.cpp \(CUDA\)~119@ 4K~3625@ 4K[hardware\-corner\.net RTX 3090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/)[Single H100 80 GB workstation](https://llmrequirements.com/hardware/h100-80)qwen3\-6\-27b\-dense@ FP16 on vLLM~110@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/Qwen/Qwen3.6-27B)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-5\-122b\-a10b\-moe@ cyankiwi AWQ\-4bit on vLLM cluster TP=2 \(aiter, ROCm\)batch~104@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_XL on llama\.cpp~101@ 65K~1171@ 0\.5K[aminrj\.com \(Qwen3\.6 on 24GB\)↗](https://aminrj.com/posts/llamacpp-qwen36-35b/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ IQ4\_XS\-Q8nextn on llama\-server\+MTP~101@ 4K—[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)kimi\-k2\-5\-1t\-moe@ INT4 \(BF16 KV, EP=8, overclocked GDDR7\) on SGLang\+MTP~101@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Kimi K2\.5 8x single\-batch decode\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-6\-27b\-dense@ INT4 \(AutoRound\) \+ MTP n=3, FP8 KV on vLLM \(flashinfer, MTP\)~100@ 262K—[GitHub lastloop\-ai↗](https://github.com/lastloop-ai/vllm-blackwell-guide)[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)glm\-51\-754b\-moe@ NVFP4\-MTP \(lukealonso/GLM\-5\.1\-NVFP4\-MTP, served as GLM\-5\) on SGLang\+MTP~100@ 4K—[local\-inference\-lab/rtx6kpro wiki \(GLM\-5 single\-batch decode; models/glm5\.md = GLM\-5\.1\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on vLLM\+DFlash~97@ 0\.5K~9090@ 0\.5K \(derived\)[GitHub AEON\-7↗](https://github.com/AEON-7/Qwen3.6-NVFP4-DFlash)[MacBook Pro M5 Max 64 GB](https://llmrequirements.com/hardware/mbp-m5-max-64)gemma\-4\-12b@ MLX NVFP4 on Ollama 0\.31 \(MLX\) \+ MTP~95@ 4K—[Ollama blog \(framework\-author first\-party; M5 Max, Aider polyglot\)↗](https://ollama.com/blog/faster-gemma-4-mlx-mtp)[Single RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-32)qwen3\-6\-27b\-dense@ NVFP4 on vLLM~92@ 200K~5300@ 47K[GitHub devnen↗](https://github.com/devnen/qwen3.6-windows-server)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on vLLM~90@ 43K~2133@ 32K \(derived\)[GitHub technigmaai/dgx\-spark↗](https://github.com/technigmaai/dgx-spark/tree/main/spark-vllm-docker/nvidia-Qwen3.6-35B-A3B-NVFP4)[Dual RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x2)qwen3\-6\-27b\-dense@ AWQ on vLLM~90@ 100K—[GitHub \- tfriedel \(RTX 3090 lab\)↗](https://github.com/tfriedel/qwen3.6-rtx3090-lab)[MacBook Pro M5 Max 64 GB](https://llmrequirements.com/hardware/mbp-m5-max-64)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~87@ 4K~2447@ 4K[oMLX Benchmark↗](https://omlx.ai/benchmarks/oykgm8sq)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 on vLLM~85@ 4K—[local\-inference\-lab/rtx6kpro wiki \(MiniMax\-M2\.5 single\-stream table\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ Q4\_0 on llama\.cpp~81@ 4K~1244@ 0\.5K[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[Quad RTX Pro 6000 Blackwell build \(384 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-FP8 on vLLM~81@ 20K—[local\-inference\-lab/rtx6kpro wiki \(MiniMax\-M2\.5 single\-stream table\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[DGX B200 — 8× B200 server \(1\.44 TB HBM3e\)](https://llmrequirements.com/hardware/b200-x8)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \+ FP8 KV on Dynamo \+ vLLM \(TP=4, expert\-parallel, MTP\)batch~80\.6@ 4K\(20\-conc\. aggregate\)—[NVIDIA ai\-dynamo/dynamo recipes \(B200 TP4\+EP, NVFP4\+FP8, MTP\)↗](https://github.com/ai-dynamo/dynamo/tree/main/recipes/nemotron-3-ultra)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-AWQ\-INT4 \(fp8 KV, EAGLE3\) on vLLMbatch~79@ 262\.144K\(4\-conc\. aggregate\)—[NVIDIA Developer Forum 375361 \(vLLM TP=4, n=4 aggregate\)↗](https://forums.developer.nvidia.com/t/minimax-m3-awq-running-tp-4-across-4x-dgx-spark-gb10-33-tok-s-full-recipe-the-gb10-build-fixes/375361)[Single AMD Radeon AI Pro R9700 32 GB build](https://llmrequirements.com/hardware/amd-r9700-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M on llama\.cpp~77@ 4K~1636@ 32K[GitHub truelies444↗](https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)kimi\-k2\-6\-1t\-moe@ FP8 on vLLM~75@ 256K—[HF \- RedHatAI \(Kimi\-K2\.6\-FP8\-BLOCK\)↗](https://huggingface.co/RedHatAI/Kimi-K2.6-FP8-BLOCK)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ MTP\-GGUF UD\-Q4\_K\_XL \(draft\-mtp n=3\) on llama\.cpp \(Vulkan RADV, MTP\)~75@ 0\.5K—[kyuz0 amd\-strix\-halo\-toolboxes MTP grid \(results\-mtp/summary\.json, 15 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/mtp.html)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-5\-122b\-a10b\-moe@ cyankiwi AWQ\-8bit on vLLM cluster TP=2 \(aiter, ROCm\)batch~74@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/aiter/cyankiwi_Qwen3.5-122B-A10B-AWQ-8bit_cluster_tp2_throughput.json)[Dual AMD Radeon AI Pro R9700 build \(64 GB\)](https://llmrequirements.com/hardware/amd-r9700-x2)qwen3\-6\-35b\-a3b\-moe@ Q6\_K on llama\.cpp~72@ 4K~3038@ 32K[GitHub truelies444↗](https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-6\-27b\-dense@ AWQ/AutoRound\-INT4 on vLLM\+MTP~72@ 32K—[GitHub devnen↗](https://github.com/devnen/qwen3.6-windows-server)[Single RTX 4090 build](https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-coder\-30b@ Q4\_K on llama\.cpp \(CUDA\)~68@ 64K~1502@ 64K[hardware\-corner\.net RTX 4090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-4090/)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ not stated \(native precision\) on vLLM \(Aidendle94/B12X\-MoE, TP=2 RoCE\) \+ DSpark spec\-decode~65@ 200K—[GitHub 0rand \(DeepSeek\-V4 DSpark serving stack\)↗](https://github.com/0rand/DeepSeek-v4-DSpark-Aidendle94-GB10-ServingStack)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ NVFP4\-KV \(nvfp4\_ds\_mla\) on vLLM\+DSpark~63@ 200K—[NVIDIA Developer Forum \(374846\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-dspark-on-2x-dgx-spark-gb10-big-single-stream-speed-boost-60-67-tok-s-1m-context-now-with-concurrency/374846)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M \(UD\) on llama\.cpp~62@ 4K~1059@ 0\.5K[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[8× Strix Halo cluster \(1024 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x8)qwen3\-6\-35b\-a3b\-moe@ Q8\_0 on llama\.cpp~62@ 4K—[GitHub \- strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~60@ 32K~6520@ 8K \(derived\)[NVIDIA Developer Forum \(366822\)↗](https://forums.developer.nvidia.com/t/qwen-qwen3-6-35b-a3b-and-fp8-has-landed/366822)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~60@ 0\.5K~1114@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[DGX H200 — 8× H200 server \(1\.13 TB HBM3e\)](https://llmrequirements.com/hardware/h200-x8)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \+ FP8 KV on Dynamo \+ vLLM \(TP=8, expert\-parallel, MTP\)batch~58\.7@ 4K\(10\-conc\. aggregate\)—[NVIDIA ai\-dynamo/dynamo recipes \(8xH200 TP8\+EP, NVFP4\+FP8, MTP\)↗](https://github.com/ai-dynamo/dynamo/tree/main/recipes/nemotron-3-ultra)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)minimax\-m2\-7\-230b\-moe@ cyankiwi AWQ\-4bit on vLLM cluster TP=2 \(aiter, ROCm\)batch~57@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/aiter/cyankiwi_MiniMax-M2.7-AWQ-4bit_cluster_tp2_throughput.json)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gpt\-oss\-120b@ MXFP4 on llama\.cpp \(Vulkan RADV\)~56@ 0\.5K~720@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[Single Intel Arc Pro B70 build](https://llmrequirements.com/hardware/intel-arc-pro-b70-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M \(UD\) on llama\.cpp \(SYCL\)~55@ 4K~615@ 0\.5K[GitHub PMZFX↗](https://github.com/PMZFX/intel-arc-pro-b70-benchmarks/blob/master/llm-benchmarks.md)[Tesla V100 32 GB SXM2 mod build](https://llmrequirements.com/hardware/tesla-v100-sxm2-mod)qwen3\-6\-35b\-a3b\-moe@ Q5\_K\_M on llama\.cpp~55@ 10K—[GitHub \- ai\-bond \(V100 flash\-attn\)↗](https://github.com/ai-bond/flash-attention-v100)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+DFlash~54@ 0\.5K~2083@ 50K \(derived\)[GitHub HeNryous \(renek\)↗](https://github.com/HeNryous/mimo-v25-dflash-dgx-spark)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gemma4\-26b\-moe@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~54@ 0\.5K~1324@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[MacBook Pro M5 Max 64 GB](https://llmrequirements.com/hardware/mbp-m5-max-64)gemma\-4\-12b@ MLX NVFP4 on Ollama 0\.31 \(MLX\), no MTP~50\.2@ 4K—[Ollama blog \(framework\-author first\-party; M5 Max\)↗](https://ollama.com/blog/faster-gemma-4-mlx-mtp)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM\+DFlash~50@ 262K~4932@ 0\.5K[GitHub ZengboJamesWang↗](https://github.com/ZengboJamesWang/dgx-spark-vllm-qwen3.6-35b-a3b-dflash)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)qwen3\-6\-35b\-a3b\-moe@ Q8\_0 on llama\.cpp~50@ 4K—[Frame\.work Community↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 \(official weights, FP8 KV, MTP n=2\) on vLLM~49\.4@ 393\.216K—[NVIDIA Developer Forum 373808 \(jasl vLLM TP=4\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-on-4x-dgx-spark-via-vllm-jasl-fork-tp-4-rdma-mtp-49-54-tok-s-single-stream-full-recipe-the-traps/373808)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM\+MTP~45\.5@ 1000K~786@ 800K[GitHub tonyd2wild↗](https://github.com/tonyd2wild/deepseek-v4-flash-2x-spark-1m)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+DFlash~45@ 131K—[NVIDIA Developer Forum \(375607\)↗](https://forums.developer.nvidia.com/t/mimo-v2-5-dflash-speculative-decoding-on-a-2x-dgx-spark-pair-22-67-tok-s-depending-on-workload/375607)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+DFlash~45@ 26K~5540@ 0\.256K \(derived\)[NVIDIA Developer Forum \(375923\)↗](https://forums.developer.nvidia.com/t/mimo-v2-5-dflash-and-a-4-bit-nvfp4-kv-cache-in-one-vllm-instance-on-the-v0-24-0-release-2x-dgx-spark/375923)[Single H100 80 GB workstation](https://llmrequirements.com/hardware/h100-80)mistral\-medium\-3\-5\-128b@ FP8 on vLLM~45@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/mistralai/Mistral-Medium-3.5-128B)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM\+MTP~40@ 500K—[GitHub tonyd2wild↗](https://github.com/tonyd2wild/deepseek-v4-flash-dgx-spark)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)mimo\-v2\-5\-pro\-1t\-moe@ NVFP4 on vLLM\+MTP~40@ 1K~1950@ 2K[NVIDIA Developer Forum \(370803\)↗](https://forums.developer.nvidia.com/t/mimo-2-5-pro-nvfp4-on-8xgb10-cluster/370803)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-6\-35b\-a3b\-moe@ Q8\_0 on llama\.cpp~40@ 4K—[Frame\.work Community↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)qwen3\-5\-397b\-a17b\-moe@ FP8 \(406 GiB\) on vLLM~39\.5@ 32K—[NVIDIA Developer Forum 369446 \(vLLM eugr fork TP=8\)↗](https://forums.developer.nvidia.com/t/kimi-2-6-and-qwen-3-5-397b-fp8-on-8xgb10-cluster/369446)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp\+DFlash~38@ 256K—[GitHub phuongncn↗](https://github.com/phuongncn/qwen3.6-27b-speedhack-gx10-dgx-spark)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 4\-bit weights \+ NVFP4 4\-bit KV cache on vLLM \(TP=2 Ray\) \+ DFlash spec\-decode~37\.8@ 1000K—[GitHub tonyd2wild \(MiMo V2\.5 DFlash 1M NVFP4\-KV\)↗](https://github.com/tonyd2wild/MiMo-V2.5-DFlash-1M-ctx-NVFP4-KV-2x-DGX-Spark)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)qwen3\-5\-397b\-a17b\-moe@ NVFP4 on vLLM~37@ 32K—[Level1Techs Forum↗](https://forum.level1techs.com/t/running-qwen3-5-397b-on-4x-dgx-spark/247523)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-MXFP4 \(bf16 KV, EAGLE3 k=2\) on vLLM~34\.8@ 262\.144K~2020@ 262\.144K[NVIDIA Developer Forum 375386 \(vLLM TP=4 EAGLE3\)↗](https://forums.developer.nvidia.com/t/minimax-m3-mxfp4-on-4x-gb10-tp-4-eagle3-262k-context-35-tok-s/375386)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+MTP~34@ 0\.5K~2609@ 2K \(derived\)[NVIDIA Developer Forum \(370459\)↗](https://forums.developer.nvidia.com/t/mimo-v2-5-nvfp4-on-2x-spark-cluster-recipe-findings-fixes-benchmarks/370459)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-AWQ\-INT4 \(fp8 KV, EAGLE3\) on vLLM~33\.7@ 262\.144K—[NVIDIA Developer Forum 375361 \(vLLM TP=4 EAGLE3\)↗](https://forums.developer.nvidia.com/t/minimax-m3-awq-running-tp-4-across-4x-dgx-spark-gb10-33-tok-s-full-recipe-the-gb10-build-fixes/375361)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM~33@ 32K~512@ 128K \(derived\)[NVIDIA Developer Forum \(370309\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-official-fp8-running-across-2x-dgx-spark-tp-2-mtp-200k-ctx-recipe-numbers/370309)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)deepseek\-v3\-671b\-moe@ FP8 \(671B, TP=8\) on vLLM~33@ 1\.024K—[dzhsurf/deepseek\-v3\-r1\-deploy\-and\-benchmarks \(8xH100 vLLM TP=8, concurrency=1\)↗](https://github.com/dzhsurf/deepseek-v3-r1-deploy-and-benchmarks)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)minimax\-m2\-7\-230b\-moe@ UD\-Q3\_K\_S on llama\.cpp \(Vulkan RADV\)~31@ 0\.5K~243@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)mistral\-medium\-3\-5\-128b@ FP8 on vLLM~30@ 98K—[HF mistralai discussion \#17↗](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/discussions/17)[Quad Tesla P40 \(96 GB\) homelab build](https://llmrequirements.com/hardware/tesla-p40-x4)gpt\-oss\-120b@ MXFP4 \(native MoE, ~65GB, fits in 96GB across 4x P40\) on llama\.cpp~28\.1@ 4K—[TinyComputers\.io \(Tesla P40 home lab, gpt\-oss\-120b MXFP4\)↗](https://tinycomputers.io/posts/repurposing-enterprise-gpus-the-tesla-p40-home-lab-story.html)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp\+MTP~28@ 2K~1084@ 2K[NVIDIA Developer Forum \(370298\)↗](https://forums.developer.nvidia.com/t/mtp-llama-cpp-a-look-at-qwen3-6-27b/370298)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)mistral\-small\-4\-119b\-moe@ NVFP4 on vLLM~27\.8@ 262\.144K—[Sebastien67 Medium \(first\-hand DGX Spark vLLM NVFP4 run\)↗](https://medium.com/@Sebastien67/running-mistral-small-4-119b-nvfp4-locally-on-a-dgx-spark-81cc2fdc4f6f)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-REAP Q4\_K\_M \(GGUF, pruned REAP variant\) on llama\.cpp \(Vulkan RADV, RPC cluster\)~26\.7@ 0\.512K~272@ 0\.512K[visorcraft/strix\-halo\-llm\-perf \(2\-node RPC llama\-bench, 2026\-02\-19\)↗](https://github.com/visorcraft/strix-halo-llm-perf/blob/main/results/2026-02-19_tb-rpc-minimax-q4km.md)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 \(modelopt\_fp4, fp8 KV\) on SGLang\+MTP~25\.5@ 196\.608K—[NVIDIA Developer Forum 373676 \(SGLang TP=4 EP=4\)↗](https://forums.developer.nvidia.com/t/minimax-m2-5-nvfp4-on-4x-dgx-spark-via-sglang-tp-4-ep-4-124-tok-s-aggregate-n-8-fixing-the-cutlass-moe-compile-oom-with-max-jobs-1/373676)[Mac Mini M4 \(24 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-24)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~25 est\.@ 4K—[maloyan\.xyz \(M4 16GB, scaled\)↗](https://maloyan.xyz/blog/running-qwen-locally-mac-mini-m4)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-REAP MXFP4\_MOE \(GGUF\) on llama\.cpp \(Vulkan RADV, RPC cluster\)~24\.5@ 0\.512K~299\.5@ 0\.512K[visorcraft/strix\-halo\-llm\-perf \(2\-node RPC llama\-bench, 2026\-02\-19\)↗](https://github.com/visorcraft/strix-halo-llm-perf/blob/main/results/2026-02-19_tb-rpc-minimax-mxfp4.md)[Single AMD Radeon AI Pro R9700 32 GB build](https://llmrequirements.com/hardware/amd-r9700-32)qwen3\-6\-27b\-dense@ Q5\_K\_M on llama\.cpp~24@ 4K~611@ 32K[GitHub truelies444↗](https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)[Dual AMD Radeon AI Pro R9700 build \(64 GB\)](https://llmrequirements.com/hardware/amd-r9700-x2)qwen3\-6\-35b\-a3b\-moe@ base weights \(fp8 KV\) on vLLM~22\.8@ 1\.036K~4600@ 1\.036K[mlai\.blog \(Qwen3\.5\-35B\-A3B on dual R9700, ROCm vLLM\)↗](https://mlai.blog/2026-04-09-qwen35a3b-r9700-rocm-vllm)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)glm\-5\-2\-753b\-moe@ AWQ\-INT4 \(cyankiwi/GLM\-5\.2\-AWQ\-INT4, 15% data\-free expert pruning\) on vLLM\+MTP~22@ 8\.192K~535@ 8\.192K[NVIDIA Developer Forum 374125 \(CosmicRaisins, AWQ\-INT4 TP=4 MTP\)↗](https://forums.developer.nvidia.com/t/glm-5-2-on-a-4x-gb10-cluster-22-tok-s-decode-256k-ctx-recipe/374125)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-5\-122b\-a10b\-moe@ UD\-Q5\_K\_XL on llama\.cpp \(Vulkan RADV\)~22@ 0\.5K~337@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-27b\-dense@ UD\-Q4\_K\_M \(draft\-mtp n=3\) on llama\.cpp \(ROCm, MTP\)~21@ 0\.5K—[Caleb Coffie \- benchmarking llama\.cpp MTP on Strix Halo↗](https://calebcoffie.com/blog/benchmarking-llama-cpp-mtp-on-strix-halo)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)llama\-4\-scout@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~20@ 0\.5K~103@ 0\.5K[hardware\-corner\.net Strix Halo optimization benchmarks↗](https://www.hardware-corner.net/strix-halo-llm-optimization/)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)kimi\-k2\-6\-1t\-moe@ NVFP4 on vLLM~18@ 32K—[NVIDIA Developer Forum↗](https://forums.developer.nvidia.com/t/kimi-2-6-and-qwen-3-5-397b-fp8-on-8xgb10-cluster/369446)[Mac Mini M4 \(16 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-16)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~17@ 4K—[maloyan\.xyz \(M4 16GB\)↗](https://maloyan.xyz/blog/running-qwen-locally-mac-mini-m4)[Tesla V100 32 GB SXM2 mod build](https://llmrequirements.com/hardware/tesla-v100-sxm2-mod)qwen3\-6\-27b\-dense@ Q5\_K\_M on llama\.cpp~17@ 32K—[hardware\-corner\.net \(V100 32GB guide\)↗](https://www.hardware-corner.net/guides/tesla-v100-32gb-for-llm/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-27b\-dense@ Q5\_K\_M on llama\.cpp~16@ 4K—[llm\-tracker\.info \(kyuz0\)↗](https://llm-tracker.info/_TOORG/Strix-Halo)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)deepseek\-v4\-flash\-284b\-moe@ IQ2\_XXS\-w2Q2K imatrix \(~80\.8 GB\) on ds4 \(antirez DeepSeek\-V4\-Flash engine\) \+ MTP, ROCm 7\.2\.4 gfx1151~15\.25@ 2K~152@ 2K \(derived\)[kyuz0 ds4 Strix Halo toolbox \(ds4\-bench, single 128GB node\); antirez ds4 engine↗](https://github.com/kyuz0/strix-halo-ds4-toolbox)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)deepseek\-v4\-flash\-284b\-moe@ Hybrid Q2/Q4 imatrix \(layers 37\-42 Q4, ~97 GB\) on ds4 \(antirez engine\) \+ MTP, ROCm 7\.2\.4 gfx1151~15\.02@ 2K~138@ 2K[kyuz0 ds4 Strix Halo toolbox \(ds4\-bench, single\-node hybrid Q2/Q4\)↗](https://github.com/kyuz0/strix-halo-ds4-toolbox)[MacBook Air M4 \(16 GB\)](https://llmrequirements.com/hardware/mba-m4-16)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~15 est\.@ 4K—[maloyan\.xyz \(M4 16GB, fanless\-adjusted\)↗](https://maloyan.xyz/blog/running-qwen-locally-mac-mini-m4)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)glm\-5\-2\-753b\-moe@ NVFP4 \(REAP\-less, high\-quality 4\-bit; cf\. nvidia/GLM\-5\.2\-NVFP4 model card\) on vLLM~15@ 131\.072K~500@ 131\.072K[NVIDIA Developer Forum 374832 \(REAP\-less NVFP4, custom vLLM fork TP=4\)↗](https://forums.developer.nvidia.com/t/fitting-a-high-quality-reap-less-glm-5-2-onto-4x-dgx-spark/374832)[Quad RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x4)devstral\-2\-123b@ IQ4\_KSS \(GGUF\) on ik\_llama\.cpp \(\-sm graph, tensor\-parallel\)~15 est\.@ 4K~300@ 2K[HF ubergarm Devstral\-2\-123B\-GGUF discussion \#2 \(phakio, ik\_llama\.cpp 4\-GPU\)↗](https://huggingface.co/ubergarm/Devstral-2-123B-Instruct-2512-GGUF/discussions/2)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)mimo\-v2\-5\-310b\-a15b\-moe@ UD\-Q4/Q5\_K\_XL GGUF \(~180\-215 GB, split across 2 nodes\) on llama\.cpp RPC \(2x Strix Halo 128GB, ROCm, USB4net secondary link\)~15@ 10K~356@ 10K \(derived\)[r/LocalLLaMA operator report \(2x Strix Halo 128GB, llama\.cpp RPC over USB4net\); AesSedai/unsloth MiMo\-V2\.5 GGUF↗](https://old.reddit.com/r/LocalLLaMA/comments/1uboiko/rollin_mimo25_on_two_halo_strixeses/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)nemotron\-3\-super\-120b\-a12b\-moe@ UD\-Q4\_K\_XL on llama\.cpp \(ROCm 7\.2\.3\)~14@ 0\.5K~276@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)kimi\-k2\-6\-1t\-moe@ NVFP4 \(60 shards, ~554 GiB, no spec\-decode\) on vLLM~13\.5@ 32\.768K—[NVIDIA Developer Forum 369446 \(vLLM eugr fork TP=8\)↗](https://forums.developer.nvidia.com/t/kimi-2-6-and-qwen-3-5-397b-fp8-on-8xgb10-cluster/369446)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)deepseek\-v4\-flash\-284b\-moe@ Q4 imatrix distributed \(~153\.3 GB, Q4 experts\) on ds4 multi\-node \(pipeline\-parallel, 2x Strix, ROCm 7\.2\.4 gfx1151\) \+ MTP~13\.01@ 2K~62@ 2K \(derived\)[kyuz0 ds4 Strix Halo toolbox \(ds4\-bench, 2\-node distributed Q4\)↗](https://github.com/kyuz0/strix-halo-ds4-toolbox)[Dual RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x2)gpt\-oss\-120b@ Q4\_K\_M \(\-\-n\-cpu\-moe 26, tensor\-split 0\.5/0\.5\) on llama\.cpp \(CUDA, 2x GPU \+ CPU\-MoE\)~13@ 4K—[LLM Garage \- GPT\-OSS\-120B on Dual RTX 3090s↗](https://llmgarage.ai/gpt-oss-120b-dual-3090/)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)glm\-5\-2\-753b\-moe@ low\-bit \(vLLM TP=2\) on vLLM \(TP over 2 Sparks\)~12@ 40K—[NVIDIA Developer Forum 374523 \(GLM\-5\.2 vLLM TP=2 update\)↗](https://forums.developer.nvidia.com/t/academic-glm-5-2-on-2x-dgx-spark-gb10-nodes-crazy-1-bit-ud-iq1-s-rpc-llama-cpp-256k-context-8-tok-s/374523)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gemma\-4\-31b@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~11@ 0\.5K~302@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)glm\-5\-2\-753b\-moe@ UD\-IQ1\_S \(~2\.3 bpw, 1\-bit\) on llama\.cpp RPC \(tensor\-split over 2 nodes\)~8@ 2K~213@ 2K[NVIDIA Developer Forum 374523 \(GLM\-5\.2 on 2x DGX Spark, 1\-bit llama\.cpp RPC\)↗](https://forums.developer.nvidia.com/t/academic-glm-5-2-on-2x-dgx-spark-gb10-nodes-crazy-1-bit-ud-iq1-s-rpc-llama-cpp-256k-context-8-tok-s/374523)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)glm\-5\-2\-753b\-moe@ IQ4\_XS \(GGUF, ~365GB across 4 nodes, DSA sparse attention active\) on llama\.cpp \(RPC multi\-node\)~6\.28@ 1048\.576K~222@ 1048\.576K[NVIDIA Developer Forum 373933 \(IQ4\_XS llama\.cpp RPC, DSA active\)↗](https://forums.developer.nvidia.com/t/glm-5-2-iq4-xs-on-4x-gb10-6-28-tok-s-dsa-active-full-recipe/373933)[RTX 3060 12 GB build](https://llmrequirements.com/hardware/rtx-3060-12)gemma\-4\-12b@ Q4\_K\_M \(GGUF\) on llama\.cpp \(Vulkan\)~5@ 4K—[Hacker News Gemma 4 12B launch thread \(community, Vulkan\-throttled\)↗](https://news.ycombinator.com/item?id=48385906)[8× Strix Halo cluster \(1024 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x8)kimi\-k2\-6\-1t\-moe@ Q5\_K\_M on llama\.cpp~5@ 4K—[Frame\.work Community↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[Quad Tesla P40 \(96 GB\) homelab build](https://llmrequirements.com/hardware/tesla-p40-x4)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp~4@ 32K—[GitHub \- llama\.cpp \#12990 \(P40 FA\)↗](https://github.com/ggml-org/llama.cpp/issues/12990)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp~3@ 4K—[llm\-tracker\.info \(kyuz0\)↗](https://llm-tracker.info/_TOORG/Strix-Halo)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)gpt\-oss\-120b@ MXFP4 \(F16 GGUF, \-\-n\-cpu\-moe offload\) on llama\.cpp \(CUDA \+ CPU\-MoE offload\)~2@ 0\.1K~48@ 0\.1K[hardware\-corner\.net gpt\-oss CPU\-MoE offloading benchmark↗](https://www.hardware-corner.net/gpt-oss-offloading-moe-layers/)[Single AMD Instinct MI50 32 GB \(used\) build](https://llmrequirements.com/hardware/amd-mi50-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M on llama\.cpp——[GitHub llama\.cpp \#19880 \(MI50 enablement\)↗](https://github.com/ggml-org/llama.cpp/issues/19880)[Quad AMD MI50 32 GB \(128 GB\) homelab build](https://llmrequirements.com/hardware/amd-mi50-x4)qwen3\-6\-35b\-a3b\-moe@ Q5\_K\_M on llama\.cpp——[aibytes\.blog \(MI50 ROCm vs Vulkan\)↗](https://aibytes.blog/comparisons/rocm-7-vs-vulkan-on-mi50-4-model-benchmark-results)[Quad AMD MI50 32 GB \(128 GB\) homelab build](https://llmrequirements.com/hardware/amd-mi50-x4)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp——[HF \- unsloth \(GGUF discussion\)↗](https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF/discussions/5)[Mac Mini M4 \(16 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-16)gemma\-4\-12b@ MLX 4\-bit on Ollama \(MLX\) / mlx\-lm——[Ollama model library \(Apple\-Silicon MLX build; runnable recipe\)\. Fit corroborated by Gemma 4 launch HN thread \(Q4\_K\_M ~6\.6 GB\)\.↗](https://ollama.com/library/gemma4:12b)[Mac Mini M4 \(24 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-24)gemma\-4\-12b@ MLX 4\-bit on Ollama \(MLX\) / mlx\-lm——[Ollama model library \(Apple\-Silicon MLX build\)\. Fit corroborated by HN launch thread\.↗](https://ollama.com/library/gemma4:12b)[MacBook Air M4 \(16 GB\)](https://llmrequirements.com/hardware/mba-m4-16)gemma\-4\-12b@ MLX 4\-bit on Ollama \(MLX\) / mlx\-lm——[Ollama model library \(Apple\-Silicon MLX build\)\. Fit corroborated by HN launch thread\.↗](https://ollama.com/library/gemma4:12b)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on SGLang\+MTP——[GitHub r0b0tlab↗](https://github.com/r0b0tlab/qwen36-35b-a3b-nvfp4-gb10-native-mtp)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ NVFP4\-KV \(nvfp4\_ds\_mla\) on vLLM\+DSpark——[HF drowzeys↗](https://huggingface.co/drowzeys/DeepSeek-V4-Flash-DSpark-NVFP4-KV-1.5M-CTX-2xDGX-Spark)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM \(Ray TP=2\)——[Medium \- Michael Peres↗](https://medium.com/@michaelperes1/turning-two-dgx-sparks-into-a-local-llm-cluster-with-vllm-ray-and-qwen3-6-7eb2a6e04ade)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM——[NVIDIA Developer Forum↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-official-fp8-running-across-2x-dgx-spark-tp-2-mtp-200k-ctx-recipe-numbers/370309)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \(FP8 KV, MTP\) on vLLM \(TP=4\)——[NVIDIA\-NeMo Nemotron Spark Deployment Guide \(4x DGX Spark, NVFP4, vLLM TP=4\)↗](https://github.com/NVIDIA-NeMo/Nemotron/blob/main/usage-cookbook/Nemotron-3-Ultra/SparkDeploymentGuide/README.md)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)deepseek\-v4\-pro\-1t6\-moe@ FP8 on vLLM——[GitHub \- vLLM \#43367↗](https://github.com/vllm-project/vllm/issues/43367)[Single Intel Arc B580 12 GB build](https://llmrequirements.com/hardware/intel-arc-b580-12)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp——[GitHub \- intel/llm\-scaler \(official\)↗](https://github.com/intel/llm-scaler)[Quad RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x4)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp——[HF \- bartowski \(GGUF\)↗](https://huggingface.co/bartowski/mistralai_Mistral-Medium-3.5-128B-GGUF)[Quad RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x4)qwen3\-6\-35b\-a3b\-moe@ FP16 on vLLM——[GitHub \- tfriedel \(RTX 3090 lab\)↗](https://github.com/tfriedel/qwen3.6-rtx3090-lab)[Single RTX 4090 build](https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-6\-27b\-dense@ AWQ on vLLM——[GitHub \- thc1006 \(Ampere/Ada spec\-decode\)↗](https://github.com/thc1006/qwen3.6-speculative-decoding-rtx3090)[Single RTX 4090 build](https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M on llama\.cpp——[GitHub \- thc1006 \(spec\-decode\)↗](https://github.com/thc1006/qwen3.6-speculative-decoding-rtx3090)[RTX 3060 12 GB build](https://llmrequirements.com/hardware/rtx-3060-12)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp——[GitHub \- llama\.cpp build docs↗](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md)[Dual RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-x2)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on vLLM——[HF \- RedHatAI \(NVFP4\)↗](https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4)[Dual RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-x2)qwen3\-6\-27b\-dense@ FP8 on vLLM——[HF \- Qwen \(official FP8\)↗](https://huggingface.co/Qwen/Qwen3.6-27B-FP8)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)mistral\-medium\-3\-5\-128b@ NVFP4 on vLLM——[HF nvidia \(NVFP4 card\)↗](https://huggingface.co/nvidia/Mistral-Medium-3.5-128B-NVFP4)[Single Tesla P40 24 GB \(used\) build](https://llmrequirements.com/hardware/tesla-p40-24)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp——[GitHub llama\.cpp \#19248 \(P40\)↗](https://github.com/ggml-org/llama.cpp/discussions/19248)[Quad Tesla P40 \(96 GB\) homelab build](https://llmrequirements.com/hardware/tesla-p40-x4)qwen3\-6\-35b\-a3b\-moe@ Q5\_K\_M on llama\.cpp——[GitHub \- llama\.cpp \#12990 \(P40 FA\)↗](https://github.com/ggml-org/llama.cpp/issues/12990)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)mimo\-v2\-5\-310b\-a15b\-moe@ UD\-IQ2\_M \(~2\.7 bpw, ~92\.8 GB\) on llama\.cpp \(Vulkan/RADV, kyuz0 container\), gfx1151—~31@ 0\.5K \(derived\)[hogeheer499 strix\-halo\-guide community evidence map \(Corsair AI WS 300, IQ2\_M, capacity row\); bartowski MiMo GGUF↗](https://huggingface.co/bartowski/MiMo-V2.5-GGUF)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gemma\-4\-12b@ UD\-Q4\_K\_XL \(GGUF\) on llama\.cpp \(ROCm 7\.2\.x, gfx1151\)——[kyuz0 Strix Halo toolboxes \(ROCm gfx1151 llama\.cpp recipe; 12B not yet in grid\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)qwen3\-5\-397b\-a17b\-moe@ Q\-family GGUF \(HIP\+RPC, np2, ctx 200k\) on llama\.cpp——[visorcraft/strix\-halo\-llm\-perf \(Qwen3\.5\-397B RPC shape\-control, 2026\-02\-21\)↗](https://github.com/visorcraft/strix-halo-llm-perf/blob/main/results/2026-02-21_qwen397b-rpc-shape-control.md)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)deepseek\-v4\-flash\-284b\-moe@ Q4\_K\-class across 4 nodes on llama\.cpp RPC \(4x Framework Desktop / Strix mainboards\)——[frame\.work llama\.cpp RPC multi\-node recipe \(extended to 4 nodes\)↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)mimo\-v2\-5\-310b\-a15b\-moe@ Q4\_K\_M \(~178 GB\) / Q5\_K\_M \(~213 GB\) on llama\.cpp RPC \(4x Framework Desktop / Strix mainboards\)——[bartowski MiMo\-V2\.5 GGUF \(Q4\_K\_M/Q5\) \+ frame\.work llama\.cpp RPC 4\-node↗](https://huggingface.co/bartowski/MiMo-V2.5-GGUF)[MacBook Pro M5 Pro 48 GB](https://llmrequirements.com/hardware/mbp-m5-pro-48)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM——[GitHub ml\-explore/mlx\-lm↗](https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md)[MacBook Pro M5 Pro 48 GB](https://llmrequirements.com/hardware/mbp-m5-pro-48)gemma\-4\-12b@ MLX 4\-bit on Ollama 0\.31 \(MLX\) \+ MTP——[Ollama blog \(framework\-author first\-party MTP recipe; M5\-family\)\. Directional only; M5 Pro not separately measured\.↗](https://ollama.com/blog/faster-gemma-4-mlx-mtp)[MacBook Pro M5 Pro 48 GB](https://llmrequirements.com/hardware/mbp-m5-pro-48)qwen3\-6\-27b\-dense@ MLX 4\-bit on mlx\-lm / Ollama \(MLX\)——[Ollama model library \(Apple\-Silicon MLX build\)\. Fit from db\.json sizeQ4=16 vs 40 GB usable\.↗](https://ollama.com/library/qwen3.6)[DGX H200 — 8× H200 server \(1\.13 TB HBM3e\)](https://llmrequirements.com/hardware/h200-x8)kimi\-k2\-7\-code\-1t\-moe@ native INT4 on vLLM \(TP=8, expert\-parallel\)——[vLLM Recipes \(official K2\.7 Code command, 8xH200 INT4\); tok/s = K2\.6 same\-box SGLang INT4 \(paxsaroffcuts\)↗](https://recipes.vllm.ai/moonshotai/Kimi-K2.7-Code)

Similar Articles

RTX Pro 4500 Blackwell - Qwen 3.6 27B?

Reddit r/LocalLLaMA

A developer shares local inference benchmarks and systemd configurations for running the Qwen3.6-27B model on an NVIDIA RTX Pro 4500 Blackwell GPU using llama.cpp. The post requests optimization tips for throughput and explores potential use cases for larger models.

RTX Pro 4500 Blackwell Performance Numbers

Reddit r/LocalLLaMA

A user shares performance benchmarks comparing the Nvidia RTX Pro 4500 Blackwell 32GB GPU against the RTX 5060 Ti 16GB for AI inference, showing 1.6-6x speed improvements depending on model size and quantization.

4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s

Reddit r/LocalLLaMA

A user shares benchmark results for running large language models (Qwen 27B-122B) on a dual-GPU setup with RTX 4090 and RTX 5060 Ti, achieving high token generation speeds (e.g., 206 t/s on 35B-A3B, 37-41 t/s on 122B). The post includes setup details and a link to a GitHub repo with scripts and raw data.