Cached at:
07/11/26, 11:40 PM
# Recipes — LLMRequirements.com
Source: [https://llmrequirements.com/recipes](https://llmrequirements.com/recipes)
[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)kimi\-k2\-5\-1t\-moe@ INT4 \(FP8 KV, DCP=8\) on vLLMbatch~900@ 40K\(100\-conc\. aggregate\)—[local\-inference\-lab/rtx6kpro wiki \(Kimi K2\.5 high concurrency, Festr\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ FP8 \(fp8 KV, MTP spec=3\) on vLLMbatch~894@ 175K\(32\-conc\. aggregate\)—[theogravity/dual\-rtx\-6000\-blackwell\-qwen3\.6\-27b\-fp8 \(coding sweep, seqs=32\)↗](https://github.com/theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)deepseek\-v3\-671b\-moe@ FP8 \(TP=8\) on vLLMbatch~620@ 1\.024K\(100\-conc\. aggregate\)—[dzhsurf/deepseek\-v3\-r1\-deploy\-and\-benchmarks \(8xH100 vLLM TP=8, ~100 concurrency\)↗](https://github.com/dzhsurf/deepseek-v3-r1-deploy-and-benchmarks)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)gemma4\-26b\-moe@ native on vLLM cluster TP=2 \(triton, ROCm\)batch~411@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(triton cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/triton/google_gemma-4-26B-A4B-it_cluster_tp2_throughput.json)[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)qwen3\-5\-397b\-a17b\-moe@ NVFP4 on SGLang\+MTP~350@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Qwen3\.5\-397B 8x single\-batch\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~320@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-6\-35b\-a3b\-moe@ AWQ\-4bit / native on vLLM cluster TP=2 \(aiter, ROCm\)batch~287@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)gpt\-oss\-120b@ MXFP4 on vLLM cluster TP=2 \(triton, ROCm\)batch~229@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(triton cluster tp2 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[Single RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-32)qwen3\-coder\-30b@ Q4\_K on llama\.cpp \(CUDA\)~226@ 4K~7093@ 4K[hardware\-corner\.net RTX 5090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-5090/)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)mistral\-medium\-3\-5\-128b@ FP8 on vLLM~220@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/mistralai/Mistral-Medium-3.5-128B)[Single RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-32)gemma4\-26b\-moe@ Q4\_K on llama\.cpp \(CUDA\)~180@ 4K~8799@ 4K[hardware\-corner\.net RTX 5090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-5090/)[Single H100 80 GB workstation](https://llmrequirements.com/hardware/h100-80)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~180@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 \(FP8 KV, MTP n=2\) on vLLMbatch~179\.9@ 393\.216K\(8\-conc\. aggregate\)—[NVIDIA Developer Forum 373808 \(jasl vLLM TP=4, n=8 aggregate\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-on-4x-dgx-spark-via-vllm-jasl-fork-tp-4-rdma-mtp-49-54-tok-s-single-stream-full-recipe-the-traps/373808)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ AWQ\-4bit / native on vLLM \(aiter, ROCm\)batch~178@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter tp1 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~170@ 262K—[GitHub lastloop\-ai↗](https://github.com/lastloop-ai/vllm-blackwell-guide)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ NVFP4 on vLLM\+MTP~156@ 262K~831@ 262K[loFT LLC↗](https://loftllc.dev/en/docs/tech/llm-research/qwen3-6-27b-nvfp4-mtp-vllm-benchmark/)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-coder\-30b@ Q4\_K on llama\.cpp \(CUDA\)~153@ 4K~2988@ 4K[hardware\-corner\.net RTX 3090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/)[Quad RTX Pro 6000 Blackwell build \(384 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)qwen3\-5\-397b\-a17b\-moe@ AWQ\-INT4 \(QuantTrio\) on SGLang\+MTP~152@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Qwen3\.5\-397B single\-batch decode\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[Dual RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x2)qwen3\-6\-35b\-a3b\-moe@ AWQ on vLLM~149@ 4K—[GitHub \- tfriedel \(RTX 3090 lab\)↗](https://github.com/tfriedel/qwen3.6-rtx3090-lab)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ ROCmFP4 \(CHADROCK\) on llama\-server\+ROCmFPX~140@ 4K—[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ FP8 \(native, fp8 KV, MTP spec=3\) on vLLM~137@ 175K—[theogravity/dual\-rtx\-6000\-blackwell\-qwen3\.6\-27b\-fp8 \(benchmark sweep\)↗](https://github.com/theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)mistral\-small\-4\-119b\-moe@ NVFP4 on vLLMbatch~131@ 262\.144K\(20\-conc\. aggregate\)—[Sebastien67 Medium \(DGX Spark vLLM NVFP4, n=20 aggregate\)↗](https://medium.com/@Sebastien67/running-mistral-small-4-119b-nvfp4-locally-on-a-dgx-spark-81cc2fdc4f6f)[Quad RTX Pro 6000 Blackwell build \(384 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)qwen3\-5\-397b\-a17b\-moe@ NVFP4 \(nvidia checkpoint\) on vLLM\+MTP~130@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Qwen3\.5\-397B MTP scaling table, concurrency=1\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)gemma\-4\-31b@ native \(bf16/fp16\) on vLLM cluster TP=2 \(triton, ROCm\)batch~128@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(triton cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/triton/google_gemma-4-31B-it_cluster_tp2_throughput.json)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 \(modelopt\_fp4, fp8 KV\) on SGLang\+MTPbatch~124@ 196\.608K\(8\-conc\. aggregate\)—[NVIDIA Developer Forum 373676 \(SGLang TP=4 EP=4, n=8 aggregate\)↗](https://forums.developer.nvidia.com/t/minimax-m2-5-nvfp4-on-4x-dgx-spark-via-sglang-tp-4-ep-4-124-tok-s-aggregate-n-8-fixing-the-cutlass-moe-compile-oom-with-max-jobs-1/373676)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-next\-80b\-moe@ Q4\_K\_M on ollama / llama\.cpp \(CUDA\)~124@ 4K~3274@ 4K[vaditaslim\.com RTX PRO 6000 Blackwell 8\-model benchmarks↗](https://www.vaditaslim.com/blog/ai/local-llm-benchmarks-rtx-pro-6000)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)gemma4\-26b\-moe@ Q4\_K on llama\.cpp \(CUDA\)~119@ 4K~3625@ 4K[hardware\-corner\.net RTX 3090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/)[Single H100 80 GB workstation](https://llmrequirements.com/hardware/h100-80)qwen3\-6\-27b\-dense@ FP16 on vLLM~110@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/Qwen/Qwen3.6-27B)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-5\-122b\-a10b\-moe@ cyankiwi AWQ\-4bit on vLLM cluster TP=2 \(aiter, ROCm\)batch~104@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, Dec 2025\)↗](https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_XL on llama\.cpp~101@ 65K~1171@ 0\.5K[aminrj\.com \(Qwen3\.6 on 24GB\)↗](https://aminrj.com/posts/llamacpp-qwen36-35b/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ IQ4\_XS\-Q8nextn on llama\-server\+MTP~101@ 4K—[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)kimi\-k2\-5\-1t\-moe@ INT4 \(BF16 KV, EP=8, overclocked GDDR7\) on SGLang\+MTP~101@ 4K—[local\-inference\-lab/rtx6kpro wiki \(Kimi K2\.5 8x single\-batch decode\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-6\-27b\-dense@ INT4 \(AutoRound\) \+ MTP n=3, FP8 KV on vLLM \(flashinfer, MTP\)~100@ 262K—[GitHub lastloop\-ai↗](https://github.com/lastloop-ai/vllm-blackwell-guide)[8× RTX Pro 6000 Blackwell server \(768 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)glm\-51\-754b\-moe@ NVFP4\-MTP \(lukealonso/GLM\-5\.1\-NVFP4\-MTP, served as GLM\-5\) on SGLang\+MTP~100@ 4K—[local\-inference\-lab/rtx6kpro wiki \(GLM\-5 single\-batch decode; models/glm5\.md = GLM\-5\.1\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on vLLM\+DFlash~97@ 0\.5K~9090@ 0\.5K \(derived\)[GitHub AEON\-7↗](https://github.com/AEON-7/Qwen3.6-NVFP4-DFlash)[MacBook Pro M5 Max 64 GB](https://llmrequirements.com/hardware/mbp-m5-max-64)gemma\-4\-12b@ MLX NVFP4 on Ollama 0\.31 \(MLX\) \+ MTP~95@ 4K—[Ollama blog \(framework\-author first\-party; M5 Max, Aider polyglot\)↗](https://ollama.com/blog/faster-gemma-4-mlx-mtp)[Single RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-32)qwen3\-6\-27b\-dense@ NVFP4 on vLLM~92@ 200K~5300@ 47K[GitHub devnen↗](https://github.com/devnen/qwen3.6-windows-server)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on vLLM~90@ 43K~2133@ 32K \(derived\)[GitHub technigmaai/dgx\-spark↗](https://github.com/technigmaai/dgx-spark/tree/main/spark-vllm-docker/nvidia-Qwen3.6-35B-A3B-NVFP4)[Dual RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x2)qwen3\-6\-27b\-dense@ AWQ on vLLM~90@ 100K—[GitHub \- tfriedel \(RTX 3090 lab\)↗](https://github.com/tfriedel/qwen3.6-rtx3090-lab)[MacBook Pro M5 Max 64 GB](https://llmrequirements.com/hardware/mbp-m5-max-64)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~87@ 4K~2447@ 4K[oMLX Benchmark↗](https://omlx.ai/benchmarks/oykgm8sq)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 on vLLM~85@ 4K—[local\-inference\-lab/rtx6kpro wiki \(MiniMax\-M2\.5 single\-stream table\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ Q4\_0 on llama\.cpp~81@ 4K~1244@ 0\.5K[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[Quad RTX Pro 6000 Blackwell build \(384 GB\)](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-FP8 on vLLM~81@ 20K—[local\-inference\-lab/rtx6kpro wiki \(MiniMax\-M2\.5 single\-stream table\)↗](https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)[DGX B200 — 8× B200 server \(1\.44 TB HBM3e\)](https://llmrequirements.com/hardware/b200-x8)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \+ FP8 KV on Dynamo \+ vLLM \(TP=4, expert\-parallel, MTP\)batch~80\.6@ 4K\(20\-conc\. aggregate\)—[NVIDIA ai\-dynamo/dynamo recipes \(B200 TP4\+EP, NVFP4\+FP8, MTP\)↗](https://github.com/ai-dynamo/dynamo/tree/main/recipes/nemotron-3-ultra)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-AWQ\-INT4 \(fp8 KV, EAGLE3\) on vLLMbatch~79@ 262\.144K\(4\-conc\. aggregate\)—[NVIDIA Developer Forum 375361 \(vLLM TP=4, n=4 aggregate\)↗](https://forums.developer.nvidia.com/t/minimax-m3-awq-running-tp-4-across-4x-dgx-spark-gb10-33-tok-s-full-recipe-the-gb10-build-fixes/375361)[Single AMD Radeon AI Pro R9700 32 GB build](https://llmrequirements.com/hardware/amd-r9700-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M on llama\.cpp~77@ 4K~1636@ 32K[GitHub truelies444↗](https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)kimi\-k2\-6\-1t\-moe@ FP8 on vLLM~75@ 256K—[HF \- RedHatAI \(Kimi\-K2\.6\-FP8\-BLOCK\)↗](https://huggingface.co/RedHatAI/Kimi-K2.6-FP8-BLOCK)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ MTP\-GGUF UD\-Q4\_K\_XL \(draft\-mtp n=3\) on llama\.cpp \(Vulkan RADV, MTP\)~75@ 0\.5K—[kyuz0 amd\-strix\-halo\-toolboxes MTP grid \(results\-mtp/summary\.json, 15 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/mtp.html)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-5\-122b\-a10b\-moe@ cyankiwi AWQ\-8bit on vLLM cluster TP=2 \(aiter, ROCm\)batch~74@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/aiter/cyankiwi_Qwen3.5-122B-A10B-AWQ-8bit_cluster_tp2_throughput.json)[Dual AMD Radeon AI Pro R9700 build \(64 GB\)](https://llmrequirements.com/hardware/amd-r9700-x2)qwen3\-6\-35b\-a3b\-moe@ Q6\_K on llama\.cpp~72@ 4K~3038@ 32K[GitHub truelies444↗](https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-6\-27b\-dense@ AWQ/AutoRound\-INT4 on vLLM\+MTP~72@ 32K—[GitHub devnen↗](https://github.com/devnen/qwen3.6-windows-server)[Single RTX 4090 build](https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-coder\-30b@ Q4\_K on llama\.cpp \(CUDA\)~68@ 64K~1502@ 64K[hardware\-corner\.net RTX 4090 LLM benchmarks \(GGUF Q4\)↗](https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-4090/)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ not stated \(native precision\) on vLLM \(Aidendle94/B12X\-MoE, TP=2 RoCE\) \+ DSpark spec\-decode~65@ 200K—[GitHub 0rand \(DeepSeek\-V4 DSpark serving stack\)↗](https://github.com/0rand/DeepSeek-v4-DSpark-Aidendle94-GB10-ServingStack)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ NVFP4\-KV \(nvfp4\_ds\_mla\) on vLLM\+DSpark~63@ 200K—[NVIDIA Developer Forum \(374846\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-dspark-on-2x-dgx-spark-gb10-big-single-stream-speed-boost-60-67-tok-s-1m-context-now-with-concurrency/374846)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M \(UD\) on llama\.cpp~62@ 4K~1059@ 0\.5K[GitHub hogeheer499\-commits/strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[8× Strix Halo cluster \(1024 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x8)qwen3\-6\-35b\-a3b\-moe@ Q8\_0 on llama\.cpp~62@ 4K—[GitHub \- strix\-halo\-guide↗](https://github.com/hogeheer499-commits/strix-halo-guide)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM~60@ 32K~6520@ 8K \(derived\)[NVIDIA Developer Forum \(366822\)↗](https://forums.developer.nvidia.com/t/qwen-qwen3-6-35b-a3b-and-fp8-has-landed/366822)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~60@ 0\.5K~1114@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[DGX H200 — 8× H200 server \(1\.13 TB HBM3e\)](https://llmrequirements.com/hardware/h200-x8)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \+ FP8 KV on Dynamo \+ vLLM \(TP=8, expert\-parallel, MTP\)batch~58\.7@ 4K\(10\-conc\. aggregate\)—[NVIDIA ai\-dynamo/dynamo recipes \(8xH200 TP8\+EP, NVFP4\+FP8, MTP\)↗](https://github.com/ai-dynamo/dynamo/tree/main/recipes/nemotron-3-ultra)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)minimax\-m2\-7\-230b\-moe@ cyankiwi AWQ\-4bit on vLLM cluster TP=2 \(aiter, ROCm\)batch~57@ 4K\(200\-conc\. aggregate\)—[kyuz0 amd\-strix\-halo\-vllm\-toolboxes \(aiter cluster tp2 throughput, 200 reqs\)↗](https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/aiter/cyankiwi_MiniMax-M2.7-AWQ-4bit_cluster_tp2_throughput.json)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gpt\-oss\-120b@ MXFP4 on llama\.cpp \(Vulkan RADV\)~56@ 0\.5K~720@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[Single Intel Arc Pro B70 build](https://llmrequirements.com/hardware/intel-arc-pro-b70-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M \(UD\) on llama\.cpp \(SYCL\)~55@ 4K~615@ 0\.5K[GitHub PMZFX↗](https://github.com/PMZFX/intel-arc-pro-b70-benchmarks/blob/master/llm-benchmarks.md)[Tesla V100 32 GB SXM2 mod build](https://llmrequirements.com/hardware/tesla-v100-sxm2-mod)qwen3\-6\-35b\-a3b\-moe@ Q5\_K\_M on llama\.cpp~55@ 10K—[GitHub \- ai\-bond \(V100 flash\-attn\)↗](https://github.com/ai-bond/flash-attention-v100)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+DFlash~54@ 0\.5K~2083@ 50K \(derived\)[GitHub HeNryous \(renek\)↗](https://github.com/HeNryous/mimo-v25-dflash-dgx-spark)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gemma4\-26b\-moe@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~54@ 0\.5K~1324@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[MacBook Pro M5 Max 64 GB](https://llmrequirements.com/hardware/mbp-m5-max-64)gemma\-4\-12b@ MLX NVFP4 on Ollama 0\.31 \(MLX\), no MTP~50\.2@ 4K—[Ollama blog \(framework\-author first\-party; M5 Max\)↗](https://ollama.com/blog/faster-gemma-4-mlx-mtp)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM\+DFlash~50@ 262K~4932@ 0\.5K[GitHub ZengboJamesWang↗](https://github.com/ZengboJamesWang/dgx-spark-vllm-qwen3.6-35b-a3b-dflash)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)qwen3\-6\-35b\-a3b\-moe@ Q8\_0 on llama\.cpp~50@ 4K—[Frame\.work Community↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 \(official weights, FP8 KV, MTP n=2\) on vLLM~49\.4@ 393\.216K—[NVIDIA Developer Forum 373808 \(jasl vLLM TP=4\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-on-4x-dgx-spark-via-vllm-jasl-fork-tp-4-rdma-mtp-49-54-tok-s-single-stream-full-recipe-the-traps/373808)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM\+MTP~45\.5@ 1000K~786@ 800K[GitHub tonyd2wild↗](https://github.com/tonyd2wild/deepseek-v4-flash-2x-spark-1m)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+DFlash~45@ 131K—[NVIDIA Developer Forum \(375607\)↗](https://forums.developer.nvidia.com/t/mimo-v2-5-dflash-speculative-decoding-on-a-2x-dgx-spark-pair-22-67-tok-s-depending-on-workload/375607)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+DFlash~45@ 26K~5540@ 0\.256K \(derived\)[NVIDIA Developer Forum \(375923\)↗](https://forums.developer.nvidia.com/t/mimo-v2-5-dflash-and-a-4-bit-nvfp4-kv-cache-in-one-vllm-instance-on-the-v0-24-0-release-2x-dgx-spark/375923)[Single H100 80 GB workstation](https://llmrequirements.com/hardware/h100-80)mistral\-medium\-3\-5\-128b@ FP8 on vLLM~45@ 128K—[vLLM Recipes↗](https://recipes.vllm.ai/mistralai/Mistral-Medium-3.5-128B)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM\+MTP~40@ 500K—[GitHub tonyd2wild↗](https://github.com/tonyd2wild/deepseek-v4-flash-dgx-spark)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)mimo\-v2\-5\-pro\-1t\-moe@ NVFP4 on vLLM\+MTP~40@ 1K~1950@ 2K[NVIDIA Developer Forum \(370803\)↗](https://forums.developer.nvidia.com/t/mimo-2-5-pro-nvfp4-on-8xgb10-cluster/370803)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-6\-35b\-a3b\-moe@ Q8\_0 on llama\.cpp~40@ 4K—[Frame\.work Community↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)qwen3\-5\-397b\-a17b\-moe@ FP8 \(406 GiB\) on vLLM~39\.5@ 32K—[NVIDIA Developer Forum 369446 \(vLLM eugr fork TP=8\)↗](https://forums.developer.nvidia.com/t/kimi-2-6-and-qwen-3-5-397b-fp8-on-8xgb10-cluster/369446)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp\+DFlash~38@ 256K—[GitHub phuongncn↗](https://github.com/phuongncn/qwen3.6-27b-speedhack-gx10-dgx-spark)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 4\-bit weights \+ NVFP4 4\-bit KV cache on vLLM \(TP=2 Ray\) \+ DFlash spec\-decode~37\.8@ 1000K—[GitHub tonyd2wild \(MiMo V2\.5 DFlash 1M NVFP4\-KV\)↗](https://github.com/tonyd2wild/MiMo-V2.5-DFlash-1M-ctx-NVFP4-KV-2x-DGX-Spark)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)qwen3\-5\-397b\-a17b\-moe@ NVFP4 on vLLM~37@ 32K—[Level1Techs Forum↗](https://forum.level1techs.com/t/running-qwen3-5-397b-on-4x-dgx-spark/247523)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-MXFP4 \(bf16 KV, EAGLE3 k=2\) on vLLM~34\.8@ 262\.144K~2020@ 262\.144K[NVIDIA Developer Forum 375386 \(vLLM TP=4 EAGLE3\)↗](https://forums.developer.nvidia.com/t/minimax-m3-mxfp4-on-4x-gb10-tp-4-eagle3-262k-context-35-tok-s/375386)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)mimo\-v2\-5\-310b\-a15b\-moe@ NVFP4 on vLLM\+MTP~34@ 0\.5K~2609@ 2K \(derived\)[NVIDIA Developer Forum \(370459\)↗](https://forums.developer.nvidia.com/t/mimo-v2-5-nvfp4-on-2x-spark-cluster-recipe-findings-fixes-benchmarks/370459)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-AWQ\-INT4 \(fp8 KV, EAGLE3\) on vLLM~33\.7@ 262\.144K—[NVIDIA Developer Forum 375361 \(vLLM TP=4 EAGLE3\)↗](https://forums.developer.nvidia.com/t/minimax-m3-awq-running-tp-4-across-4x-dgx-spark-gb10-33-tok-s-full-recipe-the-gb10-build-fixes/375361)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM~33@ 32K~512@ 128K \(derived\)[NVIDIA Developer Forum \(370309\)↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-official-fp8-running-across-2x-dgx-spark-tp-2-mtp-200k-ctx-recipe-numbers/370309)[8× H100 80 GB server](https://llmrequirements.com/hardware/h100-x8)deepseek\-v3\-671b\-moe@ FP8 \(671B, TP=8\) on vLLM~33@ 1\.024K—[dzhsurf/deepseek\-v3\-r1\-deploy\-and\-benchmarks \(8xH100 vLLM TP=8, concurrency=1\)↗](https://github.com/dzhsurf/deepseek-v3-r1-deploy-and-benchmarks)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)minimax\-m2\-7\-230b\-moe@ UD\-Q3\_K\_S on llama\.cpp \(Vulkan RADV\)~31@ 0\.5K~243@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[Dual RTX Pro 6000 Blackwell build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)mistral\-medium\-3\-5\-128b@ FP8 on vLLM~30@ 98K—[HF mistralai discussion \#17↗](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/discussions/17)[Quad Tesla P40 \(96 GB\) homelab build](https://llmrequirements.com/hardware/tesla-p40-x4)gpt\-oss\-120b@ MXFP4 \(native MoE, ~65GB, fits in 96GB across 4x P40\) on llama\.cpp~28\.1@ 4K—[TinyComputers\.io \(Tesla P40 home lab, gpt\-oss\-120b MXFP4\)↗](https://tinycomputers.io/posts/repurposing-enterprise-gpus-the-tesla-p40-home-lab-story.html)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp\+MTP~28@ 2K~1084@ 2K[NVIDIA Developer Forum \(370298\)↗](https://forums.developer.nvidia.com/t/mtp-llama-cpp-a-look-at-qwen3-6-27b/370298)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)mistral\-small\-4\-119b\-moe@ NVFP4 on vLLM~27\.8@ 262\.144K—[Sebastien67 Medium \(first\-hand DGX Spark vLLM NVFP4 run\)↗](https://medium.com/@Sebastien67/running-mistral-small-4-119b-nvfp4-locally-on-a-dgx-spark-81cc2fdc4f6f)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-REAP Q4\_K\_M \(GGUF, pruned REAP variant\) on llama\.cpp \(Vulkan RADV, RPC cluster\)~26\.7@ 0\.512K~272@ 0\.512K[visorcraft/strix\-halo\-llm\-perf \(2\-node RPC llama\-bench, 2026\-02\-19\)↗](https://github.com/visorcraft/strix-halo-llm-perf/blob/main/results/2026-02-19_tb-rpc-minimax-q4km.md)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 \(modelopt\_fp4, fp8 KV\) on SGLang\+MTP~25\.5@ 196\.608K—[NVIDIA Developer Forum 373676 \(SGLang TP=4 EP=4\)↗](https://forums.developer.nvidia.com/t/minimax-m2-5-nvfp4-on-4x-dgx-spark-via-sglang-tp-4-ep-4-124-tok-s-aggregate-n-8-fixing-the-cutlass-moe-compile-oom-with-max-jobs-1/373676)[Mac Mini M4 \(24 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-24)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~25 est\.@ 4K—[maloyan\.xyz \(M4 16GB, scaled\)↗](https://maloyan.xyz/blog/running-qwen-locally-mac-mini-m4)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-REAP MXFP4\_MOE \(GGUF\) on llama\.cpp \(Vulkan RADV, RPC cluster\)~24\.5@ 0\.512K~299\.5@ 0\.512K[visorcraft/strix\-halo\-llm\-perf \(2\-node RPC llama\-bench, 2026\-02\-19\)↗](https://github.com/visorcraft/strix-halo-llm-perf/blob/main/results/2026-02-19_tb-rpc-minimax-mxfp4.md)[Single AMD Radeon AI Pro R9700 32 GB build](https://llmrequirements.com/hardware/amd-r9700-32)qwen3\-6\-27b\-dense@ Q5\_K\_M on llama\.cpp~24@ 4K~611@ 32K[GitHub truelies444↗](https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)[Dual AMD Radeon AI Pro R9700 build \(64 GB\)](https://llmrequirements.com/hardware/amd-r9700-x2)qwen3\-6\-35b\-a3b\-moe@ base weights \(fp8 KV\) on vLLM~22\.8@ 1\.036K~4600@ 1\.036K[mlai\.blog \(Qwen3\.5\-35B\-A3B on dual R9700, ROCm vLLM\)↗](https://mlai.blog/2026-04-09-qwen35a3b-r9700-rocm-vllm)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)glm\-5\-2\-753b\-moe@ AWQ\-INT4 \(cyankiwi/GLM\-5\.2\-AWQ\-INT4, 15% data\-free expert pruning\) on vLLM\+MTP~22@ 8\.192K~535@ 8\.192K[NVIDIA Developer Forum 374125 \(CosmicRaisins, AWQ\-INT4 TP=4 MTP\)↗](https://forums.developer.nvidia.com/t/glm-5-2-on-a-4x-gb10-cluster-22-tok-s-decode-256k-ctx-recipe/374125)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-5\-122b\-a10b\-moe@ UD\-Q5\_K\_XL on llama\.cpp \(Vulkan RADV\)~22@ 0\.5K~337@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-27b\-dense@ UD\-Q4\_K\_M \(draft\-mtp n=3\) on llama\.cpp \(ROCm, MTP\)~21@ 0\.5K—[Caleb Coffie \- benchmarking llama\.cpp MTP on Strix Halo↗](https://calebcoffie.com/blog/benchmarking-llama-cpp-mtp-on-strix-halo)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)llama\-4\-scout@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~20@ 0\.5K~103@ 0\.5K[hardware\-corner\.net Strix Halo optimization benchmarks↗](https://www.hardware-corner.net/strix-halo-llm-optimization/)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)kimi\-k2\-6\-1t\-moe@ NVFP4 on vLLM~18@ 32K—[NVIDIA Developer Forum↗](https://forums.developer.nvidia.com/t/kimi-2-6-and-qwen-3-5-397b-fp8-on-8xgb10-cluster/369446)[Mac Mini M4 \(16 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-16)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~17@ 4K—[maloyan\.xyz \(M4 16GB\)↗](https://maloyan.xyz/blog/running-qwen-locally-mac-mini-m4)[Tesla V100 32 GB SXM2 mod build](https://llmrequirements.com/hardware/tesla-v100-sxm2-mod)qwen3\-6\-27b\-dense@ Q5\_K\_M on llama\.cpp~17@ 32K—[hardware\-corner\.net \(V100 32GB guide\)↗](https://www.hardware-corner.net/guides/tesla-v100-32gb-for-llm/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-27b\-dense@ Q5\_K\_M on llama\.cpp~16@ 4K—[llm\-tracker\.info \(kyuz0\)↗](https://llm-tracker.info/_TOORG/Strix-Halo)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)deepseek\-v4\-flash\-284b\-moe@ IQ2\_XXS\-w2Q2K imatrix \(~80\.8 GB\) on ds4 \(antirez DeepSeek\-V4\-Flash engine\) \+ MTP, ROCm 7\.2\.4 gfx1151~15\.25@ 2K~152@ 2K \(derived\)[kyuz0 ds4 Strix Halo toolbox \(ds4\-bench, single 128GB node\); antirez ds4 engine↗](https://github.com/kyuz0/strix-halo-ds4-toolbox)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)deepseek\-v4\-flash\-284b\-moe@ Hybrid Q2/Q4 imatrix \(layers 37\-42 Q4, ~97 GB\) on ds4 \(antirez engine\) \+ MTP, ROCm 7\.2\.4 gfx1151~15\.02@ 2K~138@ 2K[kyuz0 ds4 Strix Halo toolbox \(ds4\-bench, single\-node hybrid Q2/Q4\)↗](https://github.com/kyuz0/strix-halo-ds4-toolbox)[MacBook Air M4 \(16 GB\)](https://llmrequirements.com/hardware/mba-m4-16)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM~15 est\.@ 4K—[maloyan\.xyz \(M4 16GB, fanless\-adjusted\)↗](https://maloyan.xyz/blog/running-qwen-locally-mac-mini-m4)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)glm\-5\-2\-753b\-moe@ NVFP4 \(REAP\-less, high\-quality 4\-bit; cf\. nvidia/GLM\-5\.2\-NVFP4 model card\) on vLLM~15@ 131\.072K~500@ 131\.072K[NVIDIA Developer Forum 374832 \(REAP\-less NVFP4, custom vLLM fork TP=4\)↗](https://forums.developer.nvidia.com/t/fitting-a-high-quality-reap-less-glm-5-2-onto-4x-dgx-spark/374832)[Quad RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x4)devstral\-2\-123b@ IQ4\_KSS \(GGUF\) on ik\_llama\.cpp \(\-sm graph, tensor\-parallel\)~15 est\.@ 4K~300@ 2K[HF ubergarm Devstral\-2\-123B\-GGUF discussion \#2 \(phakio, ik\_llama\.cpp 4\-GPU\)↗](https://huggingface.co/ubergarm/Devstral-2-123B-Instruct-2512-GGUF/discussions/2)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)mimo\-v2\-5\-310b\-a15b\-moe@ UD\-Q4/Q5\_K\_XL GGUF \(~180\-215 GB, split across 2 nodes\) on llama\.cpp RPC \(2x Strix Halo 128GB, ROCm, USB4net secondary link\)~15@ 10K~356@ 10K \(derived\)[r/LocalLLaMA operator report \(2x Strix Halo 128GB, llama\.cpp RPC over USB4net\); AesSedai/unsloth MiMo\-V2\.5 GGUF↗](https://old.reddit.com/r/LocalLLaMA/comments/1uboiko/rollin_mimo25_on_two_halo_strixeses/)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)nemotron\-3\-super\-120b\-a12b\-moe@ UD\-Q4\_K\_XL on llama\.cpp \(ROCm 7\.2\.3\)~14@ 0\.5K~276@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)kimi\-k2\-6\-1t\-moe@ NVFP4 \(60 shards, ~554 GiB, no spec\-decode\) on vLLM~13\.5@ 32\.768K—[NVIDIA Developer Forum 369446 \(vLLM eugr fork TP=8\)↗](https://forums.developer.nvidia.com/t/kimi-2-6-and-qwen-3-5-397b-fp8-on-8xgb10-cluster/369446)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)deepseek\-v4\-flash\-284b\-moe@ Q4 imatrix distributed \(~153\.3 GB, Q4 experts\) on ds4 multi\-node \(pipeline\-parallel, 2x Strix, ROCm 7\.2\.4 gfx1151\) \+ MTP~13\.01@ 2K~62@ 2K \(derived\)[kyuz0 ds4 Strix Halo toolbox \(ds4\-bench, 2\-node distributed Q4\)↗](https://github.com/kyuz0/strix-halo-ds4-toolbox)[Dual RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x2)gpt\-oss\-120b@ Q4\_K\_M \(\-\-n\-cpu\-moe 26, tensor\-split 0\.5/0\.5\) on llama\.cpp \(CUDA, 2x GPU \+ CPU\-MoE\)~13@ 4K—[LLM Garage \- GPT\-OSS\-120B on Dual RTX 3090s↗](https://llmgarage.ai/gpt-oss-120b-dual-3090/)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)glm\-5\-2\-753b\-moe@ low\-bit \(vLLM TP=2\) on vLLM \(TP over 2 Sparks\)~12@ 40K—[NVIDIA Developer Forum 374523 \(GLM\-5\.2 vLLM TP=2 update\)↗](https://forums.developer.nvidia.com/t/academic-glm-5-2-on-2x-dgx-spark-gb10-nodes-crazy-1-bit-ud-iq1-s-rpc-llama-cpp-256k-context-8-tok-s/374523)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gemma\-4\-31b@ UD\-Q4\_K\_XL on llama\.cpp \(Vulkan RADV\)~11@ 0\.5K~302@ 0\.5K[kyuz0 amd\-strix\-halo\-toolboxes grid \(docs/results\.json, 16 May 2026\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)glm\-5\-2\-753b\-moe@ UD\-IQ1\_S \(~2\.3 bpw, 1\-bit\) on llama\.cpp RPC \(tensor\-split over 2 nodes\)~8@ 2K~213@ 2K[NVIDIA Developer Forum 374523 \(GLM\-5\.2 on 2x DGX Spark, 1\-bit llama\.cpp RPC\)↗](https://forums.developer.nvidia.com/t/academic-glm-5-2-on-2x-dgx-spark-gb10-nodes-crazy-1-bit-ud-iq1-s-rpc-llama-cpp-256k-context-8-tok-s/374523)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)glm\-5\-2\-753b\-moe@ IQ4\_XS \(GGUF, ~365GB across 4 nodes, DSA sparse attention active\) on llama\.cpp \(RPC multi\-node\)~6\.28@ 1048\.576K~222@ 1048\.576K[NVIDIA Developer Forum 373933 \(IQ4\_XS llama\.cpp RPC, DSA active\)↗](https://forums.developer.nvidia.com/t/glm-5-2-iq4-xs-on-4x-gb10-6-28-tok-s-dsa-active-full-recipe/373933)[RTX 3060 12 GB build](https://llmrequirements.com/hardware/rtx-3060-12)gemma\-4\-12b@ Q4\_K\_M \(GGUF\) on llama\.cpp \(Vulkan\)~5@ 4K—[Hacker News Gemma 4 12B launch thread \(community, Vulkan\-throttled\)↗](https://news.ycombinator.com/item?id=48385906)[8× Strix Halo cluster \(1024 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x8)kimi\-k2\-6\-1t\-moe@ Q5\_K\_M on llama\.cpp~5@ 4K—[Frame\.work Community↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[Quad Tesla P40 \(96 GB\) homelab build](https://llmrequirements.com/hardware/tesla-p40-x4)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp~4@ 32K—[GitHub \- llama\.cpp \#12990 \(P40 FA\)↗](https://github.com/ggml-org/llama.cpp/issues/12990)[2× Strix Halo cluster \(256 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x2)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp~3@ 4K—[llm\-tracker\.info \(kyuz0\)↗](https://llm-tracker.info/_TOORG/Strix-Halo)[Single RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-24)gpt\-oss\-120b@ MXFP4 \(F16 GGUF, \-\-n\-cpu\-moe offload\) on llama\.cpp \(CUDA \+ CPU\-MoE offload\)~2@ 0\.1K~48@ 0\.1K[hardware\-corner\.net gpt\-oss CPU\-MoE offloading benchmark↗](https://www.hardware-corner.net/gpt-oss-offloading-moe-layers/)[Single AMD Instinct MI50 32 GB \(used\) build](https://llmrequirements.com/hardware/amd-mi50-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M on llama\.cpp——[GitHub llama\.cpp \#19880 \(MI50 enablement\)↗](https://github.com/ggml-org/llama.cpp/issues/19880)[Quad AMD MI50 32 GB \(128 GB\) homelab build](https://llmrequirements.com/hardware/amd-mi50-x4)qwen3\-6\-35b\-a3b\-moe@ Q5\_K\_M on llama\.cpp——[aibytes\.blog \(MI50 ROCm vs Vulkan\)↗](https://aibytes.blog/comparisons/rocm-7-vs-vulkan-on-mi50-4-model-benchmark-results)[Quad AMD MI50 32 GB \(128 GB\) homelab build](https://llmrequirements.com/hardware/amd-mi50-x4)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp——[HF \- unsloth \(GGUF discussion\)↗](https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF/discussions/5)[Mac Mini M4 \(16 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-16)gemma\-4\-12b@ MLX 4\-bit on Ollama \(MLX\) / mlx\-lm——[Ollama model library \(Apple\-Silicon MLX build; runnable recipe\)\. Fit corroborated by Gemma 4 launch HN thread \(Q4\_K\_M ~6\.6 GB\)\.↗](https://ollama.com/library/gemma4:12b)[Mac Mini M4 \(24 GB\)](https://llmrequirements.com/hardware/mac-mini-m4-24)gemma\-4\-12b@ MLX 4\-bit on Ollama \(MLX\) / mlx\-lm——[Ollama model library \(Apple\-Silicon MLX build\)\. Fit corroborated by HN launch thread\.↗](https://ollama.com/library/gemma4:12b)[MacBook Air M4 \(16 GB\)](https://llmrequirements.com/hardware/mba-m4-16)gemma\-4\-12b@ MLX 4\-bit on Ollama \(MLX\) / mlx\-lm——[Ollama model library \(Apple\-Silicon MLX build\)\. Fit corroborated by HN launch thread\.↗](https://ollama.com/library/gemma4:12b)[NVIDIA DGX Spark \(128 GB\)](https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on SGLang\+MTP——[GitHub r0b0tlab↗](https://github.com/r0b0tlab/qwen36-35b-a3b-nvfp4-gb10-native-mtp)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ NVFP4\-KV \(nvfp4\_ds\_mla\) on vLLM\+DSpark——[HF drowzeys↗](https://huggingface.co/drowzeys/DeepSeek-V4-Flash-DSpark-NVFP4-KV-1.5M-CTX-2xDGX-Spark)[2× DGX Spark cluster \(256 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x2)qwen3\-6\-35b\-a3b\-moe@ FP8 on vLLM \(Ray TP=2\)——[Medium \- Michael Peres↗](https://medium.com/@michaelperes1/turning-two-dgx-sparks-into-a-local-llm-cluster-with-vllm-ray-and-qwen3-6-7eb2a6e04ade)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 on vLLM——[NVIDIA Developer Forum↗](https://forums.developer.nvidia.com/t/deepseek-v4-flash-official-fp8-running-across-2x-dgx-spark-tp-2-mtp-200k-ctx-recipe-numbers/370309)[4× DGX Spark cluster \(512 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x4)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \(FP8 KV, MTP\) on vLLM \(TP=4\)——[NVIDIA\-NeMo Nemotron Spark Deployment Guide \(4x DGX Spark, NVFP4, vLLM TP=4\)↗](https://github.com/NVIDIA-NeMo/Nemotron/blob/main/usage-cookbook/Nemotron-3-Ultra/SparkDeploymentGuide/README.md)[8× DGX Spark cluster \(1024 GB unified, CUDA\)](https://llmrequirements.com/hardware/dgx-spark-x8)deepseek\-v4\-pro\-1t6\-moe@ FP8 on vLLM——[GitHub \- vLLM \#43367↗](https://github.com/vllm-project/vllm/issues/43367)[Single Intel Arc B580 12 GB build](https://llmrequirements.com/hardware/intel-arc-b580-12)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp——[GitHub \- intel/llm\-scaler \(official\)↗](https://github.com/intel/llm-scaler)[Quad RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x4)mistral\-medium\-3\-5\-128b@ Q4\_K\_M on llama\.cpp——[HF \- bartowski \(GGUF\)↗](https://huggingface.co/bartowski/mistralai_Mistral-Medium-3.5-128B-GGUF)[Quad RTX 3090 \(used\) build](https://llmrequirements.com/hardware/rtx-3090-x4)qwen3\-6\-35b\-a3b\-moe@ FP16 on vLLM——[GitHub \- tfriedel \(RTX 3090 lab\)↗](https://github.com/tfriedel/qwen3.6-rtx3090-lab)[Single RTX 4090 build](https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-6\-27b\-dense@ AWQ on vLLM——[GitHub \- thc1006 \(Ampere/Ada spec\-decode\)↗](https://github.com/thc1006/qwen3.6-speculative-decoding-rtx3090)[Single RTX 4090 build](https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M on llama\.cpp——[GitHub \- thc1006 \(spec\-decode\)↗](https://github.com/thc1006/qwen3.6-speculative-decoding-rtx3090)[RTX 3060 12 GB build](https://llmrequirements.com/hardware/rtx-3060-12)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp——[GitHub \- llama\.cpp build docs↗](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md)[Dual RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-x2)qwen3\-6\-35b\-a3b\-moe@ NVFP4 on vLLM——[HF \- RedHatAI \(NVFP4\)↗](https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4)[Dual RTX 5090 build](https://llmrequirements.com/hardware/rtx-5090-x2)qwen3\-6\-27b\-dense@ FP8 on vLLM——[HF \- Qwen \(official FP8\)↗](https://huggingface.co/Qwen/Qwen3.6-27B-FP8)[Single RTX Pro 6000 Blackwell 96 GB build](https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)mistral\-medium\-3\-5\-128b@ NVFP4 on vLLM——[HF nvidia \(NVFP4 card\)↗](https://huggingface.co/nvidia/Mistral-Medium-3.5-128B-NVFP4)[Single Tesla P40 24 GB \(used\) build](https://llmrequirements.com/hardware/tesla-p40-24)qwen3\-6\-27b\-dense@ Q4\_K\_M on llama\.cpp——[GitHub llama\.cpp \#19248 \(P40\)↗](https://github.com/ggml-org/llama.cpp/discussions/19248)[Quad Tesla P40 \(96 GB\) homelab build](https://llmrequirements.com/hardware/tesla-p40-x4)qwen3\-6\-35b\-a3b\-moe@ Q5\_K\_M on llama\.cpp——[GitHub \- llama\.cpp \#12990 \(P40 FA\)↗](https://github.com/ggml-org/llama.cpp/issues/12990)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)mimo\-v2\-5\-310b\-a15b\-moe@ UD\-IQ2\_M \(~2\.7 bpw, ~92\.8 GB\) on llama\.cpp \(Vulkan/RADV, kyuz0 container\), gfx1151—~31@ 0\.5K \(derived\)[hogeheer499 strix\-halo\-guide community evidence map \(Corsair AI WS 300, IQ2\_M, capacity row\); bartowski MiMo GGUF↗](https://huggingface.co/bartowski/MiMo-V2.5-GGUF)[AMD Ryzen AI Max\+ 395 \(128 GB\)](https://llmrequirements.com/hardware/strix-halo-128)gemma\-4\-12b@ UD\-Q4\_K\_XL \(GGUF\) on llama\.cpp \(ROCm 7\.2\.x, gfx1151\)——[kyuz0 Strix Halo toolboxes \(ROCm gfx1151 llama\.cpp recipe; 12B not yet in grid\)↗](https://kyuz0.github.io/amd-strix-halo-toolboxes/)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)qwen3\-5\-397b\-a17b\-moe@ Q\-family GGUF \(HIP\+RPC, np2, ctx 200k\) on llama\.cpp——[visorcraft/strix\-halo\-llm\-perf \(Qwen3\.5\-397B RPC shape\-control, 2026\-02\-21\)↗](https://github.com/visorcraft/strix-halo-llm-perf/blob/main/results/2026-02-21_qwen397b-rpc-shape-control.md)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)deepseek\-v4\-flash\-284b\-moe@ Q4\_K\-class across 4 nodes on llama\.cpp RPC \(4x Framework Desktop / Strix mainboards\)——[frame\.work llama\.cpp RPC multi\-node recipe \(extended to 4 nodes\)↗](https://community.frame.work/t/building-a-two-node-amd-strix-halo-cluster-for-llms-with-llama-cpp-rpc-minimax-m2-glm-4-6/77583)[4× Strix Halo cluster \(512 GB unified\)](https://llmrequirements.com/hardware/strix-halo-x4)mimo\-v2\-5\-310b\-a15b\-moe@ Q4\_K\_M \(~178 GB\) / Q5\_K\_M \(~213 GB\) on llama\.cpp RPC \(4x Framework Desktop / Strix mainboards\)——[bartowski MiMo\-V2\.5 GGUF \(Q4\_K\_M/Q5\) \+ frame\.work llama\.cpp RPC 4\-node↗](https://huggingface.co/bartowski/MiMo-V2.5-GGUF)[MacBook Pro M5 Pro 48 GB](https://llmrequirements.com/hardware/mbp-m5-pro-48)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit on MLX\-LM——[GitHub ml\-explore/mlx\-lm↗](https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md)[MacBook Pro M5 Pro 48 GB](https://llmrequirements.com/hardware/mbp-m5-pro-48)gemma\-4\-12b@ MLX 4\-bit on Ollama 0\.31 \(MLX\) \+ MTP——[Ollama blog \(framework\-author first\-party MTP recipe; M5\-family\)\. Directional only; M5 Pro not separately measured\.↗](https://ollama.com/blog/faster-gemma-4-mlx-mtp)[MacBook Pro M5 Pro 48 GB](https://llmrequirements.com/hardware/mbp-m5-pro-48)qwen3\-6\-27b\-dense@ MLX 4\-bit on mlx\-lm / Ollama \(MLX\)——[Ollama model library \(Apple\-Silicon MLX build\)\. Fit from db\.json sizeQ4=16 vs 40 GB usable\.↗](https://ollama.com/library/qwen3.6)[DGX H200 — 8× H200 server \(1\.13 TB HBM3e\)](https://llmrequirements.com/hardware/h200-x8)kimi\-k2\-7\-code\-1t\-moe@ native INT4 on vLLM \(TP=8, expert\-parallel\)——[vLLM Recipes \(official K2\.7 Code command, 8xH200 INT4\); tok/s = K2\.6 same\-box SGLang INT4 \(paxsaroffcuts\)↗](https://recipes.vllm.ai/moonshotai/Kimi-K2.7-Code)