缓存时间:
2026/07/11 23:40
# 配方 — LLMRequirements.com
来源:https://llmrequirements.com/recipes
8× RTX Pro 6000 Blackwell 服务器 (768 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)kimi\-k2\-5\-1t\-moe@ INT4 (FP8 KV, DCP=8) 在 vLLM 上批次~900@ 40K(100\-并发. 聚合)—local\-inference\-lab/rtx6kpro wiki (Kimi K2.5 高并发, Festr)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)双 RTX Pro 6000 Blackwell 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ FP8 (fp8 KV, MTP spec=3) 在 vLLM 上批次~894@ 175K(32\-并发. 聚合)—theogravity/dual\-rtx\-6000\-blackwell\-qwen3\.6\-27b\-fp8 (编码扫描, seqs=32)↗ (https://github.com/theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8)8× H100 80 GB 服务器 (https://llmrequirements.com/hardware/h100-x8)deepseek\-v3\-671b\-moe@ FP8 (TP=8) 在 vLLM 上批次~620@ 1.024K(100\-并发. 聚合)—dzhsurf/deepseek\-v3\-r1\-deploy\-and\-benchmarks (8xH100 vLLM TP=8, ~100 并发)↗ (https://github.com/dzhsurf/deepseek-v3-r1-deploy-and-benchmarks)2× Strix Halo 集群 (256 GB 统一内存) (https://llmrequirements.com/hardware/strix-halo-x2)gemma4\-26b\-moe@ native 在 vLLM 集群 TP=2 上 (triton, ROCm)批次~411@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (triton 集群 tp2 吞吐量, 200 请求)↗ (https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/triton/google_gemma-4-26B-A4B-it_cluster_tp2_throughput.json)8× RTX Pro 6000 Blackwell 服务器 (768 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)qwen3\-5\-397b\-a17b\-moe@ NVFP4 在 SGLang\+MTP 上~350@ 4K—local\-inference\-lab/rtx6kpro wiki (Qwen3.5\-397B 8x 单批次)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)8× H100 80 GB 服务器 (https://llmrequirements.com/hardware/h100-x8)qwen3\-6\-35b\-a3b\-moe@ FP8 在 vLLM 上~320@ 128K—vLLM Recipes↗ (https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B)2× Strix Halo 集群 (256 GB 统一内存) (https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-6\-35b\-a3b\-moe@ AWQ\-4bit / native 在 vLLM 集群 TP=2 上 (aiter, ROCm)批次~287@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (aiter 集群 tp2 吞吐量, 2025年12月)↗ (https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)2× Strix Halo 集群 (256 GB 统一内存) (https://llmrequirements.com/hardware/strix-halo-x2)gpt\-oss\-120b@ MXFP4 在 vLLM 集群 TP=2 上 (triton, ROCm)批次~229@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (triton 集群 tp2 吞吐量, 2025年12月)↗ (https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)单 RTX 5090 配置 (https://llmrequirements.com/hardware/rtx-5090-32)qwen3\-coder\-30b@ Q4\_K 在 llama\.cpp 上 (CUDA)~226@ 4K~7093@ 4Khardware\-corner\.net RTX 5090 LLM 基准测试 (GGUF Q4)↗ (https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-5090/)8× H100 80 GB 服务器 (https://llmrequirements.com/hardware/h100-x8)mistral\-medium\-3\-5\-128b@ FP8 在 vLLM 上~220@ 128K—vLLM Recipes↗ (https://recipes.vllm.ai/mistralai/Mistral-Medium-3.5-128B)单 RTX 5090 配置 (https://llmrequirements.com/hardware/rtx-5090-32)gemma4\-26b\-moe@ Q4\_K 在 llama\.cpp 上 (CUDA)~180@ 4K~8799@ 4Khardware\-corner\.net RTX 5090 LLM 基准测试 (GGUF Q4)↗ (https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-5090/)单 H100 80 GB 工作站 (https://llmrequirements.com/hardware/h100-80)qwen3\-6\-35b\-a3b\-moe@ FP8 在 vLLM 上~180@ 128K—vLLM Recipes↗ (https://recipes.vllm.ai/Qwen/Qwen3.6-35B-A3B)4× DGX Spark 集群 (512 GB 统一内存, CUDA) (https://llmrequirements.com/hardware/dgx-spark-x4)deepseek\-v4\-flash\-284b\-moe@ FP8 (FP8 KV, MTP n=2) 在 vLLM 上批次~179.9@ 393.216K(8\-并发. 聚合)—NVIDIA Developer Forum 373808 (jasl vLLM TP=4, n=8 聚合)↗ (https://forums.developer.nvidia.com/t/deepseek-v4-flash-on-4x-dgx-spark-via-vllm-jasl-fork-tp-4-rdma-mtp-49-54-tok-s-single-stream-full-recipe-the-traps/373808)AMD Ryzen AI Max\+ 395 (128 GB) (https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ AWQ\-4bit / native 在 vLLM 上 (aiter, ROCm)批次~178@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (aiter tp1 吞吐量, 2025年12月)↗ (https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)单 RTX Pro 6000 Blackwell 96 GB 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-6\-35b\-a3b\-moe@ FP8 在 vLLM 上~170@ 262K—GitHub lastloop\-ai↗ (https://github.com/lastloop-ai/vllm-blackwell-guide)双 RTX Pro 6000 Blackwell 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ NVFP4 在 vLLM\+MTP 上~156@ 262K~831@ 262KloFT LLC↗ (https://loftllc.dev/en/docs/tech/llm-research/qwen3-6-27b-nvfp4-mtp-vllm-benchmark/)单 RTX 3090 (二手) 配置 (https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-coder\-30b@ Q4\_K 在 llama\.cpp 上 (CUDA)~153@ 4K~2988@ 4Khardware\-corner\.net RTX 3090 LLM 基准测试 (GGUF Q4)↗ (https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/)四 RTX Pro 6000 Blackwell 配置 (384 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)qwen3\-5\-397b\-a17b\-moe@ AWQ\-INT4 (QuantTrio) 在 SGLang\+MTP 上~152@ 4K—local\-inference\-lab/rtx6kpro wiki (Qwen3.5\-397B 单批次解码)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)双 RTX 3090 (二手) 配置 (https://llmrequirements.com/hardware/rtx-3090-x2)qwen3\-6\-35b\-a3b\-moe@ AWQ 在 vLLM 上~149@ 4K—GitHub \- tfriedel (RTX 3090 实验室)↗ (https://github.com/tfriedel/qwen3.6-rtx3090-lab)AMD Ryzen AI Max\+ 395 (128 GB) (https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ ROCmFP4 (CHADROCK) 在 llama\-server\+ROCmFPX 上~140@ 4K—GitHub hogeheer499\-commits/strix\-halo\-guide↗ (https://github.com/hogeheer499-commits/strix-halo-guide)双 RTX Pro 6000 Blackwell 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)qwen3\-6\-27b\-dense@ FP8 (native, fp8 KV, MTP spec=3) 在 vLLM 上~137@ 175K—theogravity/dual\-rtx\-6000\-blackwell\-qwen3\.6\-27b\-fp8 (基准扫描)↗ (https://github.com/theogravity/dual-rtx-6000-blackwell-qwen3.6-27b-fp8)NVIDIA DGX Spark (128 GB) (https://llmrequirements.com/hardware/dgx-spark-128)mistral\-small\-4\-119b\-moe@ NVFP4 在 vLLM 上批次~131@ 262.144K(20\-并发. 聚合)—Sebastien67 Medium (DGX Spark vLLM NVFP4, n=20 聚合)↗ (https://medium.com/@Sebastien67/running-mistral-small-4-119b-nvfp4-locally-on-a-dgx-spark-81cc2fdc4f6f)四 RTX Pro 6000 Blackwell 配置 (384 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)qwen3\-5\-397b\-a17b\-moe@ NVFP4 (nvidia 检查点) 在 vLLM\+MTP 上~130@ 4K—local\-inference\-lab/rtx6kpro wiki (Qwen3.5\-397B MTP 缩放表, 并发=1)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)2× Strix Halo 集群 (256 GB 统一内存) (https://llmrequirements.com/hardware/strix-halo-x2)gemma\-4\-31b@ native (bf16/fp16) 在 vLLM 集群 TP=2 上 (triton, ROCm)批次~128@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (triton 集群 tp2 吞吐量, 200 请求)↗ (https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/triton/google_gemma-4-31B-it_cluster_tp2_throughput.json)4× DGX Spark 集群 (512 GB 统一内存, CUDA) (https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 (modelopt\_fp4, fp8 KV) 在 SGLang\+MTP 上批次~124@ 196.608K(8\-并发. 聚合)—NVIDIA Developer Forum 373676 (SGLang TP=4 EP=4, n=8 聚合)↗ (https://forums.developer.nvidia.com/t/minimax-m2-5-nvfp4-on-4x-dgx-spark-via-sglang-tp-4-ep-4-124-tok-s-aggregate-n-8-fixing-the-cutlass-moe-compile-oom-with-max-jobs-1/373676)单 RTX Pro 6000 Blackwell 96 GB 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-next\-80b\-moe@ Q4\_K\_M 在 ollama / llama\.cpp 上 (CUDA)~124@ 4K~3274@ 4Kvaditaslim\.com RTX PRO 6000 Blackwell 8 模型基准测试↗ (https://www.vaditaslim.com/blog/ai/local-llm-benchmarks-rtx-pro-6000)单 RTX 3090 (二手) 配置 (https://llmrequirements.com/hardware/rtx-3090-24)gemma4\-26b\-moe@ Q4\_K 在 llama\.cpp 上 (CUDA)~119@ 4K~3625@ 4Khardware\-corner\.net RTX 3090 LLM 基准测试 (GGUF Q4)↗ (https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/)单 H100 80 GB 工作站 (https://llmrequirements.com/hardware/h100-80)qwen3\-6\-27b\-dense@ FP16 在 vLLM 上~110@ 128K—vLLM Recipes↗ (https://recipes.vllm.ai/Qwen/Qwen3.6-27B)2× Strix Halo 集群 (256 GB 统一内存) (https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-5\-122b\-a10b\-moe@ cyankiwi AWQ\-4bit 在 vLLM 集群 TP=2 上 (aiter, ROCm)批次~104@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (aiter 集群 tp2 吞吐量, 2025年12月)↗ (https://kyuz0.github.io/amd-strix-halo-vllm-toolboxes/)单 RTX 3090 (二手) 配置 (https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_XL 在 llama\.cpp 上~101@ 65K~1171@ 0.5Kaminrj\.com (Qwen3.6 在 24GB 上)↗ (https://aminrj.com/posts/llamacpp-qwen36-35b/)AMD Ryzen AI Max\+ 395 (128 GB) (https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ IQ4\_XS\-Q8nextn 在 llama\-server\+MTP 上~101@ 4K—GitHub hogeheer499\-commits/strix\-halo\-guide↗ (https://github.com/hogeheer499-commits/strix-halo-guide)8× RTX Pro 6000 Blackwell 服务器 (768 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)kimi\-k2\-5\-1t\-moe@ INT4 (BF16 KV, EP=8, 超频 GDDR7) 在 SGLang\+MTP 上~101@ 4K—local\-inference\-lab/rtx6kpro wiki (Kimi K2.5 8x 单批次解码)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)单 RTX Pro 6000 Blackwell 96 GB 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-96)qwen3\-6\-27b\-dense@ INT4 (AutoRound) \+ MTP n=3, FP8 KV 在 vLLM 上 (flashinfer, MTP)~100@ 262K—GitHub lastloop\-ai↗ (https://github.com/lastloop-ai/vllm-blackwell-guide)8× RTX Pro 6000 Blackwell 服务器 (768 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x8)glm\-51\-754b\-moe@ NVFP4\-MTP (lukealonso/GLM\-5\.1\-NVFP4\-MTP, 作为 GLM\-5 服务) 在 SGLang\+MTP 上~100@ 4K—local\-inference\-lab/rtx6kpro wiki (GLM\-5 单批次解码; models/glm5\.md = GLM\-5\.1)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)NVIDIA DGX Spark (128 GB) (https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 在 vLLM\+DFlash 上~97@ 0.5K~9090@ 0.5K (衍生)GitHub AEON\-7↗ (https://github.com/AEON-7/Qwen3.6-NVFP4-DFlash)MacBook Pro M5 Max 64 GB (https://llmrequirements.com/hardware/mbp-m5-max-64)gemma\-4\-12b@ MLX NVFP4 在 Ollama 0.31 上 (MLX) \+ MTP~95@ 4K—Ollama 博客 (框架作者第一方; M5 Max, Aider polyglot)↗ (https://ollama.com/blog/faster-gemma-4-mlx-mtp)单 RTX 5090 配置 (https://llmrequirements.com/hardware/rtx-5090-32)qwen3\-6\-27b\-dense@ NVFP4 在 vLLM 上~92@ 200K~5300@ 47KGitHub devnen↗ (https://github.com/devnen/qwen3.6-windows-server)NVIDIA DGX Spark (128 GB) (https://llmrequirements.com/hardware/dgx-spark-128)qwen3\-6\-35b\-a3b\-moe@ NVFP4 在 vLLM 上~90@ 43K~2133@ 32K (衍生)GitHub technigmaai/dgx\-spark↗ (https://github.com/technigmaai/dgx-spark/tree/main/spark-vllm-docker/nvidia-Qwen3.6-35B-A3B-NVFP4)双 RTX 3090 (二手) 配置 (https://llmrequirements.com/hardware/rtx-3090-x2)qwen3\-6\-27b\-dense@ AWQ 在 vLLM 上~90@ 100K—GitHub \- tfriedel (RTX 3090 实验室)↗ (https://github.com/tfriedel/qwen3.6-rtx3090-lab)MacBook Pro M5 Max 64 GB (https://llmrequirements.com/hardware/mbp-m5-max-64)qwen3\-6\-35b\-a3b\-moe@ MLX\-4bit 在 MLX\-LM 上~87@ 4K~2447@ 4KoMLX Benchmark↗ (https://omlx.ai/benchmarks/oykgm8sq)双 RTX Pro 6000 Blackwell 配置 (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x2)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-NVFP4 在 vLLM 上~85@ 4K—local\-inference\-lab/rtx6kpro wiki (MiniMax\-M2.5 单流表)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)AMD Ryzen AI Max\+ 395 (128 GB) (https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ Q4\_0 在 llama\.cpp 上~81@ 4K~1244@ 0.5KGitHub hogeheer499\-commits/strix\-halo\-guide↗ (https://github.com/hogeheer499-commits/strix-halo-guide)四 RTX Pro 6000 Blackwell 配置 (384 GB) (https://llmrequirements.com/hardware/rtx-pro-6000-blackwell-x4)minimax\-m2\-7\-230b\-moe@ MiniMax\-M2\.5\-FP8 在 vLLM 上~81@ 20K—local\-inference\-lab/rtx6kpro wiki (MiniMax\-M2.5 单流表)↗ (https://github.com/voipmonitor/rtx6kpro/blob/master/benchmarks/results.md)DGX B200 — 8× B200 服务器 (1.44 TB HBM3e) (https://llmrequirements.com/hardware/b200-x8)nemotron\-3\-ultra\-550b\-a55b\-moe@ NVFP4 \+ FP8 KV 在 Dynamo \+ vLLM 上 (TP=4, expert\-parallel, MTP)批次~80.6@ 4K(20\-并发. 聚合)—NVIDIA ai\-dynamo/dynamo recipes (B200 TP4\+EP, NVFP4\+FP8, MTP)↗ (https://github.com/ai-dynamo/dynamo/tree/main/recipes/nemotron-3-ultra)4× DGX Spark 集群 (512 GB 统一内存, CUDA) (https://llmrequirements.com/hardware/dgx-spark-x4)minimax\-m3\-428b\-moe@ MiniMax\-M3\-AWQ\-INT4 (fp8 KV, EAGLE3) 在 vLLM 上批次~79@ 262.144K(4\-并发. 聚合)—NVIDIA Developer Forum 375361 (vLLM TP=4, n=4 聚合)↗ (https://forums.developer.nvidia.com/t/minimax-m3-awq-running-tp-4-across-4x-dgx-spark-gb10-33-tok-s-full-recipe-the-gb10-build-fixes/375361)单 AMD Radeon AI Pro R9700 32 GB 配置 (https://llmrequirements.com/hardware/amd-r9700-32)qwen3\-6\-35b\-a3b\-moe@ Q4\_K\_M 在 llama\.cpp 上~77@ 4K~1636@ 32KGitHub truelies444↗ (https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)8× H100 80 GB 服务器 (https://llmrequirements.com/hardware/h100-x8)kimi\-k2\-6\-1t\-moe@ FP8 在 vLLM 上~75@ 256K—HF \- RedHatAI (Kimi\-K2\.6\-FP8\-BLOCK)↗ (https://huggingface.co/RedHatAI/Kimi-K2.6-FP8-BLOCK)AMD Ryzen AI Max\+ 395 (128 GB) (https://llmrequirements.com/hardware/strix-halo-128)qwen3\-6\-35b\-a3b\-moe@ MTP\-GGUF UD\-Q4\_K\_XL (draft\-mtp n=3) 在 llama\.cpp 上 (Vulkan RADV, MTP)~75@ 0.5K—kyuz0 amd\-strix\-halo\-toolboxes MTP 网格 (results\-mtp/summary\.json, 2026年5月15日)↗ (https://kyuz0.github.io/amd-strix-halo-toolboxes/mtp.html)2× Strix Halo 集群 (256 GB 统一内存) (https://llmrequirements.com/hardware/strix-halo-x2)qwen3\-5\-122b\-a10b\-moe@ cyankiwi AWQ\-8bit 在 vLLM 集群 TP=2 上 (aiter, ROCm)批次~74@ 4K(200\-并发. 聚合)—kyuz0 amd\-strix\-halo\-vllm\-toolboxes (aiter 集群 tp2 吞吐量, 200 请求)↗ (https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/benchmarks/vllm_benchmark_results/aiter/cyankiwi_Qwen3.5-122B-A10B-AWQ-8bit_cluster_tp2_throughput.json)双 AMD Radeon AI Pro R9700 配置 (64 GB) (https://llmrequirements.com/hardware/amd-r9700-x2)qwen3\-6\-35b\-a3b\-moe@ Q6\_K 在 llama\.cpp 上~72@ 4K~3038@ 32KGitHub truelies444↗ (https://github.com/truelies444/amd-radeon-ai-pro-r9700-llama-cpp-rocm-benchmarks)单 RTX 3090 (二手) 配置 (https://llmrequirements.com/hardware/rtx-3090-24)qwen3\-6\-27b\-dense@ AWQ/AutoRound\-INT4 在 vLLM\+MTP 上~72@ 32K—GitHub devnen↗ (https://github.com/devnen/qwen3.6-windows-server)单 RTX 4090 配置 (https://llmrequirements.com/hardware/rtx-4090-24)qwen3\-coder\-30b@ Q4\_K 在 llama\.cpp 上 (CUDA)~68@ 64K~1502@ 64Khardware\-corner\.net RTX 4090 LLM 基准测试 (GGUF Q4)↗ (https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-4090/)2× DGX Spark 集群 (256 GB 统一内存, CUDA) (https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ 未指定 (原生精度) 在 vLLM 上 (Aidendle94/B12X\-MoE, TP=2 RoCE) \+ DSpark spec\-decode~65@ 200K—GitHub 0rand (DeepSeek\-V4 DSpark 服务栈)↗ (https://github.com/0rand/DeepSeek-v4-DSpark-Aidendle94-GB10-ServingStack)2× DGX Spark 集群 (256 GB 统一内存, CUDA) (https://llmrequirements.com/hardware/dgx-spark-x2)deepseek\-v4\-flash\-284b\-moe@ NVFP4\-KV (