@YRSM_Simon: Crazy

X AI KOLs Timeline Models

Summary

DeepSeek-V4-Flash-DSpark achieves 328 tok/s single inference and 1.7k tok/s batch throughput on 4x RTX PRO 6000 GPUs.

Crazy
Original Article
View Cached Full Text

Cached at: 07/09/26, 09:40 AM

Crazy

AI✖️Satoshi⏩️ (@AiXsatoshi): DeepSeek-V4-Flash-DSpark on 4x RTX PRO 6000 has come quite far.

  • Single inference: ~328 tok/s
  • Batch throughput: ~1.7k tok/s

Similar Articles

@Snixtp: DeepSeek V4 Flash on a single RTX Pro 6000?

X AI KOLs Following

DeepSeek V4 Flash GGUF quantizations have been released by antirez, enabling the model to run on single GPUs like the RTX Pro 6000 and Macs with 128GB+ RAM. The quantized files are available on Hugging Face with instructions for the DS4 inference engine.

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.