@QuixiAI: I got DeepSeek v4 Flash 0731 running on 4x A100 with SlimServe. 175 tok/s for single-request 1k tok/s for 64 concurrent…

X AI KOLs Timeline News

Summary

QuixiAI reports running DeepSeek v4 Flash 0731 on 4x A100 with SlimServe, achieving 175 tok/s for single requests and 1k tok/s for 64 concurrent requests.

I got DeepSeek v4 Flash 0731 running on 4x A100 with SlimServe. 175 tok/s for single-request 1k tok/s for 64 concurrent requests. https://t.co/wFQCRVe0qb
Original Article
View Cached Full Text

Cached at: 08/11/26, 05:40 AM

I got DeepSeek v4 Flash 0731 running on 4x A100 with SlimServe.

175 tok/s for single-request 1k tok/s for 64 concurrent requests. https://t.co/wFQCRVe0qb

Similar Articles

DeepSeek-V4-Flash-0731 unsloth gguf on A100

Reddit r/LocalLLaMA

DeepSeek-V4-Flash-0731 is shown running as an unsloth GGUF quant on a single 40GB A100, with 17.7 tok/s and 6 experts loaded into VRAM, enabling a full agentic coding loop.

DeepSeek-v4-Flash-Mini 54GB GGUF running at ~20.5 t/s

Reddit r/LocalLLaMA

A community build crushes DeepSeek-V4-Flash down to a 54GB IQ2_XXS GGUF variant with aggressive 2-bit quantization, achieving ~20.5 tokens/s on local hardware while drastically reducing memory footprint.

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.