@songhan_mit: Not just fast, but iterate fast:
Summary
DeepSeek V4.1 Flash is now available on Inco AI, claiming to be the fastest provider with a speed of 532 tokens per second.
View Cached Full Text
Cached at: 09/18/26, 04:50 PM
Not just fast, but iterate fast:
Inco AI (@inco_ai): ⚡ 532 tokens/s!
DeepSeek V4.1 Flash is live on Inco, and it’s the fastest provider on Artificial Analysis. #1 output speed. Not close.
Try it: https://t.co/OaukyhhRer Results:
Similar Articles
@no_stp_on_snek: DeepSeek-V4.1-Flash on 2 Sparks. thinking high, TP=2. 55M tokens. Actual kernel work, not a demo. it read the notes and…
The article describes an optimization for the DeepSeek-V4.1-Flash AI model on Metal hardware, where the attention kernel was improved to only launch necessary tiles, reducing compute waste and enhancing inference speed.
@QuixiAI: I got DeepSeek v4 Flash 0731 running on 4x A100 with SlimServe. 175 tok/s for single-request 1k tok/s for 64 concurrent…
QuixiAI reports running DeepSeek v4 Flash 0731 on 4x A100 with SlimServe, achieving 175 tok/s for single requests and 1k tok/s for 64 concurrent requests.
DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]
DeepSeek V4 Flash (98GB) now runs up to 7 tokens per second on a single RTX 4060 Ti with CPU offloading, a 3x speed improvement over the previous week's 2 t/s.
@MiaAI_lab: DeepSeek v4 Flash has just been upgraded for your 2x DGX Sparks. 66.6 tokens per sec and up to 153.7 with 6 concurrent …
MiaAI Lab released an upgraded recipe for serving DeepSeek V4 Flash on two DGX Spark nodes using vLLM with DSpark speculative decoding and NVFP4 KV-cache, achieving up to 153.7 tokens per second with six concurrent sessions.
@LotusDecoder: DeepSeek-V4.1-Flash-0910 decode 400 token/s 😋 Would deploying this to my home DGX spark also achieve this speed?
A user on X/Twitter asks if deploying the DeepSeek-V4.1-Flash-0910 model on a home DGX Spark could achieve a decode speed of 400 tokens per second.