@runinfrai: deepseek V4.1 flash is live on runinfra 378 tok/s output. 30ms to first token $0.14 / 1M input. $0.03 cached. $0.58 out…
Summary
DeepSeek V4.1 Flash is now available on RunInfra's inference API, offering 378 tok/s output speed, 30ms time-to-first-token, and competitive pricing at $0.14/1M input tokens. The service supports OpenAI- and Anthropic-compatible endpoints with zero data retention by default.
View Cached Full Text
Cached at: 10/03/26, 07:04 PM
deepseek V4.1 flash is live on runinfra
378 tok/s output. 30ms to first token
$0.14 / 1M input. $0.03 cached. $0.58 output
1M context. 99.3% cache hit rate over the last 24h
OpenAI-compatible chat completions and Anthropic-compatible /v1/messages – drop it into claude code as is
zero data retention by default
DeepSeek V4.1 Flash API pricing and speed | RunInfra
Source: https://runinfra.ai/inference-api/deepseek-v4-1-flash RunInfraby RightNow
© 2026 RunInfra. All rights reserved.
Similar Articles
@scaling01: DeepSeek just made their inference ~5x cheaper at 50 TPS
DeepSeek has reduced inference costs by approximately 5x while maintaining 50 tokens per second throughput.
@songhan_mit: Not just fast, but iterate fast:
DeepSeek V4.1 Flash is now available on Inco AI, claiming to be the fastest provider with a speed of 532 tokens per second.
DeepSeek V4 Flash Vision is now live !
DeepSeek has released vision capabilities for its V4 Flash AI model, providing a cheaper inference option through DeepInfra compared to the official API.
DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]
DeepSeek V4 Flash (98GB) now runs up to 7 tokens per second on a single RTX 4060 Ti with CPU offloading, a 3x speed improvement over the previous week's 2 t/s.
)
DeepSeek unveiled its new V4-Pro-0813 model, priced at $0.87 per million output tokens, while claiming strong benchmark performance against Anthropic's Opus 4.8 and intensifying the AI price war with OpenAI.