@scaling01: DeepSeek just made their inference ~5x cheaper at 50 TPS
Summary
DeepSeek has reduced inference costs by approximately 5x while maintaining 50 tokens per second throughput.
View Cached Full Text
Cached at: 06/29/26, 02:26 AM
DeepSeek just made their inference ~5x cheaper at 50 TPS https://t.co/9lYUGsshdb
Lisan al Gaib (@scaling01): how can you not like deepseek
thank you lord wenfeng for continuing to make intelligence too cheap to meter
Similar Articles
@runinfrai: deepseek V4.1 flash is live on runinfra 378 tok/s output. 30ms to first token $0.14 / 1M input. $0.03 cached. $0.58 out…
DeepSeek V4.1 Flash is now available on RunInfra's inference API, offering 378 tok/s output speed, 30ms time-to-first-token, and competitive pricing at $0.14/1M input tokens. The service supports OpenAI- and Anthropic-compatible endpoints with zero data retention by default.
@rohanpaul_ai: NVIDIA's newly published report says its Blackwell inference stack cut DeepSeek V4 token costs by up to 5x in one month.
NVIDIA reported that its Blackwell inference stack reduced DeepSeek V4 token costs by up to 5x in one month.
DeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says (4 minute read)
DeepSeek's new V4-Flash AI model is reported to be the cheapest well-known model to run, costing 105 times less than Anthropic's Claude Fable 5.
)
DeepSeek unveiled its new V4-Pro-0813 model, priced at $0.87 per million output tokens, while claiming strong benchmark performance against Anthropic's Opus 4.8 and intensifying the AI price war with OpenAI.
DeepSeek Announces Permanent Price Cut of 75% after Promotion Period
DeepSeek has announced a permanent 75% price reduction following a promotional period, making its AI services significantly cheaper for users.