@scaling01: DeepSeek just made their inference ~5x cheaper at 50 TPS
Summary
DeepSeek has reduced inference costs by approximately 5x while maintaining 50 tokens per second throughput.
View Cached Full Text
Cached at: 06/29/26, 02:26 AM
DeepSeek just made their inference ~5x cheaper at 50 TPS https://t.co/9lYUGsshdb
Lisan al Gaib (@scaling01): how can you not like deepseek
thank you lord wenfeng for continuing to make intelligence too cheap to meter
Similar Articles
@rohanpaul_ai: NVIDIA's newly published report says its Blackwell inference stack cut DeepSeek V4 token costs by up to 5x in one month.
NVIDIA reported that its Blackwell inference stack reduced DeepSeek V4 token costs by up to 5x in one month.
DeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says (4 minute read)
DeepSeek's new V4-Flash AI model is reported to be the cheapest well-known model to run, costing 105 times less than Anthropic's Claude Fable 5.
)
DeepSeek unveiled its new V4-Pro-0813 model, priced at $0.87 per million output tokens, while claiming strong benchmark performance against Anthropic's Opus 4.8 and intensifying the AI price war with OpenAI.
DeepSeek Announces Permanent Price Cut of 75% after Promotion Period
DeepSeek has announced a permanent 75% price reduction following a promotional period, making its AI services significantly cheaper for users.
@cline: While DeepSeek V4-Flash is significantly cheaper on price per token, this can be misleading if the overall cost per tas…
Discusses DeepSeek V4-Flash's price per token vs. overall cost per task, citing a report that DeepSeek completes benchmark tasks at 105x lower cost than Fable.