Deepseek, please explain to me how you make a 300B parameter model that is cheaper than a 9B parameter model by SO MUCH.
Summary
A Reddit user asks how DeepSeek can make a 300B parameter model cheaper than a 9B parameter model, sparking discussion on efficiency and cost.
Similar Articles
Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?
The user asks why DeepSeek v4 Flash (284B parameters) is so cheap to run compared to smaller models like Qwen 27B, questioning if it's due to pricing dumping or architectural differences. The answer likely involves its MoE architecture and efficient inference techniques.
@che_shr_cat: 1/ Parameter scale is a brute-force crutch. What if a 35B model could beat a 1,000B model simply by scaling its search …
Explores whether a 35B parameter model could surpass a 1000B model by scaling its search horizon at test time, using structured process feedback rather than brute-force parameter scaling.
deepseek-ai/DeepSeek-V4-Pro-DSpark
DeepSeek releases preview versions of its V4 series, including DeepSeek-V4-Pro (1.6T parameters, 49B activated) and DeepSeek-V4-Flash (284B parameters, 13B activated), both supporting a one-million-token context and featuring hybrid attention, manifold-constrained hyper-connections, and a Muon optimizer.
@scaling01: DeepSeek just made their inference ~5x cheaper at 50 TPS
DeepSeek has reduced inference costs by approximately 5x while maintaining 50 tokens per second throughput.
@rohanpaul_ai: Download share by model size for DeepSeek vs. Qwen. DeepSeek dominates the 250B+ segment (47%) while Qwen leads sub-10B…
A comparison of download shares shows DeepSeek dominates the 250B+ parameter segment (47%) while Qwen leads the sub-10B segment (44%), highlighting complementary specialization between the two model families.