Deepseek, please explain to me how you make a 300B parameter model that is cheaper than a 9B parameter model by SO MUCH.

Reddit r/singularity News

Summary

A Reddit user asks how DeepSeek can make a 300B parameter model cheaper than a 9B parameter model, sparking discussion on efficiency and cost.

https://preview.redd.it/tjbwkmn4djgh1.png?width=1489&format=png&auto=webp&s=d11ec03569d082cdaf806c131b5be19e407187dd How?
Original Article

Similar Articles

Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?

Reddit r/LocalLLaMA

The user asks why DeepSeek v4 Flash (284B parameters) is so cheap to run compared to smaller models like Qwen 27B, questioning if it's due to pricing dumping or architectural differences. The answer likely involves its MoE architecture and efficient inference techniques.

deepseek-ai/DeepSeek-V4-Pro-DSpark

Hugging Face Models Trending

DeepSeek releases preview versions of its V4 series, including DeepSeek-V4-Pro (1.6T parameters, 49B activated) and DeepSeek-V4-Flash (284B parameters, 13B activated), both supporting a one-million-token context and featuring hybrid attention, manifold-constrained hyper-connections, and a Muon optimizer.