Deepseek, please explain to me how you make a 300B parameter model that is cheaper than a 9B parameter model by SO MUCH.

Reddit r/singularity News

Summary

A Reddit user asks how DeepSeek can make a 300B parameter model cheaper than a 9B parameter model, sparking discussion on efficiency and cost.

https://preview.redd.it/tjbwkmn4djgh1.png?width=1489&format=png&auto=webp&s=d11ec03569d082cdaf806c131b5be19e407187dd How?
Original Article

Similar Articles

Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?

Reddit r/LocalLLaMA

The user asks why DeepSeek v4 Flash (284B parameters) is so cheap to run compared to smaller models like Qwen 27B, questioning if it's due to pricing dumping or architectural differences. The answer likely involves its MoE architecture and efficient inference techniques.

deepseek-ai/DeepSeek-V4-Flash-0731

Simon Willison's Blog

DeepSeek released DeepSeek-V4-Flash-0731, a 304B parameter model with enhanced agentic capabilities, priced at $0.14/M input and $0.27/M output, punching above its weight and ranking as the best value-per-intelligence model according to Artificial Analysis.

DeepSeek-V3 Technical Report

Papers with Code Trending

DeepSeek-V3 is a parameter-efficient Mixture-of-Experts language model with 671B total parameters, achieving strong performance comparable to leading closed-source models while requiring only 2.788M H800 GPU hours for training.