@teortaxesTex: However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equa…
Summary
The discussion compares the training costs of various AI models, noting that Qwen 27B and DeepSeek V4 Flash are similarly costly in GPU hours per token, highlighting the role of MFU in pretraining efficiency.
View Cached Full Text
Cached at: 08/17/26, 10:43 AM
However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equally costly in GPU-hours per 1T tokens (training) I wish we started seeing these figures again. V2, V3, gpt-oss, Nemotrons. Do you know any more relevant anchors?
Tiezhen WANG (@Xianbao_QIAN): On that:
27B dense is actually big:
Kimi K3 - 2.8T A104B DeepSeek V4 Pro - 1.6T A49B GLM 5.2 - 743B A39B
MiniMax M3 - 427B A26B DeepSeek V4 Flash - 284B A19B
so in terms of compute, the new Qwen 3.8 27B model is heavier than DeepSeek V4 Flash and is comparable with MiniMax
Similar Articles
Qwen 3.8 27b vs Deepseek Flash
The post compares the open-source AI models Qwen 3.8 (27B) and Deepseek Flash, discussing benchmarks and seeking user experiences to evaluate their performance.
@VraserX: Qwen3.8-Flash reportedly needs around ONE NINTH the training cost of Qwen3.7-Plus. This is the trend I think people und…
A tweet reports that the Qwen3.8-Flash model requires about one-ninth the training cost of the Qwen3.7-Plus model, highlighting a trend of decreasing AI training expenses.
Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?
The user asks why DeepSeek v4 Flash (284B parameters) is so cheap to run compared to smaller models like Qwen 27B, questioning if it's due to pricing dumping or architectural differences. The answer likely involves its MoE architecture and efficient inference techniques.
DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark
User benchmarks DeepSeek v4 Flash against Qwen3.6-27B, Qwen3.5-122B, and Gemma 4 31B on a local coding benchmark, finding Flash wins overall but Qwen 122B performs surprisingly well with better first-try success and lower token usage.
@EMostaque: So @deepseek_ai v4.1 Flash I estimate cost $10m to train, 100x less than GPT-6 Astra. It also costs 100x less to run an…
A tweet estimates that DeepSeek AI's v4.1 Flash model cost $10 million to train, 100 times less than GPT-6 Astra, with lower running costs and similar benchmark performance, prompting questions about the bull case for frontier AI labs.