@teortaxesTex: However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equa…

X AI KOLs Following News

Summary

The discussion compares the training costs of various AI models, noting that Qwen 27B and DeepSeek V4 Flash are similarly costly in GPU hours per token, highlighting the role of MFU in pretraining efficiency.

However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equally costly in GPU-hours per 1T tokens (training) I wish we started seeing these figures again. V2, V3, gpt-oss, Nemotrons. Do you know any more relevant anchors?
Original Article
View Cached Full Text

Cached at: 08/17/26, 10:43 AM

However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equally costly in GPU-hours per 1T tokens (training) I wish we started seeing these figures again. V2, V3, gpt-oss, Nemotrons. Do you know any more relevant anchors?

Tiezhen WANG (@Xianbao_QIAN): On that:

27B dense is actually big:

Kimi K3 - 2.8T A104B DeepSeek V4 Pro - 1.6T A49B GLM 5.2 - 743B A39B

MiniMax M3 - 427B A26B DeepSeek V4 Flash - 284B A19B

so in terms of compute, the new Qwen 3.8 27B model is heavier than DeepSeek V4 Flash and is comparable with MiniMax

Similar Articles

Qwen 3.8 27b vs Deepseek Flash

Reddit r/LocalLLaMA

The post compares the open-source AI models Qwen 3.8 (27B) and Deepseek Flash, discussing benchmarks and seeking user experiences to evaluate their performance.

Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?

Reddit r/LocalLLaMA

The user asks why DeepSeek v4 Flash (284B parameters) is so cheap to run compared to smaller models like Qwen 27B, questioning if it's due to pricing dumping or architectural differences. The answer likely involves its MoE architecture and efficient inference techniques.