training-cost

Tag

Cards List
#training-cost

@teortaxesTex: However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equa…

X AI KOLs Following ↗ · 2026-08-16 Cached

The discussion compares the training costs of various AI models, noting that Qwen 27B and DeepSeek V4 Flash are similarly costly in GPU hours per token, highlighting the role of MFU in pretraining efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback