Tag
The discussion compares the training costs of various AI models, noting that Qwen 27B and DeepSeek V4 Flash are similarly costly in GPU hours per token, highlighting the role of MFU in pretraining efficiency.