Tag
The discussion compares the training costs of various AI models, noting that Qwen 27B and DeepSeek V4 Flash are similarly costly in GPU hours per token, highlighting the role of MFU in pretraining efficiency.
面壁智能 has open-sourced ForgeTrain, a pretraining framework autonomously written by an AI Agent. It achieves 44% MFU on H100, about 10% higher than the Megatron-LM baseline, marking an iteration of AI self-evolution.