mfu

Tag

Cards List
#mfu

@teortaxesTex: However, 27B dense Qwen is likely pretrained with much higher MFU than DSV4-Flash 13AB has been, so they are about equa…

X AI KOLs Following ↗ · 2026-08-16 Cached

The discussion compares the training costs of various AI models, noting that Qwen 27B and DeepSeek V4 Flash are similarly costly in GPU hours per token, highlighting the role of MFU in pretraining efficiency.

0 favorites 0 likes
#mfu

@FeitengLi: #面壁智能 Opensources #ForgeTrain, a pretraining framework autonomously written by an AI Agent, even the CUDA kernels were written by itself. On H100, MiniCPM4-0.5B reaches 44% MFU, higher than Megatron (NVidia's main push for GPT implementation) baseline by about 10%. Starting AI self-evolution iteration

X AI KOLs Timeline ↗ · 2026-05-26 Cached

面壁智能 has open-sourced ForgeTrain, a pretraining framework autonomously written by an AI Agent. It achieves 44% MFU on H100, about 10% higher than the Megatron-LM baseline, marking an iteration of AI self-evolution.

0 favorites 0 likes
← Back to home

Submit Feedback