@maximelabonne: Turns out you never really needed µP, you just needed to scale the embedding learning rate by model width I'm no nanoGP…

X AI KOLs Following News

Summary

A tweet suggests that scaling the embedding learning rate by model width can replace the need for µP (micro-parameterization), referencing Muon optimizer for hidden layers and Adam for the rest.

Turns out you never really needed µP, you just needed to scale the embedding learning rate by model width I'm no nanoGPT speedrunner, but isn't it something people stumbled into by using Muon for hidden layers + Adam for the rest? https://t.co/Ybs5C2rhlH
Original Article
View Cached Full Text

Cached at: 05/22/26, 03:49 AM

Turns out you never really needed µP, you just needed to scale the embedding learning rate by model width

I’m no nanoGPT speedrunner, but isn’t it something people stumbled into by using Muon for hidden layers + Adam for the rest? https://t.co/Ybs5C2rhlH

Similar Articles

@YRSM_Simon: Amazing!

X AI KOLs Timeline

shi3z shares optimization experience for running DeepSeek v4.1 Flash on an A100 GPU without FP4 support, boosting inference speed from 33 tok/s to 673 tok/s, surpassing the official API speed.