Tag
Soufiane Hayou and Nikhil Ghosh will teach a short course on the Theory of Scaling in Modern Deep Learning at the STAI-X 2026 conference at Harvard University on July 31, 2026. The course covers hyperparameter transfer across scale with muP and variants.
This paper develops a principled scaling theory for Mixture-of-Experts (MoE) architectures, introducing the Maximally Scale-Stable Parameterization (MSSP) that ensures stable training and hyperparameter transfer across width, depth, expert width, and number of experts, validated by experiments.