scaling-theory

Tag

Cards List
#scaling-theory

@hayou_soufiane: Teaching a short course with @nikhilghosh101 on the π“π‘πžπ¨π«π² 𝐨𝐟 π’πœπšπ₯𝐒𝐧𝐠 𝐒𝐧 𝐌𝐨𝐝𝐞𝐫𝐧 πƒπžπžπ© π‹πžπšπ«π§π’π§π β€¦

X AI KOLs Timeline β†— Β· 2026-07-02 Cached

Soufiane Hayou and Nikhil Ghosh will teach a short course on the Theory of Scaling in Modern Deep Learning at the STAI-X 2026 conference at Harvard University on July 31, 2026. The course covers hyperparameter transfer across scale with muP and variants.

0 favorites 0 likes
#scaling-theory

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

arXiv cs.LG β†— Β· 2026-05-15 Cached

This paper develops a principled scaling theory for Mixture-of-Experts (MoE) architectures, introducing the Maximally Scale-Stable Parameterization (MSSP) that ensures stable training and hyperparameter transfer across width, depth, expert width, and number of experts, validated by experiments.

0 favorites 0 likes
← Back to home

Submit Feedback