@Jianlin_S: MoE (9): The Gate Normalization Debate https://kexue.fm/archives/11782
Summary
A blog post discussing the debate on gate normalization in Mixture of Experts (MoE) models.
Similar Articles
Mixture of Experts (MoEs) in Transformers
Hugging Face blog post explaining Mixture of Experts (MoEs) architecture in Transformers, covering the shift from dense to sparse models, weight loading optimizations, expert parallelism, and training techniques for MoE-based language models.
What is the point of MoE models, beyond being faster?
A discussion about the advantages of Mixture of Experts (MoE) models over dense models beyond speed, considering RAM constraints and scaling limits.
@_avichawla: https://x.com/_avichawla/status/2100876555409039605
The article explains the engineering aspects of Mixture-of-Experts (MoE) inference, detailing token routing, expert batching, GPU distribution, and performance trade-offs for efficient serving.
@jbhuang0604: Huge! It’s amazing how often Noam’s papers end up at the center of the field. In many tutorial videos I’ve made, they’v…
The article provides a detailed explanation of Mixture of Experts (MoE) in transformers, covering routing, load balancing, and recent innovations like fine-grained experts. It also highlights the significance of Noam Shazeer's research contributions and his move from Google to OpenAI.
Output Dilution: Redundant but Fragile Representations in MoE Models
This paper reveals that Mixture-of-Experts models encode moral content as robustly as dense models in probing but are far more fragile due to output dilution, where aggregated expert outputs dilute the signal, making representations vulnerable to noise.