Tag
The article explains why replacing a dense feed-forward layer with a top-2 Mixture of Experts (MoE) in LLMs can increase inference latency due to communication overheads in multi-GPU setups, despite reducing FLOPs per token.
A new in-depth blog post explains collective communication for multiple GPUs, covering primitives like broadcast and reduce, and helps beginners understand how to scale experiments.