gpu-communication

Tag

Cards List
#gpu-communication

@_avichawla: A good technical LLM interview question: You replace a dense feed-forward layer with a top-2 MoE. The profiler confirms…

X AI KOLs Timeline ↗ · 2026-09-26 Cached

The article explains why replacing a dense feed-forward layer with a top-2 Mixture of Experts (MoE) in LLMs can increase inference latency due to communication overheads in multi-GPU setups, despite reducing FLOPs per token.

0 favorites 0 likes
#gpu-communication

@jino_rohit: new in-depth blog post for "Collective Communication for Multiple GPUs". this blog should help you understand how commu…

X AI KOLs Following ↗ · 2026-06-13

A new in-depth blog post explains collective communication for multiple GPUs, covering primitives like broadcast and reduce, and helps beginners understand how to scale experiments.

0 favorites 0 likes
← Back to home

Submit Feedback