@_avichawla: A good technical LLM interview question: You replace a dense feed-forward layer with a top-2 MoE. The profiler confirms…
Summary
The article explains why replacing a dense feed-forward layer with a top-2 Mixture of Experts (MoE) in LLMs can increase inference latency due to communication overheads in multi-GPU setups, despite reducing FLOPs per token.
View Cached Full Text
Cached at: 09/27/26, 01:21 PM
A good technical LLM interview question:
You replace a dense feed-forward layer with a top-2 MoE.
The profiler confirms that the MoE executes fewer FLOPs per token.
However, end-to-end inference latency has increased instead of decreasing.
Why did this happen?
(answer below)
A standard Transformer applies the same feed-forward network to every token in a layer.
An MoE layer replaces that network with multiple experts. Each expert is a feed-forward network with its own weights.
A router reads each token’s hidden-state vector, scores the available experts, and selects the top two. It also produces a routing weight for each selection.
Only those selected experts process the token. The model therefore contains many expert weights while executing only a small subset for each token.
The serving system must move activations when the selected experts are stored on different GPUs.
The visual follows a batch that begins on GPU 0.
The router returns two expert IDs and two routing weights per token. One selected path is shown for each example token to keep the diagram readable.
Token T1 selects E1 on GPU 0. Its activation remains in local GPU memory, where E1 processes it.
T2 selects E4 on GPU 1 in the same server. GPU 0 transfers the activation through NVLink, a high-bandwidth GPU-to-GPU connection. NVSwitch connects several GPUs through this fabric.
T3 selects E7 on GPU 2 in another server. Its activation crosses the cluster network through InfiniBand or Ethernet.
To be clear, the system transfers activation vectors, which are the numerical representations produced for a token by the preceding Transformer operations.
Each destination GPU runs its expert and returns the output to GPU 0.
GPU 0 multiplies both selected expert outputs by their routing weights, adds them together, and restores the original token order.
Different tokens can select different expert pairs. Across a batch, those assignments may involve every GPU storing experts.
Each GPU may therefore send activations to several GPUs while receiving activations for its own experts. This exchange is called all-to-all communication.
And there are several ways to optimize this.
For instance, keeping frequently selected experts within the same server reduces remote traffic.
Balanced routing prevents one GPU from delaying the layer.
Communication overlap allows local expert computation to continue while remote activations move.
In this setup, sparse routing reduces expert FLOPs, but token dispatch and network communication erase part of that saving.
If you want to dive deeper, I recently covered MoE inference engineering, including routing, expert placement, token dispatch, load balancing, communication overlap, memory, quantization, and offloading.
Read my article below.
Similar Articles
@_avichawla: https://x.com/_avichawla/status/2100876555409039605
The article explains the engineering aspects of Mixture-of-Experts (MoE) inference, detailing token routing, expert batching, GPU distribution, and performance trade-offs for efficient serving.
@_avichawla: A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on lo…
Explains why evicting 90% of KV cache tokens fails to free GPU memory when serving reasoning models on vLLM, due to paged attention fragmentation, and introduces NVIDIA's TriAttention as a solution that achieves 2.5x speedup and 10.7x memory reduction.
Multi Tier MoE Caching
Discusses multi-tier caching strategies for MoE models to improve inference speed by keeping frequently activated experts on GPU, referencing existing implementations like PowerInfer and llama.cpp branches.
Proposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p]
A llama.cpp implementation that expands MoE expert routing beyond native top-K during inference using adaptive thresholds and layer-specific linear decay, without requiring model retraining or fine-tuning.
Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
This paper proposes L0-MoE, a lightweight Mixture-of-Experts approach using L0-regularization to accelerate dense Large Language Models with up to 2.5x speedup while maintaining competitive performance.