@_avichawla: A good technical LLM interview question: You replace a dense feed-forward layer with a top-2 MoE. The profiler confirms…

X AI KOLs Timeline News

Summary

The article explains why replacing a dense feed-forward layer with a top-2 Mixture of Experts (MoE) in LLMs can increase inference latency due to communication overheads in multi-GPU setups, despite reducing FLOPs per token.

A good technical LLM interview question: You replace a dense feed-forward layer with a top-2 MoE. The profiler confirms that the MoE executes fewer FLOPs per token. However, end-to-end inference latency has increased instead of decreasing. Why did this happen? (answer below) A standard Transformer applies the same feed-forward network to every token in a layer. An MoE layer replaces that network with multiple experts. Each expert is a feed-forward network with its own weights. A router reads each token's hidden-state vector, scores the available experts, and selects the top two. It also produces a routing weight for each selection. Only those selected experts process the token. The model therefore contains many expert weights while executing only a small subset for each token. The serving system must move activations when the selected experts are stored on different GPUs. The visual follows a batch that begins on GPU 0. The router returns two expert IDs and two routing weights per token. One selected path is shown for each example token to keep the diagram readable. Token T1 selects E1 on GPU 0. Its activation remains in local GPU memory, where E1 processes it. T2 selects E4 on GPU 1 in the same server. GPU 0 transfers the activation through NVLink, a high-bandwidth GPU-to-GPU connection. NVSwitch connects several GPUs through this fabric. T3 selects E7 on GPU 2 in another server. Its activation crosses the cluster network through InfiniBand or Ethernet. To be clear, the system transfers activation vectors, which are the numerical representations produced for a token by the preceding Transformer operations. Each destination GPU runs its expert and returns the output to GPU 0. GPU 0 multiplies both selected expert outputs by their routing weights, adds them together, and restores the original token order. Different tokens can select different expert pairs. Across a batch, those assignments may involve every GPU storing experts. Each GPU may therefore send activations to several GPUs while receiving activations for its own experts. This exchange is called all-to-all communication. And there are several ways to optimize this. For instance, keeping frequently selected experts within the same server reduces remote traffic. Balanced routing prevents one GPU from delaying the layer. Communication overlap allows local expert computation to continue while remote activations move. In this setup, sparse routing reduces expert FLOPs, but token dispatch and network communication erase part of that saving. If you want to dive deeper, I recently covered MoE inference engineering, including routing, expert placement, token dispatch, load balancing, communication overlap, memory, quantization, and offloading. Read my article below.
Original Article
View Cached Full Text

Cached at: 09/27/26, 01:21 PM

A good technical LLM interview question:

You replace a dense feed-forward layer with a top-2 MoE.

The profiler confirms that the MoE executes fewer FLOPs per token.

However, end-to-end inference latency has increased instead of decreasing.

Why did this happen?

(answer below)

A standard Transformer applies the same feed-forward network to every token in a layer.

An MoE layer replaces that network with multiple experts. Each expert is a feed-forward network with its own weights.

A router reads each token’s hidden-state vector, scores the available experts, and selects the top two. It also produces a routing weight for each selection.

Only those selected experts process the token. The model therefore contains many expert weights while executing only a small subset for each token.

The serving system must move activations when the selected experts are stored on different GPUs.

The visual follows a batch that begins on GPU 0.

The router returns two expert IDs and two routing weights per token. One selected path is shown for each example token to keep the diagram readable.

Token T1 selects E1 on GPU 0. Its activation remains in local GPU memory, where E1 processes it.

T2 selects E4 on GPU 1 in the same server. GPU 0 transfers the activation through NVLink, a high-bandwidth GPU-to-GPU connection. NVSwitch connects several GPUs through this fabric.

T3 selects E7 on GPU 2 in another server. Its activation crosses the cluster network through InfiniBand or Ethernet.

To be clear, the system transfers activation vectors, which are the numerical representations produced for a token by the preceding Transformer operations.

Each destination GPU runs its expert and returns the output to GPU 0.

GPU 0 multiplies both selected expert outputs by their routing weights, adds them together, and restores the original token order.

Different tokens can select different expert pairs. Across a batch, those assignments may involve every GPU storing experts.

Each GPU may therefore send activations to several GPUs while receiving activations for its own experts. This exchange is called all-to-all communication.

And there are several ways to optimize this.

For instance, keeping frequently selected experts within the same server reduces remote traffic.

Balanced routing prevents one GPU from delaying the layer.

Communication overlap allows local expert computation to continue while remote activations move.

In this setup, sparse routing reduces expert FLOPs, but token dispatch and network communication erase part of that saving.

If you want to dive deeper, I recently covered MoE inference engineering, including routing, expert placement, token dispatch, load balancing, communication overlap, memory, quantization, and offloading.

Read my article below.

Similar Articles

Multi Tier MoE Caching

Reddit r/LocalLLaMA

Discusses multi-tier caching strategies for MoE models to improve inference speed by keeping frequently activated experts on GPU, referencing existing implementations like PowerInfer and llama.cpp branches.