Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

arXiv cs.LG Papers

Summary

This paper introduces elbow-based routing, a training-free inference-time method for MoE models that dynamically adjusts the number of active experts per token by detecting the elbow point in router probability distributions, achieving a 5.3% average latency reduction while maintaining accuracy.

arXiv:2608.04401v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models enable model scaling while maintaining low inference-time compute by activating only a subset of experts per token. However, conventional routing relies on a fixed top-k selection, forcing the model to spend the same compute regardless of how many experts are relevant. We introduce elbow-based routing, a training-free inference-time modification that dynamically adjusts the number of experts on a per-token basis. Our method examines the sorted router probability distribution and identifies an elbow point that separates high- and low-probability experts. We find that most router distributions exhibit clear inflection points suitable for this strategy, and we show both theoretically and empirically that elbow-based routing preserves expert load balance. Experiments on a state-of-the-art MoE model demonstrate an average latency reduction of 5.3% while maintaining accuracy across six benchmarks.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:49 AM

# Elbow-based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection
Source: [https://arxiv.org/html/2608.04401](https://arxiv.org/html/2608.04401)
Robin Pan1Raymond Liu1†Daniel Fang1Adelina Andrei2Rosa Wu1 1John A\. Paulson School of Engineering and Applied Sciences, Harvard University 2Department of Mathematics, Harvard University

###### Abstract

Mixture\-of\-Experts \(MoE\) models enable model scaling while maintaining low inference\-time compute by activating only a subset of experts per token\. However, conventional routing relies on a fixed top\-k selection, forcing the model to spend the same compute regardless of how many experts are relevant\. We introduce elbow\-based routing, a training\-free inference\-time modification that dynamically adjusts the number of experts on a per\-token basis\. Our method examines the sorted router probability distribution and identifies an elbow point that separates high\- and low\-probability experts\. We find that most router distributions exhibit clear inflection points suitable for this strategy, and we show both theoretically and empirically that elbow\-based routing preserves expert load balance\. Experiments on a state\-of\-the\-art MoE model demonstrate an average latency reduction of 5\.3% while maintaining accuracy across six benchmarks\.

## 1Introduction

Mixture\-of\-Experts \(MoE\) architectures scale language models efficiently by activating only a subset of experts per token, enabling models to grow in size without proportionally increasing inference cost \(Shazeeret al\.\([2017](https://arxiv.org/html/2608.04401#bib.bib6)\)\)\. During both training and inference, MoE model routers activate the top\-kkexperts with the highest logits across all tokens regardless of token\. Consequently, some tokens with only a few relevant experts consume unnecessary compute, while tokens requiring more experts may be under\-routed\.

Prior works in dynamic routing require either training a router, expensive hyperparameter searches, or both\.Huanget al\.\([2024](https://arxiv.org/html/2608.04401#bib.bib1)\)demonstrates that some inputs benefit from increased expert capacity, and uses auxiliary losses to train a top\-pprouter based on a threshold probability masspp\. However,ppis a hyperparameter that must be tuned via grid search, which is an expensive process and may be sensitive to biases in the validation data\. DynMoEGuoet al\.\([2025](https://arxiv.org/html/2608.04401#bib.bib2)\)treats routing as a multi\-label classification problem with each expert as a label and allows top\-any, but also requires router training\. This motivates our central question:Can MoE models dynamically adjust the number of active experts per token to reduce latency without retraining or additional hyperparameter tuning?

We introduce elbow\-based routing, a simple inference\-time plugin that adaptively selectskkby detecting the elbow point in the router’s sorted probability curve \(Fig\.[1](https://arxiv.org/html/2608.04401#S2.F1)\), which allows us to reduce compute without significantly affecting accuracy\. We evaluate our approach on OLMoE, a state\-of\-the\-art MoE modelMuennighoffet al\.\([2025](https://arxiv.org/html/2608.04401#bib.bib4)\)\. Our contributions are as follows:

1. 1\.Our work introduces the first training\-free and inference\-time plugin for dynamic routing that does not require altering model weights, architecture, or any additional losses\.
2. 2\.We conduct an analysis of router probability distributions and find that the vast majority of router probabilities have distinct, sharp elbows and this characteristic is largely independent of input type or router layer\.
3. 3\.We conduct a comprehensive empirical evaluation across MMLU, ARC\-Easy, ARC\-Challenge, HellaSwag, PIQA, and WinoGrande, demonstrating an average latency reduction of 5\.3% while maintaining accuracy and having minimal effect on load balancing\.

## 2Elbow\-Based Routing

![Refer to caption](https://arxiv.org/html/2608.04401v1/x1.png)Figure 1:Elbow\-based routing\. Our method adapts to each token’s probability curve, which allows us to reduce computation when a clear separation exists between high\-confidence and low\-confidence experts\. Sharper elbows \(left\) select fewer experts, and more gradual elbows \(right\) select more experts\.We consider a mixture\-of\-experts model withNNexperts and a learned router that produces logitsℓ​\(x\)∈ℝN\\ell\(x\)\\in\\mathbb\{R\}^\{N\}for each tokenxx\. Applying a softmax yields router probabilitiesp​\(x\)=softmax​\(ℓ​\(x\)\),p\(x\)=\\mathrm\{softmax\}\(\\ell\(x\)\),which we sort in descending order asp\(1\)​\(x\)≥⋯≥p\(N\)​\(x\)p\_\{\(1\)\}\(x\)\\geq\\dots\\geq p\_\{\(N\)\}\(x\)\. For each token, we compute an elbow indexe​\(x\)e\(x\)on the full sorted probability vector\{p\(i\)​\(x\)\}i=1N\\\{p\_\{\(i\)\}\(x\)\\\}\_\{i=1\}^\{N\}\. The elbow is identified using a Kneedle\-styleSatopaaet al\.\([2011](https://arxiv.org/html/2608.04401#bib.bib5)\)criterion that selects the index of maximum deviation from a reference line after normalizing indices and probabilities to\[0,1\]\[0,1\]\(Algorithm[1](https://arxiv.org/html/2608.04401#alg1)\)\. Intuitively, the elbow corresponds to the index at which the rapidly decaying head of the distribution transitions into a relatively flat tail\.

### 2\.1Capped Expert Selection

The final number of active expertsk​\(x\)k\(x\)isk​\(x\)=min⁡\(e​\(x\),K\)k\(x\)=\\min\\bigl\(e\(x\),K\\bigr\), whereKKis the number of experts activated in standard top\-KKrouting\. Given thatk​\(x\)≤Kk\(x\)\\leq Kby construction, elbow\-based routing is a monotone pruning rule on top of an existing router\. It preserves expert ordering, never introduces new experts, and requires no retraining or additional routing hyperparameters beyond the model’s fixed top\-KKconstraint\. Elbow\-based routing is fast, with a time complexity of𝒪​\(N​log⁡N\)\\mathcal\{O\}\(N\\log N\)whereNN, the number of experts, is typically on the order of 10 to 100\.

### 2\.2Signal–Noise Interpretation

![Refer to caption](https://arxiv.org/html/2608.04401v1/rand2.png)Figure 2:Tail randomization confirm signal\-noise interpretation\.Randomly exchanging tail experts with indices in\[k​\(x\),8\)\[k\(x\),8\)with experts in\[k​\(x\),64\)\[k\(x\),64\)does not affect accuracy\. Accuracy drops when randomizing experts in\[0,k​\(x\)\)\[0,k\(x\)\)\.The effectiveness of elbow\-based routing can be understood through a signal–noise interpretation of the router logits\. An input vectorxxto a router may be decomposed asx=x∥\+x⟂x=x\_\{\\parallel\}\+x\_\{\\perp\}, wherex∥x\_\{\\parallel\}denotes components aligned with routing\-relevant directions, andx⟂x\_\{\\perp\}captures residual variation orthogonal to these directions\. Relevant ’signal’ experts aligned withx∥x\_\{\\parallel\}receive large, structured logits that form the head of the sorted routing distribution, while less relevant ’noise’ experts influenced mainly byx⟂x\_\{\\perp\}exhibit small, unstructured variations, yielding a flat tail\.

Under this view, the sorted router distribution naturally exhibits a transition between a small set of high\-confidence experts and a diffuse tail\. The elbow identifies this transition point, providing a token\-specific estimate of how many experts are meaningfully engaged by the router\. When such a transition is pronounced, the elbow corresponds to a sharp change in slope, which is reflected geometrically by a large elbow angle\. In Section[2\.3](https://arxiv.org/html/2608.04401#S2.SS3)we quantify this effect by measuring elbow angles across layers and tokens, and we show that the vast majority of routing distributions exhibit sharp elbows, consistent with this signal–noise interpretation\.

We empirically confirm this interpretation by evaluating model accuracy when randomizing post\-elbow experts\. Given a router probability curve withk​\(x\)<8k\(x\)<8, we replace post\-elbow experts \(ranksk​\(x\)k\(x\)to 8\) with randomly chosen experts of rank≥k​\(x\)\\geq k\(x\)\. As shown in Fig\.[2](https://arxiv.org/html/2608.04401#S2.F2), this tail randomization matches the baseline top\-8 accuracy with no randomization on MMLU\. However, accuracy decreases when replacing pre\-elbow experts \(ranks 1 tok​\(x\)k\(x\)\) in the same way\. This confirms that the elbow marks a practical signal–noise boundary for routing\.

### 2\.3Elbow Analysis and Characterization

We analyzed the ‘elbow\-ness’ of router probabilities from OlMoE routers across multiple benchmarks\. To characterize ‘elbow\-ness’ we calculate an elbow angleθ\\thetabetween the first normalized point\(0,0\)\(0,0\), the elbow point\(xe′,pe′\)\(x^\{\\prime\}\_\{e\},p^\{\\prime\}\_\{e\}\), and the last normalized point\(1,1\)\(1,1\)\.

θ\\displaystyle\\theta=arccos⁡\(𝐯1⋅𝐯2‖𝐯1‖​‖𝐯2‖\),\\displaystyle=\\arccos\\left\(\\frac\{\\mathbf\{v\}\_\{1\}\\cdot\\mathbf\{v\}\_\{2\}\}\{\\\|\\mathbf\{v\}\_\{1\}\\\|\\,\\\|\\mathbf\{v\}\_\{2\}\\\|\}\\right\),\(1\)where𝐯1\\displaystyle\\text\{where\}\\quad\\mathbf\{v\}\_\{1\}=\(−xe′−pe′\),𝐯2=\(1−xe′1−pe′\)\.\\displaystyle=\\begin\{pmatrix\}\-x^\{\\prime\}\_\{e\}\\\\ \-p^\{\\prime\}\_\{e\}\\end\{pmatrix\},\\qquad\\mathbf\{v\}\_\{2\}=\\begin\{pmatrix\}1\-x^\{\\prime\}\_\{e\}\\\\ 1\-p^\{\\prime\}\_\{e\}\\end\{pmatrix\}\.\(2\)We analyzed over 2 million router probability curves that were produced using samples from MMLU, ARC\-Easy, and ARC\-Challenge and found that 99\.7% of them possess clear, sharp elbows \(≤135​°\\leq 135\\degree\) \(Fig\.[3](https://arxiv.org/html/2608.04401#S2.F3)a\)\.

![Refer to caption](https://arxiv.org/html/2608.04401v1/x2.png)

Figure 3:Analysis of router probability distributions\. \(a\) Elbow angle distribution showing 99\.7% of cases exhibit sharp elbows \(≤135​°\\leq$$\)\. \(b\) Mean sorted router probabilities across three benchmark datasets demonstrate consistent router logit patterns\. \(c\) Elbow angle distributions by dataset show minimal variation across datasets\. \(d\) Mean elbow angles remain stable across all 16 router layers for all three datasets, indicating that elbow structure is a robust, layer\-invariant property\.Separating these probabilities by dataset, we see that router probabilities and elbow angles of these probabilities are nearly identical across datasets \(Fig\.[3](https://arxiv.org/html/2608.04401#S2.F3)b\-c\)\. Additionally, mean elbow angles are consistent across router layers and across datasets \(Fig\.[3](https://arxiv.org/html/2608.04401#S2.F3)d\)\. Because the MMLU dataset is categorized by subject, we analyzed router elbow angles and index by subject and found very similar elbow angle and elbow index distributions across all subjects \(Supplementary Fig\.[5](https://arxiv.org/html/2608.04401#A1.F5)\)\.

### 2\.4Load\-Balance Guarantees

For each token, elbow\-based routing selects a subset of the experts chosen by standard top\-KKrouting, reducing the number of active experts without introducing new paths or changing relative ordering\. In Appendix[A\.3](https://arxiv.org/html/2608.04401#A1.SS3), we formalize the effect of this pruning on expert load and show that the induced changes to the normalized per\-expert utilization distribution are bounded\.

LetLi\(top​\-​K\)L\_\{i\}^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}andLi\(elb\)L\_\{i\}^\{\(\\mathrm\{elb\}\)\}denote the expected load of expertiiunder top\-KKand elbow\-based routing, respectively, and letδ=1−𝔼​\[k​\(x\)\]/K\\delta=1\-\\mathbb\{E\}\[k\(x\)\]/Kdenote the expected fraction of pruned expert assignments\. We prove that the normalized expert utilization distributionsq\(top​\-​K\)q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}andq\(elb\)q^\{\(\\mathrm\{elb\}\)\}satisfy

‖q\(elb\)−q\(top​\-​K\)‖1≤2​δ1−δ\.\\left\\lVert q^\{\(\\mathrm\{elb\}\)\}\-q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\\right\\rVert\_\{1\}\\;\\leq\\;\\frac\{2\\delta\}\{1\-\\delta\}\.
Table 1:Layer\-wise Load Balancing AnalysisWe measure expert load usage averaged across six benchmarks: MMLU, ARC\-Easy, ARC\-Challenge, HellaSwag, PIQA, and WinoGrande\. Across all benchmarks, elbow\-based routing consistently selects around𝔼​\[k​\(x\)\]≈7\.6\\mathbb\{E\}\[k\(x\)\]\\approx 7\.6experts per token across all layers as seen in Table[2](https://arxiv.org/html/2608.04401#S2.T2), which corresponds to an average pruning rateδ=1−𝔼​\[k​\(x\)\]K=0\.05\\delta=1\-\\frac\{\\mathbb\{E\}\[k\(x\)\]\}\{K\}=0\.05\. Empirically, we find that changes to load balance are much smaller than the worst\-case theoretical bounds at thisδ\\delta\. For each layer, we calculate theℓ1\\ell\_\{1\}distance between the normalized expert utilization distributions under top\-KKand elbow\-based\-KKrouting, averaged across all six benchmarks\. Table[1](https://arxiv.org/html/2608.04401#S2.T1)reports a mean change of0\.0340\.034in theℓ1\\ell\_\{1\}norm, which suggests that only1\.7%1\.7\\%of the total normalized utilization mass is redistributed across experts\. We calculate%\\%Change in Top\-1 Share asmax⁡\(q\(t​o​p−K\)\)−max⁡\(q\(e​l​b\)\)max\(q\(t​o​p−K\)⋅100\\frac\{\\max\(q^\{\(top\-K\)\}\)\-\\max\(q^\{\(elb\)\}\)\}\{\\max\(q^\{\(top\-K\)\}\}\\cdot 100and find the maximum expert load increases by2\.16%2\.16\\%on average\.%\\%Change in CV is calculated asC​Vt​o​p−K−C​Ve​l​bC​Vt​o​p−K⋅100\\frac\{CV^\{top\-K\}\-CV^\{elb\}\}\{CV^\{top\-K\}\}\\cdot 100, and CV increases by2\.83%2\.83\\%\. For all three metrics, empirical evidence is significantly smaller than theoretical guarantees in[A\.3](https://arxiv.org/html/2608.04401#A1.SS3)\.

Table 2:Evaluation ResultsDatasetModelAcc \(%\)k\-meanFLOPsLatency \(ms\)ARC\-EasyBase OLMoE77\.8285\.37E811\.203Elbow\-878\.247\.6234\.86E810\.497ARC\-ChallengeBase OLMoE61\.8685\.37E811\.531Elbow\-862\.297\.6344\.88E810\.974MMLUBase OLMoE51\.7485\.37E813\.287Elbow\-851\.757\.6134\.91E812\.571HellaSwagBase OLMoE47\.0285\.37E814\.630Elbow\-846\.867\.6554\.92E813\.960PIQABase OLMoE73\.1885\.37E813\.398Elbow\-872\.857\.5894\.85E812\.944WinoGrandeBase OLMoE50\.3685\.37E89\.659Elbow\-850\.287\.5764\.81E88\.948

## 3Experimental Validation

We evaluate elbow\-based routing on OLMoE under baseline top\-88routing and elbow\-based routing capped atk≤8k\\leq 8\(Elbow\-88\) across six standard multiple\-choice benchmarks: MMLU, ARC\-Easy, ARC\-Challenge, HellaSwag, PIQA, and WinoGrande\. Table[2](https://arxiv.org/html/2608.04401#S2.T2)reports accuracy, the mean number of active experts per token \(kk\-mean\), estimated FLOPs, and wall\-clock latency per forward pass\. See[A\.4](https://arxiv.org/html/2608.04401#A1.SS4)for additional implementation details\. Across datasets, Elbow\-88matches baseline accuracy \(within≤0\.33\\leq 0\.33points\) while reducing average latency by5\.3%5\.3\\%on our setup and decreasing estimated compute, with an overallkk\-mean of7\.6157\.615\. These results indicate that elbow\-based routing provides a lightweight inference\-time efficiency improvement without retraining or changes to model weights\.

## 4Conclusion

We introduced elbow\-based routing, a training\-free inference\-time plugin that selects a token\-specific number of experts by detecting an elbow in the router’s sorted probability curve\. On OLMoE with a cap ofK=8K=8, this simple monotone pruning rule reduces average latency by5\.3%5\.3\\%across six benchmarks while preserving accuracy, without retraining, architectural changes, or additional hyperparameters\. Across over 2M routing decisions, we observe that router distributions exhibit sharp and consistent elbows across layers and datasets, supporting a head–tail structure in which a small set of experts carries most routing signal\. We further provide load\-balance guarantees and empirical evidence that the induced changes in expert utilization are small, suggesting that routers trained with fixed top\-KKoften contain enough structure to enable lightweight test\-time efficiency improvements and serve as a diagnostic lens for routing behavior\.

## References

- Y\. Guo, Z\. Cheng, X\. Tang, Z\. Tu, and T\. Lin \(2025\)Dynamic mixture of experts: an auto\-tuning approach for efficient transformer models\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,Cited by:[§1](https://arxiv.org/html/2608.04401#S1.p2.3)\.
- Q\. Huang, Z\. An, N\. Zhuang, M\. Tao, C\. Zhang, Y\. Jin, K\. Xu, L\. Chen, S\. Huang, and Y\. Feng \(2024\)Harder tasks need more experts: dynamic routing in moe models\.InProceedings of the Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.04401#S1.p2.3)\.
- N\. Muennighoff, L\. Soldaini, D\. Groeneveld, K\. Lo, J\. Morrison, S\. Min, W\. Shi, E\. P\. Walsh, O\. Tafjord, N\. Lambert, Y\. Gu, S\. Arora, A\. Bhagia, D\. Schwenk, D\. Wadden, A\. Wettig, B\. Hui, T\. Dettmers, D\. Kiela, A\. Farhadi, and et al\. \(2025\)OLMoE: open mixture\-of\-experts language models\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,Cited by:[§A\.4](https://arxiv.org/html/2608.04401#A1.SS4.p1.1),[§1](https://arxiv.org/html/2608.04401#S1.p3.1)\.
- V\. Satopaa, J\. Albrecht, D\. Irwin, and B\. Raghavan \(2011\)Finding a ”kneedle” in a haystack: detecting knee points in system behavior\.InProceedings of the 2011 31st International Conference on Distributed Computing Systems Workshops,ICDCSW ’11,USA,pp\. 166–171\.Cited by:[§2](https://arxiv.org/html/2608.04401#S2.p1.8)\.
- N\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. Le, G\. Hinton, and J\. Dean \(2017\)Outrageously large neural networks: the sparsely\-gated mixture\-of\-experts layer\.External Links:1701\.06538,[Link](https://arxiv.org/abs/1701.06538)Cited by:[§1](https://arxiv.org/html/2608.04401#S1.p1.1)\.

## Appendix AAppendix

### A\.1Algorithm for Determining Elbows

Below is an approximation of the Kneedle algorithm used to determine elbow indices:

Algorithm 1Elbow Detection via Kneedle ApproximationRouter logits

l∈ℝNl\\in\\mathbb\{R\}^\{N\}over

NNexperts

p←softmax​\(l\)p\\leftarrow\\mathrm\{softmax\}\(l\)

p←sort↓​\(p\)p\\leftarrow\\mathrm\{sort\}\_\{\\downarrow\}\(p\)⊳\\trianglerightsort probabilities in descending order

pmin←min⁡\(p\)p\_\{\\min\}\\leftarrow\\min\(p\);

pmax←max⁡\(p\)p\_\{\\max\}\\leftarrow\\max\(p\)
for

i=0i=0to

N−1N\-1do⊳\\trianglerightnormalize to unit square and compute deviation

pi′←pi−pminpmax−pmin\+ϵp^\{\\prime\}\_\{i\}\\leftarrow\\dfrac\{p\_\{i\}\-p\_\{\\min\}\}\{p\_\{\\max\}\-p\_\{\\min\}\+\\epsilon\}

xi′←iN−1x^\{\\prime\}\_\{i\}\\leftarrow\\dfrac\{i\}\{N\-1\}

Di←pi′−xi′D\_\{i\}\\leftarrow p^\{\\prime\}\_\{i\}\-x^\{\\prime\}\_\{i\}

endfor

e←arg⁡maxi∈\{0,…,N−1\}⁡Die\\leftarrow\\arg\\max\_\{i\\in\\\{0,\\dots,N\-1\\\}\}D\_\{i\}

return

k←e\+1k\\leftarrow e\+1

### A\.2Additional Elbow Characterization

#### A\.2\.1Sample Elbow Logits

![Refer to caption](https://arxiv.org/html/2608.04401v1/samplelogit.png)Figure 4:Plots of sorted router probabilities with elbow indexes atk=2,4,6,8,10,12k=2,4,6,8,10,12and their respective elbow angles\. Our implementation of the Kneedle algorithm effectively identifies the ’elbow’ of sorted router probability curves\.
#### A\.2\.2Elbow Analysis by MMLU Subject

![Refer to caption](https://arxiv.org/html/2608.04401v1/elbowsbysubject.png)

Figure 5:Elbow angle and index vary little across subjects\.We plot the distribution of elbow index and elbow angle by subject above and conduct a one\-way ANOVA test and findp<0\.001,η2=0\.02p<0\.001,\\eta^\{2\}=0\.02for elbow indices andp<0\.001,η2=0\.01p<0\.001,\\eta^\{2\}=0\.01withN=1,704,247N=1,704,247\. This indicates that a very small fraction of variance is attributed to subject\.
#### A\.2\.3Elbow Angle and Elbow Index Correlation

![Refer to caption](https://arxiv.org/html/2608.04401v1/anglevindexcorrelation.png)Figure 6:Elbow indexes are correlated with elbow angle\.Elbow angle increases with elbow index, indicating that tokens requiring more active experts tend to exhibit less sharp transitions in the sorted router distribution \(Pearsonr=0\.58r=0\.58\)

### A\.3Additional Load\-Balance Guarantees Details

We analyze the impact of elbow\-KKrouting on expert load balance relative to standard top\-KKrouting\. Elbow\-KKrouting acts as a monotone pruning of top\-KKrouting: for each token, it selects a subset of the experts chosen by top\-KK\. Under this structure, we show that changes in distribution of the load balance are bounded and minimal\.

Notation\.Let𝒮top​\-​K​\(x\)\\mathcal\{S\}\_\{\\mathrm\{top\}\\text\{\-\}K\}\(x\)and𝒮elb​\-​K​\(x\)\\mathcal\{S\}\_\{\\mathrm\{elb\}\\text\{\-\}K\}\(x\)denote the experts selected by standard top\-KKrouting and elbow\-KKrouting, respectively, for tokenxx\. Letk​\(x\)=\|𝒮elb​\-​K​\(x\)\|k\(x\)=\|\\mathcal\{S\}\_\{\\mathrm\{elb\}\\text\{\-\}K\}\(x\)\|denote the number of experts selected by elbow\-KKrouting for tokenxx, and define the expected fraction of pruned assignments as

δ=1−𝔼​\[k​\(x\)\]K\.\\delta=1\-\\frac\{\\mathbb\{E\}\[k\(x\)\]\}\{K\}\.
For a routing ruler∈\{top​\-​K,elb​\-​K\}r\\in\\\{\\mathrm\{top\}\\text\{\-\}K,\\,\\mathrm\{elb\}\\text\{\-\}K\\\}, define the expected per\-expert load

Li\(r\)=𝔼x∼𝒟​\[𝟏​\{i∈𝒮r​\(x\)\}\],L\_\{i\}^\{\(r\)\}=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}\\\!\\left\[\\mathbf\{1\}\\\{i\\in\\mathcal\{S\}\_\{r\}\(x\)\\\}\\right\],and the corresponding normalized utilization distribution

qi\(r\)=Li\(r\)∑j=1NLj\(r\),q\(r\)∈ΔN−1\.q\_\{i\}^\{\(r\)\}=\\frac\{L\_\{i\}^\{\(r\)\}\}\{\\sum\_\{j=1\}^\{N\}L\_\{j\}^\{\(r\)\}\},\\qquad q^\{\(r\)\}\\in\\Delta^\{N\-1\}\.
###### Theorem 1\(Distribution Changes in Expert Utilization under Elbow\-KKRouting\)\.

For any MoE layer, letq\(top​\-​K\)q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}andq\(elb​\-​K\)q^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}denote the normalized per\-expert utilization distributions under top\-KKand elbow\-KKrouting, respectively\. Then, we claim that

‖q\(elb​\-​K\)−q\(top​\-​K\)‖1≤2​δ1−δ\.\\bigl\\\|q^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}\-q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\\bigr\\\|\_\{1\}\\leq\\frac\{2\\delta\}\{1\-\\delta\}\.

###### Proof\.

By construction,𝒮elb​\-​K​\(x\)⊆𝒮top​\-​K​\(x\)\\mathcal\{S\}\_\{\\mathrm\{elb\}\\text\{\-\}K\}\(x\)\\subseteq\\mathcal\{S\}\_\{\\mathrm\{top\}\\text\{\-\}K\}\(x\)for all tokensxx\. Therefore, working at the vector level, define the removed\-mass vectorR=L\(top​\-​K\)−L\(elb​\-​K\)∈ℝ\+N\.R=L^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\-L^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}\\in\\mathbb\{R\}^\{N\}\_\{\+\}\.

Since top\-KKselects exactlyKKexperts per token and elbow\-KKselectsk​\(x\)k\(x\)experts, we have

∑i=1NLi\(top​\-​K\)=K,∑i=1NLi\(elb​\-​K\)=𝔼​\[k​\(x\)\]=\(1−δ\)​K,∑i=1NRi=δ​K\.\\sum\_\{i=1\}^\{N\}L\_\{i\}^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}=K,\\qquad\\sum\_\{i=1\}^\{N\}L\_\{i\}^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}=\\mathbb\{E\}\[k\(x\)\]=\(1\-\\delta\)K,\\qquad\\sum\_\{i=1\}^\{N\}R\_\{i\}=\\delta K\.
Passing to the corresponding normalized removed\-mass distribution we have

r=Rδ​K,r∈ΔN−1,r=\\frac\{R\}\{\\delta K\},\\qquad r\\in\\Delta^\{N\-1\},which combined with the definitions of the normalized utilization distributions gives

q\(elb​\-​K\)=L\(elb​\-​K\)\(1−δ\)​K=L\(top​\-​K\)−R\(1−δ\)​K=q\(top​\-​K\)​K−R\(1−δ\)​K=q\(top​\-​K\)−r​δ1−δ⟹q^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}=\\frac\{L^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}\}\{\(1\-\\delta\)K\}=\\frac\{L^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\-R\}\{\(1\-\\delta\)K\}=\\frac\{q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}K\-R\}\{\(1\-\\delta\)K\}=\\frac\{q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\-r\\delta\}\{1\-\\delta\}\\implies⟹q\(elb​\-​K\)−q\(top​\-​K\)=δ1−δ​\(q\(top​\-​K\)−r\)\.\\implies q^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}\-q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}=\\frac\{\\delta\}\{1\-\\delta\}\\bigl\(q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\-r\\bigr\)\.
Sinceq\(top​\-​K\)q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}andrrare probability distributions, we have that worst\-case‖q\(top​\-​K\)−r‖1≤∑i=1N\|qi\(top​\-​K\)−ri\|≤∑i=1Nqi\(top​\-​K\)\+∑i=1Nri≤2\\\|q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\-r\\\|\_\{1\}\\leq\\sum\_\{i=1\}^\{N\}\|q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\_\{i\}\-r\_\{i\}\|\\leq\\sum\_\{i=1\}^\{N\}q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\_\{i\}\+\\sum\_\{i=1\}^\{N\}r\_\{i\}\\leq 2, which gives

‖q\(elb​\-​K\)−q\(top​\-​K\)‖1≤2​δ1−δ\.\\bigl\\\|q^\{\(\\mathrm\{elb\}\\text\{\-\}K\)\}\-q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\\bigr\\\|\_\{1\}\\leq\\frac\{2\\delta\}\{1\-\\delta\}\.∎

Corollary 1\.The above bound implies that the change in the load of the most\-utilized expert is

\|‖q\(elb−K\)‖∞−‖q\(top−K\)‖∞\|≤‖q\(elb−K\)−q\(top−K\)‖1≤2​δ1−δ\.\\big\|\\\|q^\{\(\\mathrm\{elb\}\-K\)\}\\\|\_\{\\infty\}\-\\\|q^\{\(\\mathrm\{top\}\-K\)\}\\\|\_\{\\infty\}\\big\|\\leq\\\|q^\{\(\\mathrm\{elb\}\-K\)\}\-q^\{\(\\mathrm\{top\}\-K\)\}\\\|\_\{1\}\\leq\\frac\{2\\delta\}\{1\-\\delta\}\.Thus, elbow\-KKrouting cannot substantially increase the load of the most\-utilized expert, which is often a determinant of inference\-time latency\.

Corollary 2\.The above bound implies that the coefficient of variationCV2​\(q\)=N​‖q‖22−1\\mathrm\{CV\}^\{2\}\(q\)=N\\\|q\\\|\_\{2\}^\{2\}\-1change is

\|CV2​\(q\(elb−K\)\)−CV2​\(q\(top​\-​K\)\)\|=N​\|‖q\(elb−K\)‖22−‖q\(top−K\)‖22\|≤\\Big\|\\mathrm\{CV\}^\{2\}\(q^\{\(\\mathrm\{elb\}\-K\)\}\)\-\\mathrm\{CV\}^\{2\}\(q^\{\(\\mathrm\{top\\text\{\-\}K\}\)\}\)\\Big\|=N\\left\|\\,\\\|q\\,^\{\{\(\\mathrm\{elb\}\-K\)\}\}\\\|\_\{2\}^\{2\}\-\\\|q\\,^\{\(\\mathrm\{top\-\}K\)\}\\\|^\{2\}\_\{2\}\\right\|\\leq=1\(1−δ\)2​\|\(2​δ−δ2\)​‖q\(top​\-​K\)‖22−2​δ​⟨q\(top​\-​K\),r⟩\+δ2​‖r‖22\|≤2​N​δ\(1−δ\)2,=\\frac\{1\}\{\(1\-\\delta\)^\{2\}\}\\Bigl\|\(2\\delta\-\\delta^\{2\}\)\\\|q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\\\|\_\{2\}^\{2\}\-2\\delta\\langle q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\},r\\rangle\+\\delta^\{2\}\\\|r\\\|\_\{2\}^\{2\}\\Bigr\|\\leq\\frac\{2N\\delta\}\{\(1\-\\delta\)^\{2\}\},where we used that0≤‖q\(top​\-​K\)‖22≤10\\leq\\\|q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\\\|\_\{2\}^\{2\}\\leq 1,0≤‖r‖22≤10\\leq\\\|r\\\|\_\{2\}^\{2\}\\leq 1, and0≤⟨q\(top​\-​K\),r⟩≤10\\leq\\langle q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\},r\\rangle\\leq 1sinceq\(top​\-​K\),r∈ΔN−1q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\},r\\in\\Delta^\{N\-1\}to get0≤\(2​δ−δ2\)​‖q\(top​\-​K\)‖22\+δ2​‖r‖22≤2​δ0\\leq\(2\\delta\-\\delta^\{2\}\)\\\|q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\}\\\|\_\{2\}^\{2\}\+\\delta^\{2\}\\\|r\\\|\_\{2\}^\{2\}\\leq 2\\delta\.

The alignment term⟨q\(top​\-​K\),r⟩\\langle q^\{\(\\mathrm\{top\}\\text\{\-\}K\)\},r\\ranglecaptures whether pruning removes assignments primarily from heavily utilized experts as opposed to low\-utilized ones\. Pruning aligned with high\-load experts tends to reduceCV2\\mathrm\{CV\}^\{2\}, while the opposite pattern can increase it\.

Discussion\.As we have mentioned in Section[2\.4](https://arxiv.org/html/2608.04401#S2.SS4), empirically we observed across the six benchmarks thatδ=1−𝔼​\[k​\(x\)\]K=0\.05\\delta=1\-\\frac\{\\mathbb\{E\}\[k\(x\)\]\}\{K\}=0\.05\.

Using this value, Theorem 1 guarantees that the normalized expert utilization distribution changes by at most

‖q\(elb​\-​K\)−q\(top​\-​K\)‖1≤2​δ1−δ≈0\.10\.\\\|q^\{\(\\mathrm\{elb\\text\{\-\}K\}\)\}\-q^\{\(\\mathrm\{top\\text\{\-\}K\}\)\}\\\|\_\{1\}\\;\\leq\\;\\frac\{2\\delta\}\{1\-\\delta\}\\;\\approx\\;0\.10\.This tells us that after pruning, no more than worst\-case 5% of the total normalized expert utilization mass can be redistributed across experts\. Empirically, the measuredℓ1\\ell\_\{1\}changes in normalized utilization are around0\.050\.05or less across layers \(Table[1](https://arxiv.org/html/2608.04401#S2.T1)\), indicating that the relative expert usage profiles are preserved even more closely than required by the guarantee\.

Similarly, Corollary 2 bounds the change in squared coefficient of variation as

\|C​V2​\(q\(elb​\-​K\)\)−C​V2​\(q\(top​\-​K\)\)\|≤2​N​δ\(1−δ\)2≈7\.\\bigl\|CV^\{2\}\(q^\{\(\\mathrm\{elb\\text\{\-\}K\}\)\}\)\-CV^\{2\}\(q^\{\(\\mathrm\{top\\text\{\-\}K\}\)\}\)\\bigr\|\\;\\leq\\;\\frac\{2N\\delta\}\{\(1\-\\delta\)^\{2\}\}\\;\\approx\\;7\.This is a rather loose bound\. Empirically, the observed changes in CV vary from11to8%8\\%across layers \(Table[1](https://arxiv.org/html/2608.04401#S2.T1)\), indicating that elbow\-based routing preserves not only the total utilization distribution but also its variability structure significantly more closely than required by the guarantee\.

### A\.4Implementation Details

We implemented elbow\-based routing on OLMoE\-1B\-7B\-0924\-Instruct \(Muennighoffet al\.\([2025](https://arxiv.org/html/2608.04401#bib.bib4)\)\), a sparse MoE model with 64 experts and default top\-8 routing\. We select this model for its state\-of\-the\-art performance, as well as its small model size given computational resource constraints\. All evaluations were performed with an NVIDIA H200 GPU\. We implement elbow\-based routing via a runtime monkey patch that replaces`OlmoeSparseMoeBlock\.forward`with an instrumented variant, while caching the original method for reversible restoration\.

FLOPs are estimated analytically per MoE block as the sum of \(i\) router cost—one dense linear projection \(2​N​d​E2NdE\), softmax \(5​N​E5NE\), and a top\-k/sort term \(N​E​log2⁡ENE\\log\_\{2\}E\)—and \(ii\) expert MLP cost—up and down projections \(2​N​k¯​d​m2N\\bar\{k\}dm\) each plus activation \(N​k¯​mN\\bar\{k\}m\), where N=batch×\\timesseq, d=hidden dim, m=4d, E=\# experts, andk¯\\bar\{k\}is the observed mean selected experts per token\. For dynamic routing we additionally account for elbow detection overhead as per\-token sort \(N​E​log2⁡ENE\\log\_\{2\}E\) plus linear\-time normalization, subtraction, and argmax terms \(≈\(4\+1\+1\)​N​E\\approx\(4\+1\+1\)NE\)\.

Latency was measured with wall\-clock timing around the MoE block forward pass using`time\.perf\_counter\(\)`, with an explicit`torch\.cuda\.synchronize\(\)`immediately before starting and after producing the final output tensor to ensure all queued CUDA kernels completed\.

Similar Articles

Sticky Routing: Training MoE Models for Memory-Efficient Inference

arXiv cs.LG

StickyMoE proposes a differentiable routing consistency loss that encourages adjacent tokens to activate the same experts in MoE models, reducing expert-swapping overhead and cache misses during inference on edge devices by up to 3.92× while improving perplexity.

Expert Routing for Communication-Efficient MoE via Finite Expert Banks

arXiv cs.LG

The paper introduces an information-theoretic framework for communication-efficient expert routing in sparse mixture-of-experts models, treating the gate as a stochastic channel and deriving practical mutual information estimators to analyze accuracy-rate tradeoffs over finite expert banks.

EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs

arXiv cs.AI

EntropyMoE introduces an entropy-aware Mixture-of-Experts architecture for tokenizer-free LLMs, using dynamic byte patches as routing units to enable sparse conditional computation. Experiments show it achieves the lowest held-out bits-per-byte among baselines while maintaining downstream accuracy.