Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation
Summary
This paper proposes UniF-MoE, a unified framework for token-adaptive Mixture-of-Experts computation that first shares reusable computation across experts and then routes the remaining residual demand, improving performance while reducing activated computation, latency, and memory on DomainBed and GLUE benchmarks.
View Cached Full Text
Cached at: 08/12/26, 08:29 AM
# A Unified Framework for Token-Adaptive MoE Computation
Source: [https://arxiv.org/html/2608.10392](https://arxiv.org/html/2608.10392)
## Share First, Route What Remains: A Unified Framework for Token\-Adaptive MoE Computation
###### Abstract
Mixture\-of\-experts \(MoE\) models have recently moved beyond routing a fixed number of complete experts\. Shared\-expert designs preserve reusable knowledge, fine\-grained methods vary computation within experts, and dynamic routers adapt the number of active experts\. Yet these decisions are usually made independently, overlooking a basic dependency: extracting reusable computation changes both what remains and how much expert capacity the remainder needs\. We study this dependency by decomposing sparsely upcycled feed\-forward experts into key\-value channels\. Co\-activated experts align at a subset of value positions; removing these positions changes expert preference; and greater shared coverage is associated with lower residual expert demand\. These observations lead to one principle: share first, then route what remains\. We instantiate it in UniF\-MoE, a unified framework for token\-adaptive MoE computation\. Each expert is partitioned into aligned blocks\. A shared\-demand score sets the shared block count and pathway weight, key prototypes select the shared content, and the complementary demand determines the residual expert count through cumulative routing mass\. A Gram regularizer separates and normalizes router embeddings, promoting diverse routing directions, sparse expert overlap, and a simple routing geometry\. Experiments on DomainBed and GLUE show that this unified design improves predictive performance over representative static and dynamic MoEs while reducing activated computation, inference latency, and memory\. Code is available at[https://github\.com/existence0420/UniF\-MoE](https://github.com/existence0420/UniF-MoE)\.
## IIntroduction
Mixture\-of\-experts \(MoE\) models increase parameter capacity without evaluating every parameter for every token\[[14](https://arxiv.org/html/2608.10392#bib.bib28)\]\. A router sends each token to a small set of feed\-forward networks \(FFNs\), or experts, so capacity can grow while execution remains sparse\[[28](https://arxiv.org/html/2608.10392#bib.bib1),[18](https://arxiv.org/html/2608.10392#bib.bib2)\]\. Conventional top\-kkrouting answers one question: which experts should process this token?
This direct formulation hides two allocation decisions\. It treats a complete expert as the unit of computation, even when selected experts repeat common responses, and it gives every token the same expert count, even when their semantic requirements differ\. The result can be redundant computation for simple tokens and insufficient capacity for harder ones\.
Recent work relaxes this formulation along three complementary directions\. Shared experts capture common knowledge\[[4](https://arxiv.org/html/2608.10392#bib.bib10),[36](https://arxiv.org/html/2608.10392#bib.bib36)\]; modular, nested or slimmable experts vary computation within an expert\[[26](https://arxiv.org/html/2608.10392#bib.bib7),[15](https://arxiv.org/html/2608.10392#bib.bib14),[29](https://arxiv.org/html/2608.10392#bib.bib15)\]; and dynamic routers vary the number of active experts\[[13](https://arxiv.org/html/2608.10392#bib.bib11),[12](https://arxiv.org/html/2608.10392#bib.bib12),[24](https://arxiv.org/html/2608.10392#bib.bib13),[21](https://arxiv.org/html/2608.10392#bib.bib37)\]\. Other methods encourage experts to remain distinct\[[11](https://arxiv.org/html/2608.10392#bib.bib35),[16](https://arxiv.org/html/2608.10392#bib.bib38)\]\. These advances are usually developed as separate mechanisms, but their decisions are not independent\. Once reusable computation is extracted, the residual content changes, the best experts can change, and the capacity needed from them can change as well\. A router that decides these quantities separately can therefore repeat shared work or allocate expert capacity for a response that has already been partly computed\.
To expose this dependency, we delve into the internal structure of the FFN\. An expert can be written as a sum of hidden channels, each defined by an up\-projection key and a down\-projection value\[[9](https://arxiv.org/html/2608.10392#bib.bib8)\]\. In a sparsely upcycled MoE\[[17](https://arxiv.org/html/2608.10392#bib.bib5)\], experts start from the same pretrained FFN, so matching channel positions remain comparable across experts\. This lets us test whether co\-activated experts produce reusable responses at aligned positions and what happens to routing after those responses are separated\.
Our analysis gives a consistent answer\. Co\-activated experts are neither wholly redundant nor wholly distinct: they align at a subset of value positions, and frequently co\-activated pairs tend to align more\. Removing these positions substantially changes the preferred experts\. The number of residual experts needed to recover the original output also varies across tokens and decreases as shared coverage grows\. Shared modeling, fine\-grained computation, and dynamic routing are therefore three stages of one allocation problem, not three parallel knobs\. The reusable response should be identified first; routing should then act on what remains\.
This principle leads to UniF\-MoE, a unified framework for token\-adaptive MoE computation\. Each layer contains one shared expert andKKresidual experts, all divided into aligned blocks\. For each token, a shared\-demand score first determines how much computation belongs to the shared pathway and how strongly that pathway contributes\. Key prototypes then select which blocks provide that reusable content\. Only the complementary blocks are routed onward, and their demand determines how many residual experts are activated by cumulative routing mass\. A single router thus coordinates three decisions in functional order:*how much to share*,*what to share*, and*how many experts the remainder needs*\. A Gram constraint shapes its embeddings into distinct, normalized directions, encouraging diverse expert roles and sparse overlap through a simple routing geometry\.
Our contributions are summarized as follows:
- •We reveal a routing\-conditioned dependency between reusable and token\-specific computation inside sparsely upcycled experts\. This turns shared modeling, fine\-grained computation, and dynamic routing into one ordered principle:*share first, then route what remains*\.
- •We instantiate this principle in UniF\-MoE, which couples*shared width*,*shared content*, and*residual expert count*through one token\-dependent budget\. The resulting process coordinates intra\-expert and inter\-expert sparsity instead of deciding them independently\.
- •Across vision and language benchmarks, UniF\-MoE achieves a stronger accuracy–efficiency trade\-off than representative MoEs\. Ablations and routing analyses further show that its gains arise from coordinating shared reuse with residual specialization\.
## IIMotivation
### II\-APreliminaries
Let𝐱∈ℝ1×d\\mathbf\{x\}\\in\\mathbb\{R\}^\{1\\times d\}be one input token to a sparse MoE layer withKKexperts\. The router stores one embedding per expert:
𝐖g=\[𝐖1,𝐖2,…,𝐖K\]∈ℝd×K,\\mathbf\{W\}\_\{g\}=\[\\mathbf\{W\}\_\{1\},\\mathbf\{W\}\_\{2\},\\ldots,\\mathbf\{W\}\_\{K\}\]\\in\\mathbb\{R\}^\{d\\times K\},\(1\)where𝐖j∈ℝd×1\\mathbf\{W\}\_\{j\}\\in\\mathbb\{R\}^\{d\\times 1\}is the embedding of expertjj\. The router produces an affinity distribution and selects itskklargest entries:
𝐬\(𝐱\)\\displaystyle\\mathbf\{s\}\(\\mathbf\{x\}\)=softmax\(𝐱𝐖g\)∈ℝ1×K,\\displaystyle=\\operatorname\{softmax\}\(\\mathbf\{x\}\\mathbf\{W\}\_\{g\}\)\\in\\mathbb\{R\}^\{1\\times K\},\(2\)𝒯k\(𝐱\)\\displaystyle\\mathcal\{T\}\_\{k\}\(\\mathbf\{x\}\)=TopK\(𝐬\(𝐱\),k\),\\displaystyle=\\operatorname\{TopK\}\(\\mathbf\{s\}\(\\mathbf\{x\}\),k\),\(3\)whereTopK\\operatorname\{TopK\}returns the index set of the selected entries\.
Expertjjis an FFN with up\-projection𝐊j∈ℝd×H\\mathbf\{K\}\_\{j\}\\in\\mathbb\{R\}^\{d\\times H\}and down\-projection𝐕j∈ℝH×d\\mathbf\{V\}\_\{j\}\\in\\mathbb\{R\}^\{H\\times d\}\. The intermediate widthHHis also the number of hidden channels\. Its output can be written as
Ej\(𝐱\)=∑h=1HGeLU\(𝐱𝐊j\[:,h\]\)𝐕j\[h,:\]\.E\_\{j\}\(\\mathbf\{x\}\)=\\sum\_\{h=1\}^\{H\}\\operatorname\{GeLU\}\\\!\\left\(\\mathbf\{x\}\\mathbf\{K\}\_\{j\}\[:,h\]\\right\)\\mathbf\{V\}\_\{j\}\[h,:\]\.\(4\)The column𝐊j\[:,h\]\\mathbf\{K\}\_\{j\}\[:,h\]is the key of channelhh: its inner product with𝐱\\mathbf\{x\}determines the channel activation\. The row𝐕j\[h,:\]\\mathbf\{V\}\_\{j\}\[h,:\]is the value: it determines the direction written to the output\. We omit bias terms for clarity\. A conventional top\-kkMoE returns
𝐲=∑j∈𝒯k\(𝐱\)sj\(𝐱\)Ej\(𝐱\)\.\\mathbf\{y\}=\\sum\_\{j\\in\\mathcal\{T\}\_\{k\}\(\\mathbf\{x\}\)\}s\_\{j\}\(\\mathbf\{x\}\)E\_\{j\}\(\\mathbf\{x\}\)\.\(5\)Here,sj\(𝐱\)s\_\{j\}\(\\mathbf\{x\}\)denotes entryjjof𝐬\(𝐱\)\\mathbf\{s\}\(\\mathbf\{x\}\)\. Equation \([5](https://arxiv.org/html/2608.10392#S2.E5)\) routes complete experts\. Equation \([4](https://arxiv.org/html/2608.10392#S2.E4)\), however, shows that each expert is itself a sum of channel responses\. This internal structure provides a natural place to separate shared and residual computation\.
### II\-BObservations
#### Diagnostic setup\.
We analyze layer 10 of a standard top\-2 MoE model trained on TerraIncognita\[[1](https://arxiv.org/html/2608.10392#bib.bib20)\]\. The layer contains 6 experts with 1536 hidden channels each\. All experts are copied from the same pretrained Transformer\[[31](https://arxiv.org/html/2608.10392#bib.bib4)\]FFN before fine\-tuning\. Channel indexhhtherefore refers to the same initial position across experts, which makes same\-index values directly comparable\. For expert pair\(i,j\)\(i,j\), we define
cijh\\displaystyle c\_\{ijh\}=𝐕i\[h,:\]𝐕j\[h,:\]⊤‖𝐕i\[h,:\]‖2‖𝐕j\[h,:\]‖2,\\displaystyle=\\frac\{\\mathbf\{V\}\_\{i\}\[h,:\]\\mathbf\{V\}\_\{j\}\[h,:\]^\{\\top\}\}\{\\\|\\mathbf\{V\}\_\{i\}\[h,:\]\\\|\_\{2\}\\\|\\mathbf\{V\}\_\{j\}\[h,:\]\\\|\_\{2\}\},\(6\)𝒮ij\(δ\)\\displaystyle\\mathcal\{S\}\_\{ij\}\(\\delta\)=\{h∈\{1,…,H\}:cijh≥δ\},\\displaystyle=\\left\\\{h\\in\\\{1,\\ldots,H\\\}:c\_\{ijh\}\\geq\\delta\\right\\\},ρij\(δ\)\\displaystyle\\rho\_\{ij\}\(\\delta\)=\|𝒮ij\(δ\)\|H\.\\displaystyle=\\frac\{\|\\mathcal\{S\}\_\{ij\}\(\\delta\)\|\}\{H\}\.We useδ=0\.98\\delta=0\.98\. The set𝒮ij\\mathcal\{S\}\_\{ij\}contains positions where the two experts write in nearly the same direction, andρij\\rho\_\{ij\}is the fraction of such positions\. We treat this set as a weight\-space estimate of the pair’s reusable response\. Its complementℛij=\{1,…,H\}∖𝒮ij\\mathcal\{R\}\_\{ij\}=\\\{1,\\ldots,H\\\}\\setminus\\mathcal\{S\}\_\{ij\}contains the residual positions\.

\(a\) Shared ratio and co\-activation frequency \(b\) Router rank transition\(c\) Functional\-demand CDF
Figure 1:Diagnosing the MoE module in Transformer layer 10 trained on TerraIncognita\. \(a\) Value alignment and co\-activation for each expert pair\. \(b\) Expert\-rank transitions after the aligned positions are extracted\. \(c\) Pair\-balanced cumulative distribution function \(CDF\)Fε\(m\)=Pr\[kε\(𝐱\)≤m\]F\_\{\\varepsilon\}\(m\)=\\Pr\[k\_\{\\varepsilon\}\(\\mathbf\{x\}\)\\leq m\]atε=0\.05\\varepsilon=0\.05for the five pairs with the lowest and highest shared ratios\. Shading shows the standard error across pairs\.
#### Observation 1: co\-activation identifies reusable responses\.
Figure[1](https://arxiv.org/html/2608.10392#S2.F1)\(a\) comparesρij\\rho\_\{ij\}with the fraction of tokens routed to pair\(i,j\)\(i,j\)\. Their Pearson correlation is 0\.697\. Frequently co\-activated pairs align at up to roughly 80% of their value positions, while some rarely selected pairs align at less than 2%\. The aligned positions are therefore not spread uniformly across expert pairs\. They concentrate among the pairs that the router most often asks to work together, which makes them a useful target for shared computation\.
#### Observation 2: residual computation needs a new routing decision\.
Similarity alone does not show that the aligned positions should be separated from routing\. We therefore freeze the experts, extract𝒮ij\\mathcal\{S\}\_\{ij\}for each token’s original top\-2 pair, and train a new top\-2 router on the residual positions\. Figure[1](https://arxiv.org/html/2608.10392#S2.F1)\(b\) compares the two expert rankings\. Only 5\.7% of tokens keep the same top\-2 set, and the mean diagonal mass is 26\.9%\. The original router scores the complete response, not the residual alone\. Once the shared response is handled, different experts often become better choices\.
#### Observation 3: residual demand varies across tokens\.
We next ask how many residual experts are needed after the shared response has been separated\. For the active pair\(i,j\)\(i,j\), lets~e\(𝐱\)=se\(𝐱\)/\(si\(𝐱\)\+sj\(𝐱\)\)\\widetilde\{s\}\_\{e\}\(\\mathbf\{x\}\)=s\_\{e\}\(\\mathbf\{x\}\)/\(s\_\{i\}\(\\mathbf\{x\}\)\+s\_\{j\}\(\\mathbf\{x\}\)\)be the normalized original gate weight\. We split each expert output intoEe𝒮ijE\_\{e\}^\{\\mathcal\{S\}\_\{ij\}\}andEeℛijE\_\{e\}^\{\\mathcal\{R\}\_\{ij\}\}, which retain the shared and residual positions, respectively\. Merging the shared responses once gives
𝐲ijorig\(𝐱\)\\displaystyle\\mathbf\{y\}\_\{ij\}^\{\\mathrm\{orig\}\}\(\\mathbf\{x\}\)=∑e∈\{i,j\}s~e\(𝐱\)Ee\(𝐱\),\\displaystyle=\\sum\_\{e\\in\\\{i,j\\\}\}\\widetilde\{s\}\_\{e\}\(\\mathbf\{x\}\)E\_\{e\}\(\\mathbf\{x\}\),\(7\)𝐒¯ij\(𝐱\)\\displaystyle\\overline\{\\mathbf\{S\}\}\_\{ij\}\(\\mathbf\{x\}\)=∑e∈\{i,j\}s~e\(𝐱\)Ee𝒮ij\(𝐱\),\\displaystyle=\\sum\_\{e\\in\\\{i,j\\\}\}\\widetilde\{s\}\_\{e\}\(\\mathbf\{x\}\)E\_\{e\}^\{\\mathcal\{S\}\_\{ij\}\}\(\\mathbf\{x\}\),\(8\)𝐫ij\(𝐱\)\\displaystyle\\mathbf\{r\}\_\{ij\}\(\\mathbf\{x\}\)=𝐲ijorig\(𝐱\)−𝐒¯ij\(𝐱\)\.\\displaystyle=\\mathbf\{y\}\_\{ij\}^\{\\mathrm\{orig\}\}\(\\mathbf\{x\}\)\-\\overline\{\\mathbf\{S\}\}\_\{ij\}\(\\mathbf\{x\}\)\.\(9\)The residual target is𝐫ij\(𝐱\)\\mathbf\{r\}\_\{ij\}\(\\mathbf\{x\}\)\. Letq1\(𝐱\),…,qK\(𝐱\)q\_\{1\}\(\\mathbf\{x\}\),\\ldots,q\_\{K\}\(\\mathbf\{x\}\)be the expert order produced by the residual router\. For a prefix of lengthm∈\{1,…,K\}m\\in\\\{1,\\ldots,K\\\}, we stack its residual outputs:
𝐑m\(𝐱\)=\[Eq1\(𝐱\)ℛij\(𝐱\)⋮Eqm\(𝐱\)ℛij\(𝐱\)\]∈ℝm×d\.\\mathbf\{R\}\_\{m\}\(\\mathbf\{x\}\)=\\begin\{bmatrix\}E\_\{q\_\{1\}\(\\mathbf\{x\}\)\}^\{\\mathcal\{R\}\_\{ij\}\}\(\\mathbf\{x\}\)\\\\ \\vdots\\\\ E\_\{q\_\{m\}\(\\mathbf\{x\}\)\}^\{\\mathcal\{R\}\_\{ij\}\}\(\\mathbf\{x\}\)\\end\{bmatrix\}\\in\\mathbb\{R\}^\{m\\times d\}\.\(10\)We then use ridge\-stabilized least squares to measure the best reconstruction available in the row span of𝐑m\(𝐱\)\\mathbf\{R\}\_\{m\}\(\\mathbf\{x\}\):
𝐚m⋆\(𝐱\)\\displaystyle\\mathbf\{a\}\_\{m\}^\{\\star\}\(\\mathbf\{x\}\)=argmin𝐚∈ℝm‖𝐚⊤𝐑m\(𝐱\)−𝐫ij\(𝐱\)‖22\\displaystyle=\\arg\\min\_\{\\mathbf\{a\}\\in\\mathbb\{R\}^\{m\}\}\\left\\\|\\mathbf\{a\}^\{\\top\}\\mathbf\{R\}\_\{m\}\(\\mathbf\{x\}\)\-\\mathbf\{r\}\_\{ij\}\(\\mathbf\{x\}\)\\right\\\|\_\{2\}^\{2\}\+γm\(𝐱\)‖𝐚‖22,\\displaystyle\\hskip 74\.00005pt\+\\gamma\_\{m\}\(\\mathbf\{x\}\)\\\|\\mathbf\{a\}\\\|\_\{2\}^\{2\},\(11\)𝐲^m\(𝐱\)\\displaystyle\\widehat\{\\mathbf\{y\}\}\_\{m\}\(\\mathbf\{x\}\)=𝐒¯ij\(𝐱\)\+\(𝐚m⋆\(𝐱\)\)⊤𝐑m\(𝐱\),\\displaystyle=\\overline\{\\mathbf\{S\}\}\_\{ij\}\(\\mathbf\{x\}\)\+\(\\mathbf\{a\}\_\{m\}^\{\\star\}\(\\mathbf\{x\}\)\)^\{\\top\}\\mathbf\{R\}\_\{m\}\(\\mathbf\{x\}\),\(12\)em\(𝐱\)\\displaystyle e\_\{m\}\(\\mathbf\{x\}\)=‖𝐲^m\(𝐱\)−𝐲ijorig\(𝐱\)‖2max\{‖𝐲ijorig\(𝐱\)‖2,η\},\\displaystyle=\\frac\{\\\|\\widehat\{\\mathbf\{y\}\}\_\{m\}\(\\mathbf\{x\}\)\-\\mathbf\{y\}\_\{ij\}^\{\\mathrm\{orig\}\}\(\\mathbf\{x\}\)\\\|\_\{2\}\}\{\\max\\\{\\\|\\mathbf\{y\}\_\{ij\}^\{\\mathrm\{orig\}\}\(\\mathbf\{x\}\)\\\|\_\{2\},\\eta\\\}\},\(13\)𝒜ε\(𝐱\)\\displaystyle\\mathcal\{A\}\_\{\\varepsilon\}\(\\mathbf\{x\}\)=\{m∈\{1,…,K\}:em\(𝐱\)≤ε\},\\displaystyle=\\left\\\{m\\in\\\{1,\\ldots,K\\\}:e\_\{m\}\(\\mathbf\{x\}\)\\leq\\varepsilon\\right\\\},\(14\)kε\(𝐱\)\\displaystyle k\_\{\\varepsilon\}\(\\mathbf\{x\}\)=min\(𝒜ε\(𝐱\)∪\{K\+1\}\)\.\\displaystyle=\\min\\left\(\\mathcal\{A\}\_\{\\varepsilon\}\(\\mathbf\{x\}\)\\cup\\\{K\+1\\\}\\right\)\.\(15\)We setγm\(𝐱\)=10−6tr\(𝐑m\(𝐱\)𝐑m\(𝐱\)⊤\)/m\\gamma\_\{m\}\(\\mathbf\{x\}\)=10^\{\-6\}\\operatorname\{tr\}\(\\mathbf\{R\}\_\{m\}\(\\mathbf\{x\}\)\\mathbf\{R\}\_\{m\}\(\\mathbf\{x\}\)^\{\\top\}\)/mandη=10−12\\eta=10^\{\-12\}\.ε\\varepsilonis the fidelity tolerance, and𝒜ε\(𝐱\)\\mathcal\{A\}\_\{\\varepsilon\}\(\\mathbf\{x\}\)is the set of prefixes that attain it\. The sentinelK\+1K\+1records that no available prefix reaches the target\. Thus,kε\(𝐱\)k\_\{\\varepsilon\}\(\\mathbf\{x\}\)measures how quickly the ordered residual span recovers the original output\. Figure[1](https://arxiv.org/html/2608.10392#S2.F1)\(c\) plots its CDF withε\\varepsilonset to 0\.05\. Atm=2m=2, the pair\-balanced success rates are 56\.3% for high\-shared pairs and 11\.8% for low\-shared pairs; their meankεk\_\{\\varepsilon\}values are 2\.70 and 4\.61\. Across pairs, shared ratio and mean demand correlate at−0\.673\-0\.673\(bootstrap 95% CI\[−0\.899,−0\.315\]\[\-0\.899,\-0\.315\]\), and the trend persists under uniform shared averaging and across all four held\-out environments atε∈\{0\.10,0\.20\}\\varepsilon\\in\\\{0\.10,0\.20\\\}\. Shared coverage is therefore related to the full distribution of residual demand\.
### II\-CDesign Implications
These observations show that the key question is not whether shared modeling, fine\-grained computation, and dynamic routing coexist, but in what order they act\. Shared modeling must first determine how much computation is reusable\. Fine\-grained computation must then identify which content provides that response\. Only after these choices should dynamic routing decide how many experts the remainder needs\. Reversing or separating this order would route a response that has already changed\. The next section implements this principle as*share first, then route what remains*\.
## IIIMethod
\(a\) Conventional top\-kkMoE\(b\) UniF\-MoEFigure 2:Conventional top\-kkMoE selects complete experts in one step\. UniF\-MoE first executes token\-selected blocks once in a shared expert, then routes only the complementary blocks to a token\-dependent number of residual experts\. Colored blocks are active\.Figure[2](https://arxiv.org/html/2608.10392#S3.F2)compares UniF\-MoE with conventional top\-kkMoE\. Rather than attaching independent controllers for shared computation, internal selection, and expert count, UniF\-MoE makes these decisions sequentially under one token\-dependent budget\.
### III\-ABlockwise Shared\-Residual Partitioning
A UniF\-MoE layer contains one shared expertEshrE\_\{\\mathrm\{shr\}\}andKKresidual experts\{E1,…,EK\}\\\{E\_\{1\},\\ldots,E\_\{K\}\\\}\. Each expert has intermediate widthHHand is divided intoBBaligned blocks of widthM=H/BM=H/B\. Blockbbcontains positionsℋb=\{\(b−1\)M\+1,…,bM\}\\mathcal\{H\}\_\{b\}=\\\{\(b\-1\)M\+1,\\ldots,bM\\\}and produces
Ej\(b\)\(𝐱\)=∑h=\(b−1\)M\+1bMGeLU\(𝐱𝐊j\[:,h\]\)𝐕j\[h,:\]\.E\_\{j\}^\{\(b\)\}\(\\mathbf\{x\}\)=\\sum\_\{h=\(b\-1\)M\+1\}^\{bM\}\\operatorname\{GeLU\}\\\!\\left\(\\mathbf\{x\}\\mathbf\{K\}\_\{j\}\[:,h\]\\right\)\\mathbf\{V\}\_\{j\}\[h,:\]\.\(16\)Thus,Ej\(𝐱\)=∑b=1BEj\(b\)\(𝐱\)E\_\{j\}\(\\mathbf\{x\}\)=\\sum\_\{b=1\}^\{B\}E\_\{j\}^\{\(b\)\}\(\\mathbf\{x\}\)\. All experts are initialized from the same dense FFN, so their block boundaries are aligned at the start of training\. For each token, the shared expert executes one subset of blocks, and the residual experts execute the complementary subset\. Here, “residual” refers to the channel positions left after shared\-block selection, not to the Transformer residual connection\. No hidden position is assigned to both pathways for the same token\.
### III\-BToken\-Adaptive Shared Modeling
#### Shared demand\.
The standard router𝐖g\\mathbf\{W\}\_\{g\}contains the embeddings ofKKresidual experts\. We add one learnable column𝐖shr\\mathbf\{W\}\_\{\\mathrm\{shr\}\}for the shared expert:
𝐖g⋆=\[𝐖shr,𝐖g\]∈ℝd×\(K\+1\)\.\\mathbf\{W\}\_\{g\}^\{\\star\}=\[\\mathbf\{W\}\_\{\\mathrm\{shr\}\},\\mathbf\{W\}\_\{g\}\]\\in\\mathbb\{R\}^\{d\\times\(K\+1\)\}\.\(17\)This column produces the shared\-demand score
α\(𝐱\)=τ\+\(1−2τ\)σ\(𝐱𝐖shr\),τ=B−1B2,\\alpha\(\\mathbf\{x\}\)=\\tau\+\(1\-2\\tau\)\\sigma\\\!\\left\(\\mathbf\{x\}\\mathbf\{W\}\_\{\\mathrm\{shr\}\}\\right\),\\qquad\\tau=\\frac\{B\-1\}\{B^\{2\}\},\(18\)whereσ\(⋅\)\\sigma\(\\cdot\)is thesigmoid\\operatorname\{sigmoid\}function\. We interpretα\(𝐱\)\\alpha\(\\mathbf\{x\}\)as the token’s demand for shared computation\. It controls both the number of shared blocks and the mixture weight of the shared pathway\. The shared block count is
b\(𝐱\)=round\(Bα\(𝐱\)\)\.b\(\\mathbf\{x\}\)=\\operatorname\{round\}\(B\\alpha\(\\mathbf\{x\}\)\)\.\(19\)Becauseσ\(z\)∈\(0,1\)\\sigma\(z\)\\in\(0,1\), the chosen bound givesτ<α\(𝐱\)<1−τ\\tau<\\alpha\(\\mathbf\{x\}\)<1\-\\tau\. ForB≥2B\\geq 2, rounding therefore guaranteesb\(𝐱\)∈\{1,…,B−1\}b\(\\mathbf\{x\}\)\\in\\\{1,\\ldots,B\-1\\\}: every token uses at least one shared block and leaves at least one block for residual computation\.
#### Shared block selection\.
The scoreα\(𝐱\)\\alpha\(\\mathbf\{x\}\)determines how many blocks to share, but not which blocks\. We represent shared blockbbby the mean of its up\-projection keys:
μb=1M∑h=\(b−1\)M\+1bM𝐊shr\[:,h\]∈ℝd×1\.\\mathbf\{\\mu\}\_\{b\}=\\frac\{1\}\{M\}\\sum\_\{h=\(b\-1\)M\+1\}^\{bM\}\\mathbf\{K\}\_\{\\mathrm\{shr\}\}\[:,h\]\\in\\mathbb\{R\}^\{d\\times 1\}\.\(20\)Its priority for token𝐱\\mathbf\{x\}is
ub\(𝐱\)=𝐱μb\.u\_\{b\}\(\\mathbf\{x\}\)=\\mathbf\{x\}\\mathbf\{\\mu\}\_\{b\}\.\(21\)Theb\(𝐱\)b\(\\mathbf\{x\}\)blocks with the largest priorities form the shared index set\. All other blocks form the residual index set:
ℐshr\(𝐱\)\\displaystyle\\mathcal\{I\}\_\{\\mathrm\{shr\}\}\(\\mathbf\{x\}\)=TopK\(\{ub\(𝐱\)\}b=1B,b\(𝐱\)\),\\displaystyle=\\operatorname\{TopK\}\\\!\\left\(\\\{u\_\{b\}\(\\mathbf\{x\}\)\\\}\_\{b=1\}^\{B\},b\(\\mathbf\{x\}\)\\right\),\(22\)ℐres\(𝐱\)\\displaystyle\\mathcal\{I\}\_\{\\mathrm\{res\}\}\(\\mathbf\{x\}\)=\{1,…,B\}∖ℐshr\(𝐱\)\.\\displaystyle=\\\{1,\\ldots,B\\\}\\setminus\\mathcal\{I\}\_\{\\mathrm\{shr\}\}\(\\mathbf\{x\}\)\.\(23\)Two tokens can therefore use the same shared width while choosing different content\. This second decision prevents shared demand from collapsing into a fixed prefix or a scalar width rule\.
### III\-CCumulative Residual\-Expert Routing
The residual embeddings produce an affinity distribution
𝐬\(𝐱\)=softmax\(𝐱𝐖g\)∈ℝ1×K\.\\mathbf\{s\}\(\\mathbf\{x\}\)=\\operatorname\{softmax\}\(\\mathbf\{x\}\\mathbf\{W\}\_\{g\}\)\\in\\mathbb\{R\}^\{1\\times K\}\.\(24\)Letq1\(𝐱\),…,qK\(𝐱\)q\_\{1\}\(\\mathbf\{x\}\),\\ldots,q\_\{K\}\(\\mathbf\{x\}\)be the expert indices sorted by decreasing affinity\. We write the sorted values aspi\(𝐱\)=sqi\(𝐱\)\(𝐱\)p\_\{i\}\(\\mathbf\{x\}\)=s\_\{q\_\{i\}\(\\mathbf\{x\}\)\}\(\\mathbf\{x\}\):
p1\(𝐱\)≥p2\(𝐱\)≥⋯≥pK\(𝐱\),pi\(𝐱\)=sqi\(𝐱\)\(𝐱\)\.p\_\{1\}\(\\mathbf\{x\}\)\\geq p\_\{2\}\(\\mathbf\{x\}\)\\geq\\cdots\\geq p\_\{K\}\(\\mathbf\{x\}\),\\qquad p\_\{i\}\(\\mathbf\{x\}\)=s\_\{q\_\{i\}\(\\mathbf\{x\}\)\}\(\\mathbf\{x\}\)\.\(25\)The residual demand is the complement of the shared demand:
β\(𝐱\)=1−α\(𝐱\)\.\\beta\(\\mathbf\{x\}\)=1\-\\alpha\(\\mathbf\{x\}\)\.\(26\)Thus, the two pathways divide one budget instead of estimating their importance independently\. We activate the smallest prefix whose cumulative affinity covers this demand:
k\(𝐱\)=min\{n∈\{1,…,K\}:∑i=1npi\(𝐱\)≥β\(𝐱\)\}\.k\(\\mathbf\{x\}\)=\\min\\left\\\{n\\in\\\{1,\\ldots,K\\\}:\\sum\_\{i=1\}^\{n\}p\_\{i\}\(\\mathbf\{x\}\)\\geq\\beta\(\\mathbf\{x\}\)\\right\\\}\.\(27\)Equation \([27](https://arxiv.org/html/2608.10392#S3.E27)\) makes expert count conditional on the preceding shared allocation\. The valueβ\(𝐱\)\\beta\(\\mathbf\{x\}\)says how much computation remains, while the shape of𝐬\(𝐱\)\\mathbf\{s\}\(\\mathbf\{x\}\)says how concentrated that demand is across experts\. A large residual can still be covered by one expert when the distribution is sharp, whereas a diffuse distribution recruits more experts\.
### III\-DShared\-Residual Output Merging
The selected output of the shared expert and the output of residual expertjjare
Eshr𝒮\(𝐱\)\\displaystyle E\_\{\\mathrm\{shr\}\}^\{\\mathcal\{S\}\}\(\\mathbf\{x\}\)=∑b∈ℐshr\(𝐱\)Eshr\(b\)\(𝐱\),\\displaystyle=\\sum\_\{b\\in\\mathcal\{I\}\_\{\\mathrm\{shr\}\}\(\\mathbf\{x\}\)\}E\_\{\\mathrm\{shr\}\}^\{\(b\)\}\(\\mathbf\{x\}\),\(28\)Ejℛ\(𝐱\)\\displaystyle E\_\{j\}^\{\\mathcal\{R\}\}\(\\mathbf\{x\}\)=∑b∈ℐres\(𝐱\)Ej\(b\)\(𝐱\)\.\\displaystyle=\\sum\_\{b\\in\\mathcal\{I\}\_\{\\mathrm\{res\}\}\(\\mathbf\{x\}\)\}E\_\{j\}^\{\(b\)\}\(\\mathbf\{x\}\)\.\(29\)The final layer output is
𝐲=α\(𝐱\)Eshr𝒮\(𝐱\)\+∑i=1k\(𝐱\)pi\(𝐱\)Eqi\(𝐱\)ℛ\(𝐱\)\.\\mathbf\{y\}=\\alpha\(\\mathbf\{x\}\)E\_\{\\mathrm\{shr\}\}^\{\\mathcal\{S\}\}\(\\mathbf\{x\}\)\+\\sum\_\{i=1\}^\{k\(\\mathbf\{x\}\)\}p\_\{i\}\(\\mathbf\{x\}\)E\_\{q\_\{i\}\(\\mathbf\{x\}\)\}^\{\\mathcal\{R\}\}\(\\mathbf\{x\}\)\.\(30\)The shared output is weighted byα\(𝐱\)\\alpha\(\\mathbf\{x\}\)\. Selected residual experts keep their affinities\. DefinePn\(𝐱\)=∑i=1npi\(𝐱\)P\_\{n\}\(\\mathbf\{x\}\)=\\sum\_\{i=1\}^\{n\}p\_\{i\}\(\\mathbf\{x\}\)\. By the definition of the smallest sufficient prefix,
β\(𝐱\)\\displaystyle\\beta\(\\mathbf\{x\}\)≤Pk\(𝐱\)\(𝐱\)<β\(𝐱\)\+pk\(𝐱\)\(𝐱\),\\displaystyle\\leq P\_\{k\(\\mathbf\{x\}\)\}\(\\mathbf\{x\}\)<\\beta\(\\mathbf\{x\}\)\+p\_\{k\(\\mathbf\{x\}\)\}\(\\mathbf\{x\}\),\(31\)1\\displaystyle 1≤α\(𝐱\)\+Pk\(𝐱\)\(𝐱\)<1\+pk\(𝐱\)\(𝐱\)\.\\displaystyle\\leq\\alpha\(\\mathbf\{x\}\)\+P\_\{k\(\\mathbf\{x\}\)\}\(\\mathbf\{x\}\)<1\+p\_\{k\(\\mathbf\{x\}\)\}\(\\mathbf\{x\}\)\.\(32\)The residual routing mass is at leastβ\(𝐱\)\\beta\(\\mathbf\{x\}\)and exceeds it by less than the affinity of the final expert\. Becauseα\(𝐱\)\+β\(𝐱\)=1\\alpha\(\\mathbf\{x\}\)\+\\beta\(\\mathbf\{x\}\)=1, the total coefficient mass has a controlled overshoot while preserving the router’s confidence\.
### III\-ETraining Objective and Compute
The shared and residual embeddings should remain distinct rather than collapse into interchangeable routes\. We therefore initialize the columns of𝐖g⋆\\mathbf\{W\}\_\{g\}^\{\\star\}orthonormally and regularize their Gram matrix:
ℒdiv\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{div\}\}=‖\(𝐖g⋆\)⊤𝐖g⋆−𝐈K\+1‖F,\\displaystyle=\\left\\\|\(\\mathbf\{W\}\_\{g\}^\{\\star\}\)^\{\\top\}\\mathbf\{W\}\_\{g\}^\{\\star\}\-\\mathbf\{I\}\_\{K\+1\}\\right\\\|\_\{F\},\(33\)ℒ\\displaystyle\\mathcal\{L\}=ℒtask\+λdivℒdiv,\\displaystyle=\\mathcal\{L\}\_\{\\mathrm\{task\}\}\+\\lambda\_\{\\mathrm\{div\}\}\\mathcal\{L\}\_\{\\mathrm\{div\}\},\(34\)where𝐈K\+1\\mathbf\{I\}\_\{K\+1\}is the identity matrix\. The diagonal terms keep each embedding at unit scale, yielding a simple normalized routing geometry\. The off\-diagonal terms separate routing directions, increasing expert diversity\. Orthonormal initialization starts from the same geometry whenK\+1≤dK\+1\\leq d\. Together, these effects promote diverse roles and sparse expert overlap without forcing the expert outputs themselves to be orthogonal\.
Measured in blocks, token𝐱\\mathbf\{x\}activates
CB\(𝐱\)=b\(𝐱\)\+k\(𝐱\)\[B−b\(𝐱\)\]C\_\{B\}\(\\mathbf\{x\}\)=b\(\\mathbf\{x\}\)\+k\(\\mathbf\{x\}\)\[B\-b\(\\mathbf\{x\}\)\]\(35\)blocks\. A conventional top\-kklayer executeskBkBblocks\. Whenk\(𝐱\)=1k\(\\mathbf\{x\}\)=1, the shared and residual blocks together equal one complete FFN\. Every additional residual expert adds only theB−b\(𝐱\)B\-b\(\\mathbf\{x\}\)blocks that remain\. The ordered decomposition therefore performs reusable work once and confines repeated expert cost to the residual pathway\.
## IVExperiments
TABLE I:Out\-of\-domain accuracy \(%\) on DomainBed\. – and†\\daggerdenote unreported and reproduced results, respectively\.TABLE II:In\-domain GLUE task scores \(%\)\. MoE Best is the per\-task oracle upper envelope of the fixed top\-kkvariants\.### IV\-AExperimental Setup
#### Benchmarks and implementation details\.
We evaluate UniF\-MoE on vision and language tasks\. For vision, the backbone is DeiT\-S/16\[[30](https://arxiv.org/html/2608.10392#bib.bib23)\]pretrained on ImageNet\[[5](https://arxiv.org/html/2608.10392#bib.bib22)\]\. We use the train\-validation selection criterion of DomainBed\[[10](https://arxiv.org/html/2608.10392#bib.bib16)\]and report out\-of\-domain accuracy on PACS\[[20](https://arxiv.org/html/2608.10392#bib.bib17)\], VLCS\[[8](https://arxiv.org/html/2608.10392#bib.bib18)\], OfficeHome\[[32](https://arxiv.org/html/2608.10392#bib.bib19)\], TerraIncognita\[[1](https://arxiv.org/html/2608.10392#bib.bib20)\], and DomainNet\[[25](https://arxiv.org/html/2608.10392#bib.bib21)\]\.KKandBBare set to 6 and 8, respectively\. For language, the backbone is BERT\-large\[[6](https://arxiv.org/html/2608.10392#bib.bib24)\]\. We evaluate CoLA\[[34](https://arxiv.org/html/2608.10392#bib.bib26)\], MRPC\[[7](https://arxiv.org/html/2608.10392#bib.bib27)\], QNLI\[[27](https://arxiv.org/html/2608.10392#bib.bib29)\], MNLI\[[35](https://arxiv.org/html/2608.10392#bib.bib30)\], and RTE\[[2](https://arxiv.org/html/2608.10392#bib.bib31)\]from GLUE\[[33](https://arxiv.org/html/2608.10392#bib.bib25)\], usingK=16K=16andB=16B=16\. Learning rate and other settings follow the corresponding baselines\. Our results are averaged over three runs\.
#### Comparative baselines\.
Vision baselines include dense DeiT\-S/16; static MoEs GMoE\[[19](https://arxiv.org/html/2608.10392#bib.bib6)\], EMoE, and EMoE\-L \(EMoE\-learn\)\[[26](https://arxiv.org/html/2608.10392#bib.bib7)\]; domain\-generalization methods LFME\[[3](https://arxiv.org/html/2608.10392#bib.bib32)\], DMDA\[[22](https://arxiv.org/html/2608.10392#bib.bib33)\], and PC\-MoE\[[23](https://arxiv.org/html/2608.10392#bib.bib34)\]; and dynamic MoEs DynMoE\[[12](https://arxiv.org/html/2608.10392#bib.bib12)\]and MASS\[[24](https://arxiv.org/html/2608.10392#bib.bib13)\]\. Language baselines include dense BERT\-large, fixed top\-kkMoEs withk∈\{1,2,4,8\}k\\in\\\{1,2,4,8\\\}, DynMoE, and MASS\. For fixed routing, we also report the average acrosskkand the per\-task best result\.
### IV\-BMain Results
#### Vision tasks\.
Table[I](https://arxiv.org/html/2608.10392#S4.T1)shows that UniF\-MoE obtains the best average result on DomainBed, leading on PACS, VLCS, and DomainNet and tying the best result on OfficeHome\. These four datasets mix category evidence with strong style or source variation, a setting that plays to UniF\-MoE’s strength: reusable blocks preserve transferable cues, while residual routing directs the remaining appearance\-specific computation to the patches that need it\. TerraIncognita is dominated by location\-dependent backgrounds and camera viewpoints, where LFME’s explicit domain specialization remains stronger\. The overall pattern supports the value of coordinating reusable and residual computation without imposing one budget on every patch\.
#### Language tasks\.
Table[II](https://arxiv.org/html/2608.10392#S4.T2)shows that UniF\-MoE performs best on all five GLUE tasks, spanning linguistic acceptability, paraphrase detection, and several forms of textual inference\. The strongest fixedkkchanges across tasks, yet even their per\-task oracle envelope remains below UniF\-MoE\. Choosingkkper dataset still assigns one operating point to every token\. UniF\-MoE instead identifies reusable linguistic computation before adapting the residual expert count within each sequence, replacing a global expert budget with a token\-level dependency\.
### IV\-CToken\-Adaptive Computation and Cost
#### Element\-conditioned pathway composition\.
Figure[3](https://arxiv.org/html/2608.10392#S4.F3)groups image patches by semantic element and records their routing choices in Transformer layer 10\. Shared blocks recur across several elements, whereas residual experts show more selective preferences\. Patches that reuse the same shared block can still activate different residual experts, confirming the dependency at the center of our design: shared selection removes common computation, but does not predetermine who should process the remainder\.
#### Patch\-wise resource allocation\.
Figure[4](https://arxiv.org/html/2608.10392#S4.F4)shows alarm clocks from different domains\. High\-compute patches fall on the display in Product and Real World, and on the bells, ticks, and hands in Clipart\. The Art example assigns minimal computation to a watermark and several strong but irrelevant strokes\. The joint budget therefore responds neither to domain nor contrast alone; it reserves residual capacity for regions that carry useful class evidence after reusable features are handled\.
#### Measured cost comparison\.
Table[III](https://arxiv.org/html/2608.10392#S4.T3)compares UniF\-MoE with representative dense, static\-MoE, and dynamic\-MoE designs\. EMoE has the lowest activated parameter count and FLOPs among the MoEs, whereas UniF\-MoE achieves the lowest inference time and memory\. Relative to top\-2 GMoE, it activates 9\.1% fewer parameters and uses 16\.1% fewer FLOPs, reducing inference time and memory by 45\.2% and 52\.7%; it also undercuts DynMoE on every compute and runtime statistic\. UniF\-MoE does not minimize every analytical proxy, but sharing once and repeating only residual blocks yields the strongest measured inference profile among the MoE baselines\.
\(a\) Original image\(b\) Patch semantics

\(c\) Shared\-block and residual\-expert usage
Figure 3:Element\-conditioned routing for a*Person*example of PACS\. Darker cells indicate higher relative use within each semantic element\.\(a\) Product\(b\) Art\(c\) Real World\(d\) ClipartFigure 4:Patch\-level block usageCB\(𝐱\)C\_\{B\}\(\\mathbf\{x\}\)for the*Alarm Clock*class across the four OfficeHome domains\. High\-demand patches withCB\(𝐱\)≥20C\_\{B\}\(\\mathbf\{x\}\)\\geq 20are outlined\.TABLE III:Cost comparison on VLCS\.NNdenotes the total parameter count, andNAN\_\{A\}denotes the average activated parameter count across samples; both are reported in units of2202^\{20\}parameters\. FLOPs are reported per image in units of2302^\{30\}operations\. T\-TPS and I\-TPS denote training and inference time per step, respectively, in seconds; T\-RTM and I\-RTM denote the corresponding run\-time memory in GiB\. The best MoE values are bold\.
### IV\-DAblation Study and Hyperparameter Analysis
#### Contribution of the three adaptive decisions\.
TABLE IV:Ablating the three adaptive decisions\.×\\timesindicates that the corresponding mechanism is replaced by its non\-adaptive counterpart: fixedα=0\.4\\alpha=0\.4, top\-2 residual\-expert activation, or prefix\-based block selection, respectively\. The full block counts are 56 on DomainBed and 272 on GLUE\.Table[IV](https://arxiv.org/html/2608.10392#S4.T4)breaks the ordered process one decision at a time\. Every replacement lowers accuracy and increases computation on both benchmarks\. Fixingα\\alphais most disruptive because an incorrect shared\-residual split propagates to both later decisions\. Prefix\-based selection loses token\-specific shared content, while top\-2 residual\-expert activation ignores how much demand remains\. Their failures show that the gain comes from coordinating the three stages, not merely adding three forms of sparsity\.
#### Effect of diversity regularization\.

\(a\) Residual expert co\-activation  \(b\) Diversity\-loss weight
Figure 5:Effect of diversity regularization\. \(a\) Residual expert co\-activation on PACS without and withℒdiv\\mathcal\{L\}\_\{\\mathrm\{div\}\}\. \(b\) GLUE results with different values ofλdiv\\lambda\_\{\\mathrm\{div\}\}\.Figure[5](https://arxiv.org/html/2608.10392#S4.F5)examines the qualitative and quantitative effects of the diversity constraint\. As shown in Figure[5](https://arxiv.org/html/2608.10392#S4.F5)\(a\), it reduces mean pair co\-activation by 62\.9% and moves the router embeddings close to orthonormality\. The graph does not become empty: useful expert cooperation remains, but routing is no longer concentrated in the same few pairs\. Figure[5](https://arxiv.org/html/2608.10392#S4.F5)\(b\) further favors the intermediate settingλdiv=0\.01\\lambda\_\{\\mathrm\{div\}\}=0\.01on average\. A weaker constraint leaves routing directions overly correlated, whereas an excessive one can overwhelm the task loss\. Together, the two panels show that moderate regularization preserves distinct residual routes without prohibiting useful cooperation\.
#### Shared\-block granularity\.
Figure 6:DomainBed results with different values ofBB\.Figure[6](https://arxiv.org/html/2608.10392#S4.F6)shows thatB=8B=8or1616performs best on four of the five DomainBed datasets\. WithB=4B=4, one selection changes a quarter of the FFN, which is too coarse to localize the shared response\. WithB=32B=32, each prototype summarizes a narrow group of keys and cannot cover enough semantics\. We useB=8B=8as a stable balance between selection granularity and complexity\.
## VRelated Work
### V\-AShared Modeling and Expert Specialization
Trained experts may contain overlapping computations\. DeepSeekMoE finely segments experts into smaller ones and adds dedicated shared experts\[[4](https://arxiv.org/html/2608.10392#bib.bib10)\], while Union\-of\-Experts builds a virtual shared expert from routing\-neuron outputs\[[36](https://arxiv.org/html/2608.10392#bib.bib36)\]\. A complementary line encourages expert specialization: orthogonality and variance objectives reduce overlap\[[11](https://arxiv.org/html/2608.10392#bib.bib35)\], and MP\-MoE uses inter\-expert covariance to select diverse expert sets\[[16](https://arxiv.org/html/2608.10392#bib.bib38)\]\. These methods improve reuse or specialization, but do not use token\-specific shared allocation to define the computation that specialization should handle\. UniF\-MoE makes that dependency explicit: the shared pathway handles reusable block responses first, and the Gram constraint preserves distinct routes for what remains\.
### V\-BDynamic Routing and Fine\-grained Computation
Dynamic MoEs vary expert count by accumulating routing confidence\[[13](https://arxiv.org/html/2608.10392#bib.bib11)\], growing or pruning the expert pool\[[12](https://arxiv.org/html/2608.10392#bib.bib12)\], combining cumulative mass with expert expansion\[[24](https://arxiv.org/html/2608.10392#bib.bib13)\], or distributing a constrained budget across layers and tokens\[[21](https://arxiv.org/html/2608.10392#bib.bib37)\]\. Fine\-grained methods instead focus inside the FFN: Emergent MoE uses key centroids to expose modular structure\[[26](https://arxiv.org/html/2608.10392#bib.bib7)\], while nested and slimmable experts vary the executed width\[[15](https://arxiv.org/html/2608.10392#bib.bib14),[29](https://arxiv.org/html/2608.10392#bib.bib15)\]\. These approaches adapt expert count or width, but generally make the two allocations independently and without conditioning either on a shared response\. UniF\-MoE orders them through one shared\-residual budget: the shared demand determines shared width and content, while the residual demand determines the residual expert count\.
## VIConclusion
Shared modeling, fine\-grained computation, and dynamic routing need not be separate mechanisms\. By exposing their shared\-residual dependency, this work turns them into one ordered rule: identify reusable computation first, then route what remains\. UniF\-MoE realizes this rule through token\-adaptive shared width, shared content, and residual expert count\. Across vision and language tasks, the resulting allocation improves predictive performance while reducing activated computation and measured inference overhead\. Ablations and routing analyses confirm that the three stages are complementary, establishing shared\-residual decomposition as a practical basis for jointly controlling intra\-expert and inter\-expert computation\.
## Acknowledgment
This work was funded in part by the National Natural Science Foundation of China grant under number 62536004, 62222603, in part by the Key\-Area Research and Development Program of Guangdong Province under number 2023B0303030001, in part by the Program for Guangdong Introducing Innovative and Entrepreneurial Teams \(2019ZT08X214\), and in part by the Science and Technology Program of Guangzhou under number 2024A04J6310, and in part by the Fundamental Research Funds for the Central Universities 2025ZYGXZR021\.
## References
- \[1\]S\. Beery, G\. Van Horn, and P\. Perona\(2018\)Recognition in terra incognita\.InProceedings of the European Conference on Computer Vision,pp\. 456–473\.Cited by:[§II\-B](https://arxiv.org/html/2608.10392#S2.SS2.SSS0.Px1.p1.2),[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[2\]L\. Bentivogli, I\. Dagan, H\. T\. Dang, D\. Giampiccolo, and B\. Magnini\(2009\)The fifth PASCAL recognizing textual entailment challenge\.InProceedings of the Second Text Analysis Conference,External Links:[Link](https://tac.nist.gov/publications/2009/additional.papers/RTE5_overview.proceedings.pdf)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[3\]L\. Chen, Y\. Zhang, Y\. Song, Z\. Shen, and L\. Liu\(2024\)LFME: a simple framework for learning from multiple experts in domain generalization\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 102919–102947\.External Links:[Document](https://dx.doi.org/10.52202/079017-3269)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3)\.
- \[4\]D\. Dai, C\. Deng, C\. Zhao, R\. X\. Xu, H\. Gao, D\. Chen, J\. Li, W\. Zeng, X\. Yu, Y\. Wu, Z\. Xie, Y\. K\. Li, P\. Huang, F\. Luo, C\. Ruan, Z\. Sui, and W\. Liang\(2024\)DeepSeekMoE: towards ultimate expert specialization in mixture\-of\-experts language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1280–1297\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.70)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-A](https://arxiv.org/html/2608.10392#S5.SS1.p1.1)\.
- \[5\]J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and F\. Li\(2009\)ImageNet: a large\-scale hierarchical image database\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2009.5206848)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[6\]J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova\(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),pp\. 4171–4186\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1423)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[7\]W\. B\. Dolan and C\. Brockett\(2005\)Automatically constructing a corpus of sentential paraphrases\.InProceedings of the Third International Workshop on Paraphrasing \(IWP2005\),External Links:[Link](https://aclanthology.org/I05-5002/)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[8\]C\. Fang, Y\. Xu, and D\. N\. Rockmore\(2013\)Unbiased metric learning: on the utilization of multiple datasets and web images for softening bias\.InProceedings of the IEEE International Conference on Computer Vision,pp\. 1657–1664\.Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[9\]M\. Geva, R\. Schuster, J\. Berant, and O\. Levy\(2021\)Transformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5484–5495\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.446)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p4.1)\.
- \[10\]I\. Gulrajani and D\. Lopez\-Paz\(2021\)In search of lost domain generalization\.InInternational Conference on Learning Representations,Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[11\]H\. Guo, H\. Lu, G\. Nan, B\. Chu, J\. Zhuang, Y\. Yang, W\. Che, X\. Cao, S\. Leng, Q\. Cui, and X\. Jiang\(2025\)Advancing expert specialization for better MoE\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-A](https://arxiv.org/html/2608.10392#S5.SS1.p1.1)\.
- \[12\]Y\. Guo, Z\. Cheng, X\. Tang, Z\. Tu, and T\. Lin\(2025\)Dynamic mixture of experts: an auto\-tuning approach for efficient transformer models\.InInternational Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[13\]Q\. Huang, Z\. An, N\. Zhuang, M\. Tao, C\. Zhang, Y\. Jin, K\. Xu, L\. Chen, S\. Huang, and Y\. Feng\(2024\)Harder task needs more experts: dynamic routing in MoE models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12883–12895\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.696)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[14\]R\. A\. Jacobs, M\. I\. Jordan, S\. J\. Nowlan, and G\. E\. Hinton\(1991\)Adaptive mixtures of local experts\.Neural Computation3\(1\),pp\. 79–87\.External Links:[Document](https://dx.doi.org/10.1162/neco.1991.3.1.79)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p1.1)\.
- \[15\]G\. Jain, N\. Hegde, A\. Kusupati, A\. Nagrani, S\. Buch, P\. Jain, A\. Arnab, and S\. Paul\(2024\)Mixture of nested experts: adaptive processing of visual tokens\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 58480–58497\.External Links:[Document](https://dx.doi.org/10.52202/079017-1863)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[16\]X\. Kang, D\. Xue, Z\. Wang, C\. Du, X\. Chen, H\. Zhou, H\. Chen, and C\. Meng\(2026\)Breaking the echo chamber: a dynamic ensemble pruning perspective on MoE\.InInternational Conference on Machine Learning,Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-A](https://arxiv.org/html/2608.10392#S5.SS1.p1.1)\.
- \[17\]A\. Komatsuzaki, J\. Puigcerver, J\. Lee\-Thorp, C\. Riquelme Ruiz, B\. Mustafa, J\. Ainslie, Y\. Tay, M\. Dehghani, and N\. Houlsby\(2023\)Sparse upcycling: training mixture\-of\-experts from dense checkpoints\.InInternational Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p4.1)\.
- \[18\]D\. Lepikhin, H\. Lee, Y\. Xu, D\. Chen, O\. Firat, Y\. Huang, M\. Krikun, N\. Shazeer, and Z\. Chen\(2021\)GShard: scaling giant models with conditional computation and automatic sharding\.InInternational Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p1.1)\.
- \[19\]B\. Li, Y\. Shen, J\. Yang, Y\. Wang, J\. Ren, T\. Che, J\. Zhang, and Z\. Liu\(2023\)Sparse mixture\-of\-experts are domain generalizable learners\.InInternational Conference on Learning Representations,Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3)\.
- \[20\]D\. Li, Y\. Yang, Y\. Song, and T\. M\. Hospedales\(2017\)Deeper, broader and artier domain generalization\.InProceedings of the IEEE International Conference on Computer Vision,pp\. 5542–5550\.Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[21\]B\. Liu, K\. Tian, W\. Wang, Z\. Zhang, L\. Qiao, and D\. Li\(2026\)Alloc\-MoE: budget\-aware expert activation allocation for efficient mixture\-of\-experts inference\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9653–9667\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.437)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[22\]S\. Long, Q\. Zhou, C\. Ying, L\. Ma, and Y\. Luo\(2024\)Rethinking domain generalization: discriminability and generalizability\.IEEE Transactions on Circuits and Systems for Video Technology34\(11\),pp\. 11783–11797\.External Links:[Document](https://dx.doi.org/10.1109/TCSVT.2024.3422887)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3)\.
- \[23\]H\. Nguyen, P\. Akbarian, H\. T\. Pham, T\. T\. N\. Vu, S\. Zhang, and N\. Ho\(2025\)Statistical advantages of perturbing cosine router in mixture of experts\.InInternational Conference on Learning Representations,Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3)\.
- \[24\]S\. Park and N\. Park\(2026\)How many experts are enough? towards optimal semantic specialization for mixture\-of\-experts\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 24792–24800\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i29.39665)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[25\]X\. Peng, Q\. Bai, X\. Xia, Z\. Huang, K\. Saenko, and B\. Wang\(2019\)Moment matching for multi\-source domain adaptation\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 1406–1415\.Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[26\]Z\. Qiu, Z\. Huang, and J\. Fu\(2024\)Unlocking emergent modularity in large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 2638–2660\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.144)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px2.p1.3),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[27\]P\. Rajpurkar, J\. Zhang, K\. Lopyrev, and P\. Liang\(2016\)SQuAD: 100,000\+ questions for machine comprehension of text\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,pp\. 2383–2392\.External Links:[Document](https://dx.doi.org/10.18653/v1/D16-1264)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[28\]N\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. Le, G\. Hinton, and J\. Dean\(2017\)Outrageously large neural networks: the sparsely\-gated mixture\-of\-experts layer\.InInternational Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p1.1)\.
- \[29\]N\. Tastan, S\. Laskaridis, K\. Nandakumar, and S\. Horvath\(2026\)MoSE: mixture of slimmable experts for efficient and adaptive language models\.InProceedings of the Forty\-third International Conference on Machine Learning \(ICML\),External Links:[Link](https://openreview.net/forum?id=18C6xMcD96)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-B](https://arxiv.org/html/2608.10392#S5.SS2.p1.1)\.
- \[30\]H\. Touvron, M\. Cord, M\. Douze, F\. Massa, A\. Sablayrolles, and H\. Jégou\(2021\)Training data\-efficient image transformers and distillation through attention\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 10347–10357\.Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[31\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 5998–6008\.Cited by:[§II\-B](https://arxiv.org/html/2608.10392#S2.SS2.SSS0.Px1.p1.2)\.
- \[32\]H\. Venkateswara, J\. Eusebio, S\. Chakraborty, and S\. Panchanathan\(2017\)Deep hashing network for unsupervised domain adaptation\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 5018–5027\.Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[33\]A\. Wang, A\. Singh, J\. Michael, F\. Hill, O\. Levy, and S\. R\. Bowman\(2019\)GLUE: a multi\-task benchmark and analysis platform for natural language understanding\.InInternational Conference on Learning Representations,Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[34\]A\. Warstadt, A\. Singh, and S\. R\. Bowman\(2019\)Neural network acceptability judgments\.Transactions of the Association for Computational Linguistics7,pp\. 625–641\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00290)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[35\]A\. Williams, N\. Nangia, and S\. R\. Bowman\(2018\)A broad\-coverage challenge corpus for sentence understanding through inference\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\),pp\. 1112–1122\.External Links:[Document](https://dx.doi.org/10.18653/v1/N18-1101)Cited by:[§IV\-A](https://arxiv.org/html/2608.10392#S4.SS1.SSS0.Px1.p1.4)\.
- \[36\]S\. Wu, A\. Lv, R\. Xie, X\. Sun, D\. Wang, R\. Yan, and Y\. Lin\(2026\)Union\-of\-experts: neurons in mixture\-of\-experts are secretly routers\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 36193–36206\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1675)Cited by:[§I](https://arxiv.org/html/2608.10392#S1.p3.1),[§V\-A](https://arxiv.org/html/2608.10392#S5.SS1.p1.1)\.
## Appendix AAdditional Experimental Details
### A\-AEvaluation Protocol and Statistical Reporting
#### DomainBed\.
We follow the train\-validation selection criterion\. Each domain, referred to as an environment in DomainBed, is split into an 80% in\-split and a 20% out\-split\. For each held\-out test environment, the in\-splits of the remaining environments are used for training, and the checkpoint with the highest mean out\-split accuracy across these training environments is selected\. Its accuracy on the held\-out environment’s in\-split is reported as the test result\. We repeat this procedure with every environment held out, average the resulting environment accuracies to obtain one dataset score, and report the mean and standard deviation over three independently seeded runs\. The run seed controls the data split and the Python, NumPy, and PyTorch random generators; deterministic cuDNN execution is enabled\.
#### GLUE\.
We use the official training and validation splits and select the best epoch according to the validation score\. CoLA is evaluated by Matthews correlation, MRPC by the mean of accuracy and F1, and QNLI, MNLI, and RTE by accuracy; MNLI uses the matched validation split\. For each task and seed, the learning rate is selected from\{2×10−5,3×10−5,5×10−5\}\\\{2\\times 10^\{\-5\},3\\times 10^\{\-5\},5\\times 10^\{\-5\}\\\}using validation performance\. We report the mean and standard deviation over three independently seeded runs\. The same seed is passed to the Python, NumPy, and PyTorch generators through Hugging Face Accelerate\.
#### Performance Stability\.
Tables[V](https://arxiv.org/html/2608.10392#A1.T5)and[VI](https://arxiv.org/html/2608.10392#A1.T6)expose the variation terms omitted from the compact main tables\. For UniF\-MoE, each entry is the mean±\\pmstandard deviation over three runs\. Baseline variation is retained when it is available from the cited or reproduced result; an entry without a±\\pmterm was reported without a standard deviation, and “–” denotes an unavailable result\. These variation estimates describe run\-to\-run stability\.
TABLE V:Out\-of\-domain accuracy \(%\) with the variation available for the DomainBed main results\.TABLE VI:GLUE task scores \(%\) with the standard deviations omitted from the compact main table\.
### A\-BArchitecture and Optimization Details
#### Vision configuration\.
The vision backbone is ImageNet\-pretrained DeiT\-S/16 with model dimensiond=384d=384and FFN widthH=1536H=1536\. Transformer layers 8 and 10 under zero\-based indexing are converted to UniF\-MoE layers\. Each converted layer contains one shared expert,K=6K=6residual experts, andB=8B=8blocks per expert\. We use hidden dropout 0\.1, stochastic\-depth rate 0\.1\. Images are normalized with ImageNet statistics\. Training\-domain images use random resized cropping to224×224224\\times 224, horizontal flipping, color jitter, and random grayscale; validation and test images are resized directly to224×224224\\times 224\. Table[VII](https://arxiv.org/html/2608.10392#A1.T7)lists the final DomainBed optimization settings\.
TABLE VII:Final DomainBed optimization settings\. All runs use Adam, a batch size of 32 per training environment, and an evaluation batch size of 64\.
#### Held\-out environments\.
PACS contains Art, Cartoon, Photo, and Sketch; VLCS contains Caltech101, LabelMe, SUN09, and VOC2007; OfficeHome contains Art, Clipart, Product, and Real World; TerraIncognita contains locations L100, L38, L43, and L46; and DomainNet contains Clipart, Infograph, Painting, Quickdraw, Real, and Sketch\. Each domain is used once as the held\-out test environment\.
#### Language configuration\.
The language backbone is BERT\-large\-cased\. Transformer layers 20 and 22 under zero\-based indexing are converted to UniF\-MoE layers withK=16K=16residual experts andB=16B=16blocks\. We train for at most 10 epochs with FP16 mixed precision, dynamic padding, maximum sequence length 128, training and evaluation batch sizes of 32, and no gradient accumulation\. Optimization uses AdamW with zero weight decay, a linear learning\-rate schedule, and no warm\-up\. The best validation epoch is retained, and training stops after four consecutive epochs without improvement\.
### A\-CComputing Environment
Experiments are run on Ubuntu 20\.04\.2 LTS workers equipped with an AMD EPYC 75F3 CPU, 503 GiB of host memory, and an NVIDIA GeForce RTX 3090 GPU\. The software environment uses Python 3\.8\.20, PyTorch 2\.4\.1 with CUDA 12\.1 and cuDNN 9\.1, torchvision 0\.19\.1, timm 1\.0\.20, Tutel 0\.2, Transformers 4\.46\.3, Datasets 3\.1\.0, Evaluate 0\.4\.6, and Accelerate 1\.0\.1\.
## Appendix BGenerative AI Use
Generative AI tools were used solely for language editing and polishing\. The authors reviewed and verified all resulting text and take full responsibility for the content of the paper\.Similar Articles
@_avichawla: https://x.com/_avichawla/status/2100876555409039605
The article explains the engineering aspects of Mixture-of-Experts (MoE) inference, detailing token routing, expert batching, GPU distribution, and performance trade-offs for efficient serving.
ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression
ConMoE proposes a train-free prototype remapping framework for Mixture-of-Experts (MoE) compression, which selects a subset of experts as reusable prototypes and deterministically remaps original expert calls to them, reducing memory usage without weight updates or fine-tuning.
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
UniMoMo compresses MoE-based recommendation models by merging experts based on functional behavior and routing traffic, preserving quality while speeding up inference.
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
Mix-MoE proposes a mixed Mixture-of-Experts framework with specialized expert groups and Fourier-transform-enhanced routing to mitigate parameter interference in multilingual machine translation, achieving significant improvements over baselines.
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
EntropyMoE introduces an entropy-aware Mixture-of-Experts architecture for tokenizer-free LLMs, using dynamic byte patches as routing units to enable sparse conditional computation. Experiments show it achieves the lowest held-out bits-per-byte among baselines while maintaining downstream accuracy.