MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

arXiv cs.CL Papers

Summary

This paper proposes MoEGen, a parameter-efficient fine-tuning framework that uses mixture-of-experts to generate instance-adaptive LoRA updates via expert codes and a lightweight hypernetwork, improving performance on commonsense reasoning benchmarks without storing separate adapters per expert.

arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool. We ask whether MoE-based PEFT can produce instance-specific adaptations without explicitly storing a separate LoRA module for each expert. To address this gap, we propose MoEGen, an adaptation framework that shifts MoE-based PEFT from expert selection to expert-conditioned parameter generation. Instead of storing each expert as a full LoRA adapter, MoEGen represents each expert as a small learnable vector, termed an expert code. It routes each input over these vectors and uses their weighted combination to condition a lightweight hypernetwork that generates input-specific low-rank updates. This design decouples expert capacity from adapter storage while enabling instance-conditioned adaptation. Experiments on eight commonsense reasoning benchmarks show consistent improvements over strong static and MoE-based PEFT baselines across three backbones. MoEGen also performs strongly in joint medical and legal-domain adaptation.
Original Article
View Cached Full Text

Cached at: 08/05/26, 07:44 AM

# Mixture-of-Experts for Instance-Adaptive LoRA Generation
Source: [https://arxiv.org/html/2608.03275](https://arxiv.org/html/2608.03275)
Yiming Zeng1,∗,Lei Lu2,∗,Zexin Li3,Zhuochun Li4,Shuoqiu Li5, Shuyi Liao1,Xidong Wu6,Zeyu Zhang7,Minmei Wang1, Yu Zhao8,Tingting Yu1,†,Shangqian Gao5,† 1University of Connecticut2Northeastern University3Nanyang Technological University 4University of Pittsburgh5Florida State University6Google 7Amazon AGI8University of Cincinnati

###### Abstract

Parameter\-efficient fine\-tuning \(PEFT\) enables efficient adaptation of large language models, but existing MoE\-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool\. We ask whether MoE\-based PEFT can produce instance\-specific adaptations without explicitly storing a separate LoRA module for each expert\. To address this gap, we proposeMoEGen, an adaptation framework that shifts MoE\-based PEFT from expert selection to expert\-conditioned parameter generation\. Instead of storing each expert as a full LoRA adapter, MoEGen represents each expert as a small learnable vector, termed an expert code\. It routes each input over these vectors and uses their weighted combination to condition a lightweight hypernetwork that generates input\-specific low\-rank updates\. This design decouples expert capacity from adapter storage while enabling instance\-conditioned adaptation\. Experiments on eight commonsense reasoning benchmarks show consistent improvements over strong static and MoE\-based PEFT baselines across three backbones\. MoEGen also performs strongly in joint medical and legal\-domain adaptation\.

MoEGen: Mixture\-of\-Experts for Instance\-Adaptive LoRA Generation

11footnotetext:Equal contribution\.22footnotetext:Corresponding authors\.![Refer to caption](https://arxiv.org/html/2608.03275v1/x1.png)Figure 1:Comparison between traditional MoE and MoEGen\. MoEGen replaces full feed\-forward experts with compact style embeddings and a shared hypernetwork for input\-specific LoRA generation\.## 1Introduction

Modern pretrained LLMs demonstrate remarkable performance across a wide spectrum of natural language processing tasksFeduset al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib43)\); Touvronet al\.\([2023a](https://arxiv.org/html/2608.03275#bib.bib1)\); Chowdheryet al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib42)\); OLMoet al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib45)\)\. However, for challenging domain\-specific downstream tasks, e\.g\., code generation and mathematical reasoningRozièreet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib50)\); Lewkowyczet al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib46)\), model adaptation is a standard way to further improve task performance\.OpenAIet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib4)\); Touvronet al\.\([2023a](https://arxiv.org/html/2608.03275#bib.bib1)\)\. While full supervised fine\-tuning is an effective model adaption method, it is computationally expensive and requires storing separate task\-specific model copies, making it less practical in resource\-constrained settingsAghajanyanet al\.\([2021](https://arxiv.org/html/2608.03275#bib.bib37)\); Lialinet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib36)\)\.

Parameter\-efficient fine\-tuning \(PEFT\) methods provide practical alternatives for model adaptation by updating only a small subset of model parametersLiet al\.\([2023](https://arxiv.org/html/2608.03275#bib.bib40)\); Tianet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib48)\); Huet al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib25)\); Liuet al\.\([2024b](https://arxiv.org/html/2608.03275#bib.bib29)\)\. LoRAHuet al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib25)\), for example, learns low\-rank weight updates while keeping the backbone parameters frozen\. Recent MoE\-LoRA methods further expand adaptation capacity by selecting or combining multiple LoRA experts through a learned routerDouet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib32)\); Liuet al\.\([2024a](https://arxiv.org/html/2608.03275#bib.bib34)\); Gaoet al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib31)\)\. However, because each expert stores a complete set of low\-rank parameters, adapter storage grows linearly with the size of the expert pool, making a large and diverse adaptation space costly to maintain\.

To address these limitations, we proposeMoEGen, a lightweight plug\-and\-play adaptation framework that synthesizes input\-specific LoRA parameters from compositional latent expert codes\. Instead of representing each expert as a full LoRA module, MoEGen encodes experts as small learnable vectors, termed expert codes, which are independent of the targeted module dimensions\. Prompt\-conditioned routing selects and combines multiple expert codes in this low dimensional latent space, and a shared hypernetwork synthesizes input\-specific LoRA parameter updates from the resultingexpert\-codemixture\. Unlike conventional MoE\-LoRA methods, where expert combinations are formed directly over stored LoRA parameter modules, MoEGen performs composition before parameter generation\. This decouples expert capacity from explicit backbone\-specific adapter storage and enables a larger adaptation space through lightweightexpert\-codecombinations\.

Overall, our contributions are as follows:

- •Lightweight expert parameterization\.MoEGen represents each expert as a low\-dimensional learnable code rather than a complete LoRA adapter, substantially reducing the parameter and storage costs of maintaining multiple experts\.
- •Dynamic input\-conditioned LoRA generation\.MoEGen dynamically generates an input\-specific low\-rank matrix for each adapted component\. This enables more targeted adaptation by tailoring the low\-rank update to the semantic characteristics of each input\.
- •Consistent empirical performance\.MoEGen improves over the strongest competing baselines by 0\.6–1\.1 points across eight commonsense reasoning benchmarks while updating only 0\.44–0\.52% of model parameters\. It further achieves improvements of 4\.6–6\.2 points on the cross\-domain joint benchmark\.

## 2Related Work

### 2\.1Parameter\-Efficient Fine\-Tuning

Parameter\-efficient fine\-tuning \(PEFT\) adapts pretrained models by updating only a small fraction of their parameters\. Adapter\-based methods\(Houlsbyet al\.,[2019](https://arxiv.org/html/2608.03275#bib.bib33); Pfeifferet al\.,[2021](https://arxiv.org/html/2608.03275#bib.bib39)\)insert trainable bottleneck modules into frozen transformers, while prefix\-tuning\(Li and Liang,[2021](https://arxiv.org/html/2608.03275#bib.bib35)\)prepends learnable tokens\. These methods may introduce inference latency or have limited expressiveness on complex tasksHanet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib49)\)\. LoRA\(Huet al\.,[2022](https://arxiv.org/html/2608.03275#bib.bib25)\)instead learns low\-rank weight updates in parallel with frozen weights\. Its extensions include adaptive rank allocation in AdaLoRA\(Zhanget al\.,[2023](https://arxiv.org/html/2608.03275#bib.bib30)\), directional and magnitude decomposition in DoRA\(Liuet al\.,[2024b](https://arxiv.org/html/2608.03275#bib.bib29)\), and quantization\-aware initialization in LoftQ\(Liet al\.,[2023](https://arxiv.org/html/2608.03275#bib.bib40)\)\. Sparse alternatives such as SpIELAnsellet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib8)\)and SMTHeet al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib24)\)update selected parameters or sub\-matrices\. However, these methods generally learn fixed updates that are applied uniformly across inputs, whereas MoEGen generates input\-conditioned updates\.

### 2\.2Mixture\-of\-Experts

Mixture\-of\-Experts \(MoE\) increases model capacity by sparsely routing inputs to expert modules\(Shazeeret al\.,[2017](https://arxiv.org/html/2608.03275#bib.bib41)\)\. Switch TransformerFeduset al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib43)\)and GLaMDuet al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib44)\)demonstrate the effectiveness of this sparse computation\. MoELoRALiuet al\.\([2024a](https://arxiv.org/html/2608.03275#bib.bib34)\)and LoRAMoEDouet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib32)\)extend MoE to PEFT by routing among multiple LoRA experts, while HydraLoRATianet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib48)\)introduces asymmetric expert structures\. These methods store full LoRA experts and perform routing inside transformer layers, coupling expert capacity with adapter storage\. In contrast, MoEGen performs external routing over lightweight expert codes and generates input\-conditioned LoRA parameters through a shared hypernetwork\.

![Refer to caption](https://arxiv.org/html/2608.03275v1/x2.png)Figure 2:Overview of MoEGen\. A frozen sentence encoder maps the input promptxxto a semantic embeddingzz\. For each adapted componentll,zzis concatenated with a learnable component embeddingℓ\(l\)\\ell^\{\(l\)\}and projected ash\(l\)=Wr​\[z;ℓ\(l\)\]h^\{\(l\)\}=W\_\{r\}\[z;\\ell^\{\(l\)\}\]for sparse expert\-code routing\. The selected expert codes, together with component\-local coordinatespj\(l\)p\_\{j\}^\{\(l\)\}, condition a shared hypernetwork to generate candidate LoRAAAmatrices\. These candidates are mixed intoAmix\(l\)A\_\{\\mathrm\{mix\}\}^\{\(l\)\}and combined with the static matrixB\(l\)B^\{\(l\)\}to form the dynamic low\-rank update applied to the frozen weightW\(l\)W^\{\(l\)\}\.

## 3Method

### 3\.1Overview

MoEGen consists of four modules: a frozen pretrained LLM, a frozen prompt encoder, a sparse MoE router, and a shared LoRA hypernetwork\. For each input, the prompt encoder produces a context embedding, which is used by the router to select a small number of expert codes\. These expert codes do not store LoRA parameters directly\. Instead, they serve as compact conditioning vectors that guide the hypernetwork to generate dynamic LoRA matrices\.

We apply MoEGen to the attention projections of each decoder block, including the query, key, value, and output projections; we refer to each adapted projection as acomponent\. For a component with frozen weightW\(l\)∈ℝdout×dinW^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}, MoEGen generates a low\-rank updateΔ​W\(l\)∈ℝdout×din\\Delta W^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}and computes the adapted output as

y=W\(l\)​x\+αr​Δ​W\(l\)​x,y=W^\{\(l\)\}x\+\\frac\{\\alpha\}\{r\}\\Delta W^\{\(l\)\}x,\(1\)wherex∈ℝdinx\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\},y∈ℝdouty\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\},rris the LoRA rank, andα\\alphais the LoRA scaling factor\. The low\-rank update is factorized asΔ​W\(l\)=B\(l\)​\(A\(l\)\)⊤\\Delta W^\{\(l\)\}=B^\{\(l\)\}\(A^\{\(l\)\}\)^\{\\top\}withB\(l\)∈ℝdout×rB^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times r\}andA\(l\)∈ℝdin×rA^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\\times r\}, wherer≪min⁡\(din,dout\)r\\ll\\min\(d\_\{\\mathrm\{in\}\},d\_\{\\mathrm\{out\}\}\)\. The original LLM parameters are never modified\.

### 3\.2Per\-Component Semantic Routing with Compact Expert Codes

MoEGen maintains a single pool ofNNtrainable expert codes\{E1,…,EN\}\\\{E\_\{1\},\\ldots,E\_\{N\}\\\}withEi∈ℝdhE\_\{i\}\\in\\mathbb\{R\}^\{d\_\{h\}\}, shared across allLLcomponents\. The routing parameters remain compact because all components share a common gating projectionWgW\_\{g\}\.

For each input promptxx, we use a frozen sentence encoderZhanget al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib10)\)fenc:𝒳→ℝdzf\_\{\\mathrm\{enc\}\}:\\mathcal\{X\}\\rightarrow\\mathbb\{R\}^\{d\_\{z\}\}to obtain a semantic context embeddingz=fenc​\(x\)z=f\_\{\\mathrm\{enc\}\}\(x\)\. The encoder remains fixed during training to preserve general semantic representations and reduce optimization cost\. A lightweight routerfrf\_\{\\mathrm\{r\}\}then maps this shared semantic embedding into component\-specific routing features:

fr​\(z\)\\displaystyle f\_\{\\mathrm\{r\}\}\(z\)=\[h\(1\),…,h\(L\)\],\\displaystyle=\[h^\{\(1\)\},\\ldots,h^\{\(L\)\}\],\(2\)H\\displaystyle H=\[\(h\(1\)\)⊤⋯\(h\(L\)\)⊤\]⊤∈ℝL×dh\.\\displaystyle=\\begin\{bmatrix\}\(h^\{\(1\)\}\)^\{\\top\}&\\cdots&\(h^\{\(L\)\}\)^\{\\top\}\\end\{bmatrix\}^\{\\top\}\\in\\mathbb\{R\}^\{L\\times d\_\{h\}\}\.
whereLLis the number of adapted LoRA components \(i\.e\., target projections across decoder layers\), and each rowh\(l\)∈ℝdhh^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{h\}\}denotes the routing feature for componentll\.

To enable component\-specific specialization, we associate each component with a learnable embeddingℓ\(l\)∈ℝdℓ\\ell^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\ell\}\}and concatenate it with the shared prompt embedding:

h\(l\)=Wr​\[z;ℓ\(l\)\],h^\{\(l\)\}=W\_\{r\}\[z;\\ell^\{\(l\)\}\],\(3\)where\[⋅;⋅\]\[\\cdot;\\cdot\]denotes vector concatenation andWr∈ℝdh×\(dz\+dℓ\)W\_\{r\}\\in\\mathbb\{R\}^\{d\_\{h\}\\times\(d\_\{z\}\+d\_\{\\ell\}\)\}is a shared projection matrix\. This formulation allows the router to model interactions between prompt\-level semantics and component\-specific identities, while requiring onlydh​\(dz\+dℓ\)\+L​dℓd\_\{h\}\(d\_\{z\}\+d\_\{\\ell\}\)\+Ld\_\{\\ell\}parameters\.

Each expert codeEiE\_\{i\}is a compact conditioning vector rather than a full LoRA adapter: it encodes an adaptation style that the hypernetwork \(Sec\.[3\.3](https://arxiv.org/html/2608.03275#S3.SS3)\) later expands into component\-specific LoRA matrices\. For each componentll, the routing featureh\(l\)h^\{\(l\)\}is first passed through a nonlinear block and then projected to expert logits:

h~\(l\)\\displaystyle\\tilde\{h\}^\{\(l\)\}=GELU​\(LN​\(h\(l\)\)\),\\displaystyle=\\mathrm\{GELU\}\\\!\\left\(\\mathrm\{LN\}\\\!\\left\(h^\{\(l\)\}\\right\)\\right\),\(4\)g\(l\)\\displaystyle g^\{\(l\)\}=Wg​h~\(l\)\+bg\.\\displaystyle=W\_\{g\}\\tilde\{h\}^\{\(l\)\}\+b\_\{g\}\.We then perform top\-kkselection and normalize the selected logits:

𝒯\(l\)\\displaystyle\\mathcal\{T\}^\{\(l\)\}=TopK⁡\(g\(l\),k\),\\displaystyle=\\operatorname\{TopK\}\\\!\\left\(g^\{\(l\)\},k\\right\),\(5\)wi\(l\)\\displaystyle w\_\{i\}^\{\(l\)\}=exp⁡\(gi\(l\)\)∑j∈𝒯\(l\)exp⁡\(gj\(l\)\),i∈𝒯\(l\)\.\\displaystyle=\\frac\{\\exp\\\!\\left\(g\_\{i\}^\{\(l\)\}\\right\)\}\{\\sum\_\{j\\in\\mathcal\{T\}^\{\(l\)\}\}\\exp\\\!\\left\(g\_\{j\}^\{\(l\)\}\\right\)\},\\quad i\\in\\mathcal\{T\}^\{\(l\)\}\.Here,𝒯\(l\)⊆\{1,…,N\}\\mathcal\{T\}^\{\(l\)\}\\subseteq\\\{1,\\ldots,N\\\}is the set of indices corresponding to thekklargest entries ofg\(l\)g^\{\(l\)\}\. Experts outside𝒯\(l\)\\mathcal\{T\}^\{\(l\)\}are assigned zero weight\.

### 3\.3Expert\-Conditioned LoRA Generation

For each adapted componentll, MoEGen dynamically generates the LoRAAAmatrix while maintaining a trainable staticBBmatrix\. This asymmetric design keeps the generated part lightweight: the input\-conditioned part is synthesized by the hypernetwork, while the output projection side is shared across inputs\.

In MoEGen,expert codesare used as conditioning vectors for parameter generation\. We introduce a component\-local embeddingP\(l\)∈ℝdin×dhP^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\\times d\_\{h\}\}for each adapted component, where thejj\-th rowpj\(l\)∈ℝdhp\_\{j\}^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{h\}\}serves as a learnable coordinate for thejj\-th input dimension\.

Given a selected expert codeEi∈ℝdhE\_\{i\}\\in\\mathbb\{R\}^\{d\_\{h\}\}, we use a shared hypernetwork to generate the corresponding LoRAAAmatrix row by row\. The hypernetwork is shared across all components, experts, and input dimensions\. For thejj\-th input dimension of thell\-th target component, it takes the concatenation of a learnable coordinate embeddingpj\(l\)∈ℝdhp\_\{j\}^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{h\}\}and the expert codeEiE\_\{i\}, and outputs one row of the LoRAAAmatrix:

ai,j\(l\)=Linear​\(\[pj\(l\);Ei\]\),ai,j\(l\)∈ℝr\.a\_\{i,j\}^\{\(l\)\}=\\mathrm\{Linear\}\\left\(\[p\_\{j\}^\{\(l\)\};E\_\{i\}\]\\right\),\\qquad a\_\{i,j\}^\{\(l\)\}\\in\\mathbb\{R\}^\{r\}\.\(6\)The coordinatepj\(l\)p\_\{j\}^\{\(l\)\}identifies which row is being generated, whileEiE\_\{i\}provides the expert\-specific conditioning signal\. The hypernetwork is implemented as a lightweight block consisting of LayerNorm, GELU activation, and a linear projectionℝ2​dh→ℝr\\mathbb\{R\}^\{2d\_\{h\}\}\\\!\\to\\\!\\mathbb\{R\}^\{r\}, contributing𝒪​\(dh​r\)\\mathcal\{O\}\(d\_\{h\}r\)parameters that do not scale with depth,dind\_\{\\mathrm\{in\}\},doutd\_\{\\mathrm\{out\}\}, or the number of expertsNN\. Stacking rows yields the candidate matrix for expertii:

Ai\(l\)=\[ai,1\(l\),…,ai,din\(l\)\]⊤∈ℝdin×r\.A\_\{i\}^\{\(l\)\}=\[a\_\{i,1\}^\{\(l\)\},\\ldots,a\_\{i,d\_\{\\mathrm\{in\}\}\}^\{\(l\)\}\]^\{\\top\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{in\}\}\\times r\}\.\(7\)
For the selected expert set𝒯\(l\)\\mathcal\{T\}^\{\(l\)\}, the candidate matrices are mixed according to the routing weights:

A\(l\)=∑i∈𝒯\(l\)wi\(l\)​Ai\(l\)\.A^\{\(l\)\}=\\sum\_\{i\\in\\mathcal\{T\}^\{\(l\)\}\}w\_\{i\}^\{\(l\)\}A\_\{i\}^\{\(l\)\}\.\(8\)The final low\-rank update for componentllis computed as

Δ​W\(l\)=B\(l\)​\(A\(l\)\)⊤,\\Delta W^\{\(l\)\}=B^\{\(l\)\}\\left\(A^\{\(l\)\}\\right\)^\{\\top\},\(9\)whereB\(l\)∈ℝdout×rB^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times r\}andΔ​W\(l\)∈ℝdout×din\\Delta W^\{\(l\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}\. Following LoRAHuet al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib25)\),B\(l\)B^\{\(l\)\}is initialized to zero for stable training\.

Overall, the routed expert codes specify the adaptation style, and the component\-local coordinates specify where this style is applied\. As a result, increasing the number of experts expands the conditioning space without adding per\-component LoRA parameters for each new expert\. To see this, consider the per\-component parameter cost of the expert pool\. In standard MoE\-LoRA, each expert stores a full low\-rank adapter, giving a per\-component cost ofN⋅r⋅\(din\+dout\)N\\cdot r\\cdot\(d\_\{\\mathrm\{in\}\}\+d\_\{\\mathrm\{out\}\}\)that scales linearly withNN\. In MoEGen, the expert pool contributes onlyN⋅dhN\\cdot d\_\{h\}parameters, since each expert is a compact code of dimensiondhd\_\{h\}, and the component\-specific costdin⋅dh\+dout⋅rd\_\{\\mathrm\{in\}\}\\cdot d\_\{h\}\+d\_\{\\mathrm\{out\}\}\\cdot rfromP\(l\)P^\{\(l\)\}andB\(l\)B^\{\(l\)\}is independent ofNN\. Withdh≪r⋅dind\_\{h\}\\ll r\\cdot d\_\{\\mathrm\{in\}\}, expanding the expert pool grows the conditioning space at negligible cost, while the per\-component footprint stays constant\.

### 3\.4Training Objective

MoEGen is trained end\-to\-end with the standard causal language modeling objective\. We optimize three groups of parameters: \(i\) the router parametersWrW\_\{r\},\{ℓ\(l\)\}l=1L\\\{\\ell^\{\(l\)\}\\\}\_\{l=1\}^\{L\},WgW\_\{g\},bgb\_\{g\}; \(ii\) the expert codes\{Ei\}i=1N\\\{E\_\{i\}\\\}\_\{i=1\}^\{N\}; and \(iii\) the hypernetwork parametersϕLinear\\phi\_\{\\mathrm\{Linear\}\},\{P\(l\),B\(l\)\}l=1L\\\{P^\{\(l\)\},B^\{\(l\)\}\\\}\_\{l=1\}^\{L\}\. The base model weights remain frozen throughout training\. To avoid routing collapse and encourage balanced expert utilization, we add a load\-balancing regularization loss followingFeduset al\.\([2022](https://arxiv.org/html/2608.03275#bib.bib43)\)\. The full training objective is

ℒ=ℒLM\+λlb​ℒload,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{LM\}\}\+\\lambda\_\{\\mathrm\{lb\}\}\\mathcal\{L\}\_\{\\mathrm\{load\}\},\(10\)
where the language modeling loss isℒLM=−∑t=1Tlog⁡pθ​\(yt∣y<t,X\)\.\\mathcal\{L\}\_\{\\mathrm\{LM\}\}=\-\\sum\_\{t=1\}^\{T\}\\log p\_\{\\theta\}\(y\_\{t\}\\mid y\_\{<t\},X\)\.given an input\-output pair\(X,Y\)\(X,Y\),λlb\\lambda\_\{\\mathrm\{lb\}\}controls the strength of the load\-balancing regularization\. The load\-balancing loss is defined as

ℒload=NL​∑l=1L∑i=1N\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{load\}\}=\\frac\{N\}\{L\}\\sum\_\{l=1\}^\{L\}\\sum\_\{i=1\}^\{N\}\(1B​k​∑b=1B𝟏​\[i∈𝒯b\(l\)\]\)\\displaystyle\\left\(\\frac\{1\}\{Bk\}\\sum\_\{b=1\}^\{B\}\\mathbf\{1\}\[i\\in\\mathcal\{T\}\_\{b\}^\{\(l\)\}\]\\right\)\(11\)×\(1B​∑b=1Bwb,i\(l\)\)\.\\displaystyle\\times\\left\(\\frac\{1\}\{B\}\\sum\_\{b=1\}^\{B\}w\_\{b,i\}^\{\(l\)\}\\right\)\.
where𝒯b\(l\)\\mathcal\{T\}\_\{b\}^\{\(l\)\}denotes the top\-kkexpert set selected for thebb\-th input at componentll, andwb,i\(l\)w\_\{b,i\}^\{\(l\)\}is the normalized routing weight for expertii, withwb,i\(l\)=0w\_\{b,i\}^\{\(l\)\}=0ifi∉𝒯b\(l\)i\\notin\\mathcal\{T\}\_\{b\}^\{\(l\)\}\. This regularizer couples the discrete routing decisions with the differentiable routing probabilities, encouraging more balanced expert usage across both inputs and components\.

Table 1:Comparison of different parameter\-efficient fine\-tuning methods across commonsense reasoning benchmarks\. Bold indicates the best result per base model\. All methods use the Alpaca prompt format\. \#Params\(%\) denotes the percentage of trainable parameters relative to the base model\.

## 4Experiment

### 4\.1Experiment Setup

#### Datasets\.

FollowingHeet al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib24)\), we useCommonsense\-170K, which combines the training splits of eight commonsense reasoning benchmarks: BoolQClarket al\.\([2019](https://arxiv.org/html/2608.03275#bib.bib23)\), PIQABisket al\.\([2019](https://arxiv.org/html/2608.03275#bib.bib21)\), SIQASapet al\.\([2019](https://arxiv.org/html/2608.03275#bib.bib20)\), HellaSwagZellerset al\.\([2019](https://arxiv.org/html/2608.03275#bib.bib19)\), WinoGrandeSakaguchiet al\.\([2019](https://arxiv.org/html/2608.03275#bib.bib18)\), ARC\-EasyClarket al\.\([2018](https://arxiv.org/html/2608.03275#bib.bib17)\), ARC\-ChallengeClarket al\.\([2018](https://arxiv.org/html/2608.03275#bib.bib17)\), and OpenBookQAMihaylovet al\.\([2018](https://arxiv.org/html/2608.03275#bib.bib15)\)\. We evaluate each method on the held\-out test set of each benchmark\.

We also use a mixed\-task dataset composed of four biomedical and legal\-domain tasks: MedNLIRomanov and Shivade \([2018](https://arxiv.org/html/2608.03275#bib.bib14)\), PubMedQAJinet al\.\([2019](https://arxiv.org/html/2608.03275#bib.bib13)\), HQSBen Abacha and Demner\-Fushman \([2019](https://arxiv.org/html/2608.03275#bib.bib12)\), and BillSumKornilova and Eidelman \([2019](https://arxiv.org/html/2608.03275#bib.bib11)\)\. MedNLI and PubMedQA are evaluated by accuracy, while HQS and BillSum are evaluated by ROUGE\-1, ROUGE\-2, and ROUGE\-LLin \([2004](https://arxiv.org/html/2608.03275#bib.bib9)\)\. For this setting, theAvgcolumn averages one primary metric per task: accuracy for MedNLI and PubMedQA, and ROUGE\-L for HQS and BillSum\.

#### Implementation details\.

MoEGen uses a frozen Qwen3\-Embedding\-0\.6BZhanget al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib10)\)sentence encoder\. We useN=8N\{=\}8expert codes with code dimensiondh=32d\_\{h\}\{=\}32, top\-k=2k\{=\}2routing, and LoRA rankr=64r\{=\}64, and apply the generated adapters to the Q/K/V/O projections of every decoder block\. All methods are trained for33epochs with AdamWLoshchilov and Hutter \([2017](https://arxiv.org/html/2608.03275#bib.bib7)\)and bfloat16 precision\. The effective batch size is1616for commonsense reasoning and6464for mixed\-task joint training\. Additional details are provided in Appendix\.

### 4\.2Commonsense Reasoning

Table[1](https://arxiv.org/html/2608.03275#S3.T1)reports the results across LLaMA\-2\-7BTouvronet al\.\([2023b](https://arxiv.org/html/2608.03275#bib.bib28)\), LLaMA\-3\-8BGrattafioriet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib26)\), and Qwen3\-8B\-baseYanget al\.\([2025](https://arxiv.org/html/2608.03275#bib.bib5)\)\. MoEGen achieves the highest average score on all three backbones, reaching84\.284\.2,87\.987\.9, and89\.989\.9, respectively\. Compared with the strongest static PEFT baseline on each backbone, MoEGen improves the average score by2\.42\.4points over SMT on LLaMA\-2\-7B, by1\.11\.1points over SMT on LLaMA\-3\-8B, and by0\.80\.8points over DoRA on Qwen3\-8B\-base\. MoEGen trains only0\.52%0\.52\\%,0\.44%0\.44\\%, and0\.49%0\.49\\%of the parameters on the three backbones, respectively, showing that these improvements do not result from a larger trainable parameter budget\.

Compared with standard LoRA, MoEGen achieves a higher score on every evaluated benchmark across all three backbones\. On LLaMA\-2\-7B, it improves the average score from77\.677\.6to84\.284\.2\. On LLaMA\-3\-8B, the average score increases from80\.880\.8to87\.987\.9\. On Qwen3\-8B\-base, where LoRA achieves an average score of87\.487\.4, MoEGen reaches89\.989\.9\. These results show that input\-conditioned adaptation provides consistent benefits across model families with different baseline accuracies\. Notably, the improvement remains positive even on Qwen3\-8B\-Base, where standard LoRA already performs strongly, suggesting that MoEGen captures useful input\-dependent specialization beyond a single static low\-rank update\.

MoEGen also achieves higher average scores than the MoE\-based PEFT methods MoELoRA and MoLA\. MoELoRA and MoLA increase adapter capacity by introducing multiple LoRA experts, with each expert corresponding to a fixed set of learned parameters\. In contrast, MoEGen routes among compact expert codes and uses the resulting conditioning signal to synthesize input\-specific LoRA matrices\. MoEGen improves over MoELoRA by0\.60\.6,1\.21\.2, and1\.91\.9points on LLaMA\-2\-7B, LLaMA\-3\-8B, and Qwen3\-8B\-base, respectively\. Compared with MoLA, the corresponding improvements are4\.84\.8,4\.64\.6, and1\.31\.3points\. The results indicate that generating input\-specific adapter parameters yields higher average accuracy than selecting or mixing a fixed set of adapter experts in these experiments\.

The comparison with full fine\-tuning on LLaMA\-2\-7B further demonstrates the parameter efficiency of MoEGen\. Full fine\-tuning updates all model parameters and achieves an average score of82\.282\.2, whereas MoEGen reaches84\.284\.2while training only0\.52%0\.52\\%of the parameters\. This comparison shows that input\-adaptive parameter updates can outperform a global update of the entire model on the evaluated commonsense reasoning benchmarks\.

Table 2:Comparison on the NLP Joint benchmark with LLaMA\-3\-8B, Qwen3\-8B\-Base, and LLaMA\-2\-13B at epoch 3\. Best results per backbone are inbold\. Avg averages Acc for MedNLI/PubMedQA and R\-L for BillSum/HQS\.
### 4\.3Cross\-Domain Joint Training

To evaluate MoEGen under heterogeneous multi\-task adaptation, we jointly train on MedNLI, PubMedQA, HQS, and BillSum followingLuet al\.\([2024](https://arxiv.org/html/2608.03275#bib.bib2)\)\. These tasks cover natural language inference, question answering, and clinical and legal summarization\. All methods use the same training data and optimization protocol and are evaluated separately on each task\. Table[2](https://arxiv.org/html/2608.03275#S4.T2)reports the results\.

On LLaMA\-3\-8B, MoEGen achieves the highest score on every metric and improves the average score by6\.26\.2points over MoLA, from55\.3%55\.3\\%to61\.5%61\.5\\%\. On BillSum, MoEGen reaches43\.7%43\.7\\%ROUGE\-L, compared with30\.3%30\.3\\%for the best baseline\. Plain LoRA performs worst: it drops to57\.1%57\.1\\%on MedNLI and15\.8%15\.8\\%on HQS, which indicates that a single static update struggles to accommodate the competing adaptation requirements of all four tasks\.

The advantage holds on Qwen3\-8B\-Base\. MoEGen reaches the highest average of60\.4%60\.4\\%, compared with55\.8%55\.8\\%for the strongest baseline \(MoLA\), obtains the best MedNLI accuracy, and matches the best PubMedQA accuracy\. On BillSum it reaches42\.4%42\.4\\%ROUGE\-L against25\.2%25\.2\\%for the best baseline, while on HQS it stays within0\.70\.7points of the best result\.

MoEGen also achieves the highest average score on LLaMA\-2\-13B, reaching58\.8%58\.8\\%compared with53\.8%53\.8\\%for MoELoRA and MoLA\. It obtains the highest MedNLI and PubMedQA accuracy and improves BillSum ROUGE\-L from18\.8%18\.8\\%to38\.6%38\.6\\%\. On HQS, its scores remain within1\.21\.2points of the best baseline across all three ROUGE metrics\. Together, the results on all three backbones show that MoEGen consistently improves overall performance under cross\-domain joint training\. These results suggest that the bottleneck in this setting is not adapter capacity but how the adapter is conditioned on the input: the baselines remain within a narrow BillSum range despite their differing adapter structures, while MoEGen generates input\-conditioned updates and improves ROUGE\-L on every backbone\.

![Refer to caption](https://arxiv.org/html/2608.03275v1/figure/fig_3_q.png)Figure 3:Case study of MoEGen routing on eight commonsense reasoning test sets with 50 prompts per dataset and 400 cases in total\.\(a\)t\-SNE of the routed fused style produced by the hypernetwork\.\(b\)t\-SNE of the router latent representation used by the gate\.
### 4\.4Ablation Study

#### Number of Expert\.

We study the effect of the expert\-code pool sizeN∈\{2,4,8,16\}N\\in\\\{2,4,8,16\\\}on the cross\-domain joint benchmark using LLaMA\-3\-8B\. All variants follow the setup in Section[4\.3](https://arxiv.org/html/2608.03275#S4.SS3), withr=64r\{=\}64,dh=32d\_\{h\}\{=\}32, top\-22routing, an effective batch size of6464, and33training epochs\. As shown in Table[3](https://arxiv.org/html/2608.03275#S4.T3),N=8N\{=\}8achieves the highest average score and performs best on three of the four tasks\. Smaller pools may provide insufficient capacity for heterogeneous domains, whereas a larger pool does not yield further improvements\. Notably, whenN=2N\{=\}2under top\-22routing, both codes are always selected, reducing the router to weighting a fixed code pair\. Its lower performance suggests that input\-dependent code selection contributes to the effectiveness of MoEGen\. We therefore useN=8N\{=\}8for the cross\-domain benchmark\. An ablation study on the routing design and shared expert is also provided in Appendix[B](https://arxiv.org/html/2608.03275#A2)\.

Table 3:Effect of the number of expert codes on the cross\-domain joint benchmark with LLaMA\-3\-8B\.
#### Context encoder\.

We next ablate the source of the routing contextzz\. By default,zzis produced by a frozen Qwen3\-Embedding\-0\.6B sentence encoder\. TheMoEGen \(w/o embedding\)variant removes this external encoder entirely and instead reuses the frozen backbone’s last\-layer final\-token representation of the prompt aszz, leaving all other components unchanged\. As shown in Table[4](https://arxiv.org/html/2608.03275#S4.T4)\(a\), the w/o embedding variant remains competitive but consistently trails the full model, reducing the average score from61\.5%61\.5\\%to60\.8%60\.8\\%, with the largest drops on PubMedQA \(77\.8%→76\.4%77\.8\\%\\rightarrow 76\.4\\%\) and HQS \(33\.3%→32\.4%33\.3\\%\\rightarrow 32\.4\\%\)\. This suggests that the dedicated sentence encoder provides a cleaner semantic routing signal than the backbone’s own hidden states\. Table[4](https://arxiv.org/html/2608.03275#S4.T4)\(b\) compares the computational cost of the two variants\. Keeping Qwen3\-Embedding\-0\.6B is not a bottleneck on either side: training is8%8\\%faster with the external encoder, since encoding prompts with the 0\.6B model is cheaper than the additional forward pass through the 8B backbone required by the w/o embedding variant, and per\-example inference latency differs by only0\.60\.6ms \(\+2\.4%\+2\.4\\%\) for the same reason\. The only advantage of removing the encoder is deployment footprint, saving its0\.60\.6B parameters and2\.22\.2GB of GPU memory\. Given the consistent quality gains at negligible time cost, we retain the external sentence encoder as the default source of the routing context\.

\(a\)Task performance\.
\(b\)Computational cost\.

Table 4:Ablation study on the context encoder\. The w/o sentence encoder variant removes the frozen Qwen3\-Embedding\-0\.6B encoder and routes on the backbone’s own last\-token representation\.

### 4\.5Qualitative Study

We visualize the routing behavior of MoEGen on eight commonsense reasoning benchmarks\. For each task, we randomly sample 50 test prompts and apply t\-SNE to the router inputs and the fused expert representations\.

As shown in Figure[3](https://arxiv.org/html/2608.03275#S4.F3), the router inputs already exhibit task\-dependent structure\. Distinctive tasks such as SocialIQA, WinoGrande, and BoolQ form relatively compact regions, whereas science QA benchmarks such as ARC\-Challenge, ARC\-Easy, and OpenBookQA are more closely distributed\.

After expert mixing, the fused representations preserve this organization while connecting related tasks more smoothly\. This suggests that the router captures both task\-specific adaptation patterns and shared structures across similar reasoning tasks\.

Additionally, Table[7](https://arxiv.org/html/2608.03275#A2.T7)presents a representative example from BillSum\. The reference summary describes the bill by naming the act and outlining its main provisions\. In contrast, LoRA, MoE\-LoRA, and MoLA generate short legal fragments copied from the bill text, failing to capture the main actions of the legislation\. The outputs of MoE\-LoRA and MoLA are also highly similar, suggesting that simply adding MoE\-style adapters may still struggle to model diverse task behaviors under joint multi\-task adaptation\.

MoEGen produces a more complete summary\. It correctly identifies the bill title and recovers the key legislative actions, including the repeal of the excise tax and changes to fuel\-related taxes\. As a result, its prediction is much closer to the reference summary and achieves a substantially higher ROUGE\-L score\. Additional examples are provided in the Appendix[C](https://arxiv.org/html/2608.03275#A3)\.

#### Inference Efficiency\.

Table 5:Per\-sample inference latency \(s\)\. Max Tok\. denotes the maximum number of generated tokens used during evaluation\.We further evaluate the inference cost of MoEGen on the cross\-domain joint benchmark\. Table[5](https://arxiv.org/html/2608.03275#S4.T5)reports the per\-sample latency of each method using LLaMA\-2\-13B on a single B200 GPU with batch size 32\. MoEGen introduces an input\-dependent routing and generation cost, but the overhead remains small on longer\-generation tasks\. The largest relative gap appears on MedNLI, where the output is very short and the fixed routing cost is less amortized\. As the generation length increases, this gap becomes much smaller: MoEGen is only 9\.1% slower than LoRA on HQS, 8\.1% slower on PubMedQA, and 4\.1% slower on BillSum\. It also remains substantially faster than MoLA on the longer\-generation tasks\.

## 5Conclusion

This paper proposes an instance\-adaptive framework for parameter\-efficient LLM adaptation\. It aligns LoRA updates with input semantics through MoE\-guided routing and dynamic parameter generation driven by a lightweight hypernetwork\. By replacing full LoRA experts with compact expert codes, MoEGen expands the adaptation space while keeping the parameter cost low\. Experimental results show that MoEGen consistently outperforms static LoRA variants and existing MoE\-based PEFT methods across commonsense, biomedical, and legal\-domain benchmarks, demonstrating its advantages in accuracy and adaptation flexibility\.

## Limitations

MoEGen adds a lightweight routing and parameter\-generation module on top of standard LoRA\. Although this design keeps the number of trainable parameters small and introduces only modest inference overhead in our experiments, it still requires implementing an additional module beyond conventional static PEFT methods\. In addition, this work mainly studies instance\-adaptive LoRA generation under supervised fine\-tuning settings\. Extending the same idea to other adaptation scenarios, such as preference tuning or continual learning, is an interesting direction for future work\.

## References

- Intrinsic dimensionality explains the effectiveness of language model fine\-tuning\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 7319–7328\.External Links:[Link](https://aclanthology.org/2021.acl-long.568/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.568)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- A\. Ansell, I\. Vulić, H\. Sterz, A\. Korhonen, and E\. M\. Ponti \(2024\)Scaling sparse fine\-tuning to large language models\.External Links:2401\.16405,[Link](https://arxiv.org/abs/2401.16405)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- A\. Ben Abacha and D\. Demner\-Fushman \(2019\)On the summarization of consumer health questions\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 2228–2234\.External Links:[Link](https://aclanthology.org/P19-1215/),[Document](https://dx.doi.org/10.18653/v1/P19-1215)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p2.1)\.
- Y\. Bisk, R\. Zellers, R\. L\. Bras, J\. Gao, and Y\. Choi \(2019\)PIQA: reasoning about physical commonsense in natural language\.External Links:1911\.11641,[Link](https://arxiv.org/abs/1911.11641)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- A\. Chowdhery, S\. Narang, J\. Devlin, M\. Bosma, G\. Mishra, A\. Roberts, P\. Barham, H\. W\. Chung, C\. Sutton, S\. Gehrmann, P\. Schuh, K\. Shi, S\. Tsvyashchenko, J\. Maynez, A\. Rao, P\. Barnes, Y\. Tay, N\. Shazeer, V\. Prabhakaran, E\. Reif, N\. Du, B\. Hutchinson, R\. Pope, J\. Bradbury, J\. Austin, M\. Isard, G\. Gur\-Ari, P\. Yin, T\. Duke, A\. Levskaya, S\. Ghemawat, S\. Dev, H\. Michalewski, X\. Garcia, V\. Misra, K\. Robinson, L\. Fedus, D\. Zhou, D\. Ippolito, D\. Luan, H\. Lim, B\. Zoph, A\. Spiridonov, R\. Sepassi, D\. Dohan, S\. Agrawal, M\. Omernick, A\. M\. Dai, T\. S\. Pillai, M\. Pellat, A\. Lewkowycz, E\. Moreira, R\. Child, O\. Polozov, K\. Lee, Z\. Zhou, X\. Wang, B\. Saeta, M\. Diaz, O\. Firat, M\. Catasta, J\. Wei, K\. Meier\-Hellstern, D\. Eck, J\. Dean, S\. Petrov, and N\. Fiedel \(2022\)PaLM: scaling language modeling with pathways\.External Links:2204\.02311,[Link](https://arxiv.org/abs/2204.02311)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- C\. Clark, K\. Lee, M\. Chang, T\. Kwiatkowski, M\. Collins, and K\. Toutanova \(2019\)BoolQ: exploring the surprising difficulty of natural yes/no questions\.External Links:1905\.10044,[Link](https://arxiv.org/abs/1905.10044)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord \(2018\)Think you have solved question answering? try arc, the ai2 reasoning challenge\.arXiv:1803\.05457v1\.Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- S\. Dou, E\. Zhou, Y\. Liu, S\. Gao, W\. Shen, L\. Xiong, Y\. Zhou, X\. Wang, Z\. Xi, X\. Fan, S\. Pu, J\. Zhu, R\. Zheng, T\. Gui, Q\. Zhang, and X\. Huang \(2024\)LoRAMoE: alleviating world knowledge forgetting in large language models via MoE\-style plugin\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 1932–1945\.External Links:[Link](https://aclanthology.org/2024.acl-long.106/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.106)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.03275#S2.SS2.p1.1)\.
- N\. Du, Y\. Huang, A\. M\. Dai, S\. Tong, D\. Lepikhin, Y\. Xu, M\. Krikun, Y\. Zhou, A\. W\. Yu, O\. Firat, B\. Zoph, L\. Fedus, M\. Bosma, Z\. Zhou, T\. Wang, Y\. E\. Wang, K\. Webster, M\. Pellat, K\. Robinson, K\. Meier\-Hellstern, T\. Duke, L\. Dixon, K\. Zhang, Q\. V\. Le, Y\. Wu, Z\. Chen, and C\. Cui \(2022\)GLaM: efficient scaling of language models with mixture\-of\-experts\.External Links:2112\.06905,[Link](https://arxiv.org/abs/2112.06905)Cited by:[§2\.2](https://arxiv.org/html/2608.03275#S2.SS2.p1.1)\.
- W\. Fedus, B\. Zoph, and N\. Shazeer \(2022\)Switch transformers: scaling to trillion parameter models with simple and efficient sparsity\.External Links:2101\.03961,[Link](https://arxiv.org/abs/2101.03961)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.03275#S2.SS2.p1.1),[§3\.4](https://arxiv.org/html/2608.03275#S3.SS4.p1.7)\.
- C\. Gao, K\. Chen, J\. Rao, R\. Liu, B\. Sun, Y\. Zhang, D\. Peng, X\. Guo, and V\. Subrahmanian \(2025\)MoLA: MoE LoRA with layer\-wise expert allocation\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 5112–5127\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.284/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.284),ISBN 979\-8\-89176\-195\-7Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.2](https://arxiv.org/html/2608.03275#S4.SS2.p1.9)\.
- Z\. Han, C\. Gao, J\. Liu, J\. Zhang, and S\. Q\. Zhang \(2024\)Parameter\-efficient fine\-tuning for large models: a comprehensive survey\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=lIsCS8b6zj)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- H\. He, J\. B\. Li, X\. Jiang, and H\. Miller \(2025\)SMT: fine\-tuning large language models with sparse matrices\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=GbgCRJedQ7)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. de Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly \(2019\)Parameter\-efficient transfer learning for nlp\.External Links:1902\.00751,[Link](https://arxiv.org/abs/1902.00751)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1),[§3\.3](https://arxiv.org/html/2608.03275#S3.SS3.p4.5)\.
- Q\. Jin, B\. Dhingra, Z\. Liu, W\. W\. Cohen, and X\. Lu \(2019\)PubMedQA: a dataset for biomedical research question answering\.External Links:1909\.06146,[Link](https://arxiv.org/abs/1909.06146)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p2.1)\.
- A\. Kornilova and V\. Eidelman \(2019\)BillSum: a corpus for automatic summarization of US legislation\.InProceedings of the 2nd Workshop on New Frontiers in Summarization,L\. Wang, J\. C\. K\. Cheung, G\. Carenini, and F\. Liu \(Eds\.\),Hong Kong, China,pp\. 48–56\.External Links:[Link](https://aclanthology.org/D19-5406/),[Document](https://dx.doi.org/10.18653/v1/D19-5406)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p2.1)\.
- A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo, Y\. Wu, B\. Neyshabur, G\. Gur\-Ari, and V\. Misra \(2022\)Solving quantitative reasoning problems with language models\.External Links:2206\.14858,[Link](https://arxiv.org/abs/2206.14858)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- X\. L\. Li and P\. Liang \(2021\)Prefix\-tuning: optimizing continuous prompts for generation\.External Links:2101\.00190,[Link](https://arxiv.org/abs/2101.00190)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- Y\. Li, Y\. Yu, C\. Liang, P\. He, N\. Karampatziakis, W\. Chen, and T\. Zhao \(2023\)LoftQ: lora\-fine\-tuning\-aware quantization for large language models\.External Links:2310\.08659,[Link](https://arxiv.org/abs/2310.08659)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- V\. Lialin, V\. Deshpande, X\. Yao, and A\. Rumshisky \(2024\)Scaling down to scale up: a guide to parameter\-efficient fine\-tuning\.External Links:2303\.15647,[Link](https://arxiv.org/abs/2303.15647)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- C\. Lin \(2004\)ROUGE: a package for automatic evaluation of summaries\.InText Summarization Branches Out,Barcelona, Spain,pp\. 74–81\.External Links:[Link](https://aclanthology.org/W04-1013/)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p2.1)\.
- Q\. Liu, X\. Wu, X\. Zhao, Y\. Zhu, D\. Xu, F\. Tian, and Y\. Zheng \(2024a\)When moe meets llms: parameter efficient fine\-tuning for multi\-task medical applications\.External Links:2310\.18339,[Link](https://arxiv.org/abs/2310.18339)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.03275#S2.SS2.p1.1)\.
- S\. Liu, C\. Wang, H\. Yin, P\. Molchanov, Y\. F\. Wang, K\. Cheng, and M\. Chen \(2024b\)DoRA: weight\-decomposed low\-rank adaptation\.External Links:2402\.09353,[Link](https://arxiv.org/abs/2402.09353)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- I\. Loshchilov and F\. Hutter \(2017\)Decoupled weight decay regularization\.arXiv preprint arXiv:1711\.05101\.Cited by:[Appendix A](https://arxiv.org/html/2608.03275#A1.p1.1),[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px2.p1.7)\.
- L\. Lu, Z\. Wang, R\. Bao, M\. Wang, F\. Li, Y\. Wu, W\. Jiang, J\. Xu, Y\. Wang, and S\. Gao \(2024\)All\-in\-one tuning and structural pruning for domain\-specific llms\.External Links:2412\.14426,[Link](https://arxiv.org/abs/2412.14426)Cited by:[§4\.3](https://arxiv.org/html/2608.03275#S4.SS3.p1.1)\.
- T\. Mihaylov, P\. Clark, T\. Khot, and A\. Sabharwal \(2018\)Can a suit of armor conduct electricity? a new dataset for open book question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 2381–2391\.External Links:[Link](https://aclanthology.org/D18-1260/),[Document](https://dx.doi.org/10.18653/v1/D18-1260)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- T\. OLMo, P\. Walsh, L\. Soldaini, D\. Groeneveld, K\. Lo, S\. Arora, A\. Bhagia, Y\. Gu, S\. Huang, M\. Jordan, N\. Lambert, D\. Schwenk, O\. Tafjord, T\. Anderson, D\. Atkinson, F\. Brahman, C\. Clark, P\. Dasigi, N\. Dziri, A\. Ettinger, M\. Guerquin, D\. Heineman, H\. Ivison, P\. W\. Koh, J\. Liu, S\. Malik, W\. Merrill, L\. J\. V\. Miranda, J\. Morrison, T\. Murray, C\. Nam, J\. Poznanski, V\. Pyatkin, A\. Rangapur, M\. Schmitz, S\. Skjonsberg, D\. Wadden, C\. Wilhelm, M\. Wilson, L\. Zettlemoyer, A\. Farhadi, N\. A\. Smith, and H\. Hajishirzi \(2025\)2 olmo 2 furious\.External Links:2501\.00656,[Link](https://arxiv.org/abs/2501.00656)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- OpenAI, J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat, R\. Avila, I\. Babuschkin, S\. Balaji, V\. Balcom, P\. Baltescu, H\. Bao, M\. Bavarian, J\. Belgum, I\. Bello, J\. Berdine, G\. Bernadett\-Shapiro, C\. Berner, L\. Bogdonoff, O\. Boiko, M\. Boyd, A\. Brakman, G\. Brockman, T\. Brooks, M\. Brundage, K\. Button, T\. Cai, R\. Campbell, A\. Cann, B\. Carey, C\. Carlson, R\. Carmichael, B\. Chan, C\. Chang, F\. Chantzis, D\. Chen, S\. Chen, R\. Chen, J\. Chen, M\. Chen, B\. Chess, C\. Cho, C\. Chu, H\. W\. Chung, D\. Cummings, J\. Currier, Y\. Dai, C\. Decareaux, T\. Degry, N\. Deutsch, D\. Deville, A\. Dhar, D\. Dohan, S\. Dowling, S\. Dunning, A\. Ecoffet, A\. Eleti, T\. Eloundou, D\. Farhi, L\. Fedus, N\. Felix, S\. P\. Fishman, J\. Forte, I\. Fulford, L\. Gao, E\. Georges, C\. Gibson, V\. Goel, T\. Gogineni, G\. Goh, R\. Gontijo\-Lopes, J\. Gordon, M\. Grafstein, S\. Gray, R\. Greene, J\. Gross, S\. S\. Gu, Y\. Guo, C\. Hallacy, J\. Han, J\. Harris, Y\. He, M\. Heaton, J\. Heidecke, C\. Hesse, A\. Hickey, W\. Hickey, P\. Hoeschele, B\. Houghton, K\. Hsu, S\. Hu, X\. Hu, J\. Huizinga, S\. Jain, S\. Jain, J\. Jang, A\. Jiang, R\. Jiang, H\. Jin, D\. Jin, S\. Jomoto, B\. Jonn, H\. Jun, T\. Kaftan, Ł\. Kaiser, A\. Kamali, I\. Kanitscheider, N\. S\. Keskar, T\. Khan, L\. Kilpatrick, J\. W\. Kim, C\. Kim, Y\. Kim, J\. H\. Kirchner, J\. Kiros, M\. Knight, D\. Kokotajlo, Ł\. Kondraciuk, A\. Kondrich, A\. Konstantinidis, K\. Kosic, G\. Krueger, V\. Kuo, M\. Lampe, I\. Lan, T\. Lee, J\. Leike, J\. Leung, D\. Levy, C\. M\. Li, R\. Lim, M\. Lin, S\. Lin, M\. Litwin, T\. Lopez, R\. Lowe, P\. Lue, A\. Makanju, K\. Malfacini, S\. Manning, T\. Markov, Y\. Markovski, B\. Martin, K\. Mayer, A\. Mayne, B\. McGrew, S\. M\. McKinney, C\. McLeavey, P\. McMillan, J\. McNeil, D\. Medina, A\. Mehta, J\. Menick, L\. Metz, A\. Mishchenko, P\. Mishkin, V\. Monaco, E\. Morikawa, D\. Mossing, T\. Mu, M\. Murati, O\. Murk, D\. Mély, A\. Nair, R\. Nakano, R\. Nayak, A\. Neelakantan, R\. Ngo, H\. Noh, L\. Ouyang, C\. O’Keefe, J\. Pachocki, A\. Paino, J\. Palermo, A\. Pantuliano, G\. Parascandolo, J\. Parish, E\. Parparita, A\. Passos, M\. Pavlov, A\. Peng, A\. Perelman, F\. de Avila Belbute Peres, M\. Petrov, H\. P\. de Oliveira Pinto, Michael, Pokorny, M\. Pokrass, V\. H\. Pong, T\. Powell, A\. Power, B\. Power, E\. Proehl, R\. Puri, A\. Radford, J\. Rae, A\. Ramesh, C\. Raymond, F\. Real, K\. Rimbach, C\. Ross, B\. Rotsted, H\. Roussez, N\. Ryder, M\. Saltarelli, T\. Sanders, S\. Santurkar, G\. Sastry, H\. Schmidt, D\. Schnurr, J\. Schulman, D\. Selsam, K\. Sheppard, T\. Sherbakov, J\. Shieh, S\. Shoker, P\. Shyam, S\. Sidor, E\. Sigler, M\. Simens, J\. Sitkin, K\. Slama, I\. Sohl, B\. Sokolowsky, Y\. Song, N\. Staudacher, F\. P\. Such, N\. Summers, I\. Sutskever, J\. Tang, N\. Tezak, M\. B\. Thompson, P\. Tillet, A\. Tootoonchian, E\. Tseng, P\. Tuggle, N\. Turley, J\. Tworek, J\. F\. C\. Uribe, A\. Vallone, A\. Vijayvergiya, C\. Voss, C\. Wainwright, J\. J\. Wang, A\. Wang, B\. Wang, J\. Ward, J\. Wei, C\. Weinmann, A\. Welihinda, P\. Welinder, J\. Weng, L\. Weng, M\. Wiethoff, D\. Willner, C\. Winter, S\. Wolrich, H\. Wong, L\. Workman, S\. Wu, J\. Wu, M\. Wu, K\. Xiao, T\. Xu, S\. Yoo, K\. Yu, Q\. Yuan, W\. Zaremba, R\. Zellers, C\. Zhang, M\. Zhang, S\. Zhao, T\. Zheng, J\. Zhuang, W\. Zhuk, and B\. Zoph \(2024\)GPT\-4 technical report\.External Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- J\. Pfeiffer, A\. Kamath, A\. Rücklé, K\. Cho, and I\. Gurevych \(2021\)AdapterFusion: non\-destructive task composition for transfer learning\.External Links:2005\.00247,[Link](https://arxiv.org/abs/2005.00247)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- A\. Romanov and C\. Shivade \(2018\)Lessons from natural language inference in the clinical domain\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 1586–1596\.External Links:[Link](https://aclanthology.org/D18-1187/),[Document](https://dx.doi.org/10.18653/v1/D18-1187)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p2.1)\.
- B\. Rozière, J\. Gehring, F\. Gloeckle, S\. Sootla, I\. Gat, X\. E\. Tan, Y\. Adi, J\. Liu, R\. Sauvestre, T\. Remez, J\. Rapin, A\. Kozhevnikov, I\. Evtimov, J\. Bitton, M\. Bhatt, C\. C\. Ferrer, A\. Grattafiori, W\. Xiong, A\. Défossez, J\. Copet, F\. Azhar, H\. Touvron, L\. Martin, N\. Usunier, T\. Scialom, and G\. Synnaeve \(2024\)Code llama: open foundation models for code\.External Links:2308\.12950,[Link](https://arxiv.org/abs/2308.12950)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- K\. Sakaguchi, R\. L\. Bras, C\. Bhagavatula, and Y\. Choi \(2019\)WinoGrande: an adversarial winograd schema challenge at scale\.External Links:1907\.10641,[Link](https://arxiv.org/abs/1907.10641)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- M\. Sap, H\. Rashkin, D\. Chen, R\. Le Bras, and Y\. Choi \(2019\)Social IQa: commonsense reasoning about social interactions\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 4463–4473\.External Links:[Link](https://aclanthology.org/D19-1454/),[Document](https://dx.doi.org/10.18653/v1/D19-1454)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- N\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. Le, G\. Hinton, and J\. Dean \(2017\)Outrageously large neural networks: the sparsely\-gated mixture\-of\-experts layer\.External Links:1701\.06538,[Link](https://arxiv.org/abs/1701.06538)Cited by:[§2\.2](https://arxiv.org/html/2608.03275#S2.SS2.p1.1)\.
- C\. Tian, Z\. Shi, Z\. Guo, L\. Li, and C\. Xu \(2024\)HydraLoRA: an asymmetric lora architecture for efficient fine\-tuning\.External Links:2404\.19245,[Link](https://arxiv.org/abs/2404.19245)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.03275#S2.SS2.p1.1)\.
- H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar, A\. Rodriguez, A\. Joulin, E\. Grave, and G\. Lample \(2023a\)LLaMA: open and efficient foundation language models\.External Links:2302\.13971,[Link](https://arxiv.org/abs/2302.13971)Cited by:[§1](https://arxiv.org/html/2608.03275#S1.p1.1)\.
- H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale, D\. Bikel, L\. Blecher, C\. C\. Ferrer, M\. Chen, G\. Cucurull, D\. Esiobu, J\. Fernandes, J\. Fu, W\. Fu, B\. Fuller, C\. Gao, V\. Goswami, N\. Goyal, A\. Hartshorn, S\. Hosseini, R\. Hou, H\. Inan, M\. Kardas, V\. Kerkez, M\. Khabsa, I\. Kloumann, A\. Korenev, P\. S\. Koura, M\. Lachaux, T\. Lavril, J\. Lee, D\. Liskovich, Y\. Lu, Y\. Mao, X\. Martinet, T\. Mihaylov, P\. Mishra, I\. Molybog, Y\. Nie, A\. Poulton, J\. Reizenstein, R\. Rungta, K\. Saladi, A\. Schelten, R\. Silva, E\. M\. Smith, R\. Subramanian, X\. E\. Tan, B\. Tang, R\. Taylor, A\. Williams, J\. X\. Kuan, P\. Xu, Z\. Yan, I\. Zarov, Y\. Zhang, A\. Fan, M\. Kambadur, S\. Narang, A\. Rodriguez, R\. Stojnic, S\. Edunov, and T\. Scialom \(2023b\)Llama 2: open foundation and fine\-tuned chat models\.External Links:2307\.09288,[Link](https://arxiv.org/abs/2307.09288)Cited by:[§4\.2](https://arxiv.org/html/2608.03275#S4.SS2.p1.9)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.2](https://arxiv.org/html/2608.03275#S4.SS2.p1.9)\.
- R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. Choi \(2019\)HellaSwag: can a machine really finish your sentence?\.External Links:1905\.07830,[Link](https://arxiv.org/abs/1905.07830)Cited by:[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px1.p1.1)\.
- Q\. Zhang, M\. Chen, A\. Bukharin, N\. Karampatziakis, P\. He, Y\. Cheng, W\. Chen, and T\. Zhao \(2023\)AdaLoRA: adaptive budget allocation for parameter\-efficient fine\-tuning\.External Links:2303\.10512,[Link](https://arxiv.org/abs/2303.10512)Cited by:[§2\.1](https://arxiv.org/html/2608.03275#S2.SS1.p1.1)\.
- Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin, F\. Huang, and J\. Zhou \(2025\)Qwen3 embedding: advancing text embedding and reranking through foundation models\.External Links:2506\.05176,[Link](https://arxiv.org/abs/2506.05176)Cited by:[§3\.2](https://arxiv.org/html/2608.03275#S3.SS2.p2.4),[§4\.1](https://arxiv.org/html/2608.03275#S4.SS1.SSS0.Px2.p1.7)\.

## Appendix AImplementation Details

![Refer to caption](https://arxiv.org/html/2608.03275v1/figure/qwen_sota_loss_curves.png)Figure 4:Training Loss Example\.Table 6:Ablation study on routing design choices\.We keep the backbone LLM frozen for all PEFT methods and update only the adapter\-related parameters\. Unless otherwise specified, all methods are trained for 3 epochs with AdamWLoshchilov and Hutter \([2017](https://arxiv.org/html/2608.03275#bib.bib7)\), a learning rate of2×10−42\\times 10^\{\-4\}, cosine learning\-rate decay, and the Alpaca prompt format\.

For static PEFT baselines, including LoRA, DoRA, and SMT, we follow the experimental setup of SMT for training\-related hyperparameters, including the learning rate, number of epochs, warmup steps, scheduler, and prompt format\. We keep each method’s adaptation modules consistent with the SMT setting whenever applicable: LoRA and DoRA are applied to the Q/K/V/O attention projections and the Gate/Up/Down feed\-forward projections, while SMT is applied to the Q/K/V attention projections\.

For the MoE\-LoRA and MoLA baselines, we generally follow the hyperparameter settings used in their original papers and keep the implementation consistent with their routing designs\. Specifically, MoE\-LoRA usesN=4N\{=\}4LoRA experts with rankr=32r\{=\}32and a task\-conditioned soft mixture, while MoLA usesN=8N\{=\}8LoRA experts with rankr=32r\{=\}32and token\-level top\-kkrouting withk=2k\{=\}2\. Both baselines are applied to the Q/K/V/O attention projections and the Gate/Up/Down feed\-forward projections\.

## Appendix BAblation Study on Routing Design Choices

We ablate two routing design choices in MoEGen: the number of activated experts and the use of an always\-on shared expert\. All variants are evaluated on the cross\-domain joint benchmark with LLaMA\-2\-13B as the backbone, following the same training setup as Section[4\.3](https://arxiv.org/html/2608.03275#S4.SS3)\. They use LoRA rankr=64r\{=\}64, scaling factorα=128\\alpha\{=\}128, expert dimension3232, and33epochs of supervised fine\-tuning\. The variants differ only in the routing strategy\.

Table 7:Qualitative example on BillSum \(Heavy Truck Tax Fairness Act of 2010\)\.Table 8:Additional BillSum qualitative examples illustrating two distinct failure modes of the static LoRA\-based baselines: a repetitive\-loop failure on the Comprehensive Immunosuppressive Drug Coverage Act \(top\) and a mid\-clause\-boilerplate failure on the Federal Election Law Reform Commission Act \(bottom\)\. All methods are trained jointly on the cross\-domain mixture \(MedNLI, PubMedQA, HQS, BillSum\) with Llama\-2\-13B\.MethodR\-LPredicted summaryExample 1: Comprehensive Immunosuppressive Drug Coverage Act \(repetitive\-loop failure\)\.Reference–Comprehensive Immunosuppressive Drug Coverage for Kidney Transplant Patients Act of 2016\. This bill amends titles II and XVIII \(Medicare\) of the Social Security Act to indefinitely extend Medicare coverage of immunosuppressive drugs for kidney transplant recipients\. Under current law, such coverage is limited to 36 months following a transplant\.LoRA0\.182\(c\) of this section shall apply to such an individual in the same manner as they apply to an individual who is entitled to benefits under part A\.MoE\-LoRA0\.047\(c\) of this section shall not apply to such an individual\. The provisions of subsection \(d\) of this section shall not apply to such an individual…The provisions of subsection \(e\) of this section shall not apply to such an individual…\(the pattern repeats for subsections \(f\), \(g\), …\)MoLA0\.069\(c\) shall apply to such individual’s coverage period in the same manner as they apply to the coverage period of an individual who is entitled to benefits under part A\. The provisions of subsection \(d\) shall apply to such individual’s coverage period…The provisions of subsection \(e\) shall apply…\(the pattern repeats for subsections \(f\), \(g\), …\)MoEGen \(Ours\)0\.571Comprehensive Immunosuppressive Drug Coverage for Kidney Transplant Patients Act of 2016\. This bill amends title XVIII \(Medicare\) of the Social Security Act to require the Medicare program to provide immunosuppressive drug coverage for kidney transplant recipients\. The bill also requires the program to provide immunosuppressive drug coverage for individuals whose entitlement to Medicare Part A hospital insurance benefits has ended by reason of the 36\-month limit on such benefits\.Example 2: Federal Election Law Reform Commission Act \(mid\-clause\-boilerplate failure\)\.Reference–Federal Election Law Reform Commission Act \- Establishes the Federal Election Law Reform Commission to recommend Federal election law reform\. Requires a Federal election bill to be introduced in the House and the Senate within a specified time of the Commission’s submittal of draft legislation\. Authorizes appropriations\.LoRA0\.135…be considered as part of the rules of each House, respectively, and such rules shall supersede other rules only to the extent that they are inconsistent therewith\.MoE\-LoRA0\.135…be considered as part of the rules of each House, respectively, and such rules shall supersede other rules only to the extent that they are inconsistent therewith\.MoLA0\.139…be considered a part of the rules of each House, respectively, and shall supersede other rules only to the extent that they are inconsistent therewith\.MoEGen \(Ours\)0\.522Federal Election Law Reform Commission Act \- Establishes the Federal Election Law Reform Commission to study and recommend reforms in Federal election laws\.Table[6](https://arxiv.org/html/2608.03275#A1.T6)shows that top\-2 routing achieves the best overall performance, improving the average score from0\.58440\.5844to0\.58770\.5877over top\-1 routing\. The gains on MedNLI, BillSum, and PubMedQA suggest that combining two routed expert codes provides a more flexible adaptation space than selecting a single expert\. HQS is the only exception, where top\-1 routing gives a slightly higher ROUGE\-L score, possibly because its summarization format is more constrained and benefits from stronger routing sparsity\.

Adding an always\-on shared expert does not improve performance\. Compared with standard top\-2 routing, the shared\-expert variant reduces the average score from0\.58770\.5877to0\.58270\.5827, with drops on BillSum, HQS, and PubMedQA\. This suggests that, under a fixed parameter budget, the shared expert may absorb part of the learning signal and weaken the specialization of routed experts\. Therefore, we use top\-2 routing without a shared expert as the default configuration\.

## Appendix CAdditional Qualitative Study

We provide two additional BillSum predictions in Table[8](https://arxiv.org/html/2608.03275#A2.T8), drawn from the same joint cross\-domain training run as Table[7](https://arxiv.org/html/2608.03275#A2.T7)\. The two examples illustrate distinct failure modes of the static LoRA\-based baselines on legal long\-document inputs: collapse into a repetitive “subsection X does not apply” loop in Example 1, and emission of mid\-clause rules\-of\-the\-House boilerplate in Example 2\. In both cases MoEGen opens with the act name followed by “Amends … to…”—the canonical BillSum register—and enumerates the substantive actions present in the reference\. None of these behaviors require extra parameters: MoEGen uses fewer trainable parameters than LoRA and MoLA and roughly the same as MoE\-LoRA in Table[2](https://arxiv.org/html/2608.03275#S4.T2), so the difference is attributable to the input\-conditioned parameterization rather than to capacity\. In particular, MoLA uses more than twice as many trainable parameters as MoEGen yet exhibits the same qualitative failures as the other two static baselines on these inputs\.

We further evaluate the inference cost of MoEGen on the cross\-domain joint benchmark\. Table[5](https://arxiv.org/html/2608.03275#S4.T5)reports the per\-sample latency of each method using LLaMA\-2\-13B on a single B200 GPU with batch size 32\. MoEGen introduces an input\-dependent routing and generation cost, but the overhead remains small on longer\-generation tasks\. The largest relative gap appears on MedNLI, where the output is very short and the fixed routing cost is less amortized\. As the generation length increases, this gap becomes much smaller: MoEGen is only 9\.1% slower than LoRA on HQS, 8\.1% slower on PubMedQA, and 4\.1% slower on BillSum\. It also remains substantially faster than MoLA on the longer\-generation tasks\.

Similar Articles

MobileMoE: Scaling On-Device Mixture of Experts

Hugging Face Daily Papers

MobileMoE introduces efficient on-device mixture-of-experts language models with sub-billion parameters, achieving better performance and efficiency than dense baselines and existing MoE models. The models are trained on open-source datasets and demonstrate significant speedups on commodity smartphones.