MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation
Summary
MoE2-LoRA introduces a dual-channel Routing-Conditioned Projection and a global LoRA expert pool to enable MoE-style low-rank adaptation for fine-tuning MoE models, achieving state-of-the-art accuracy while retaining general capabilities.
View Cached Full Text
Cached at: 07/27/26, 07:39 AM
# MoE2-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation
Source: [https://arxiv.org/html/2607.21978](https://arxiv.org/html/2607.21978)
Qingyu Yang1,2,∗Haonan He1,3,∗,†Minglei Li1,4Jingqi Ye1,3 Tao Chen1Lei Bai1Peng Ye1,4,5,‡ 1Shanghai Artificial Intelligence Laboratory2KTH Royal Institute of Technology 3University of Science and Technology of China4Fudan University 5The Chinese University of Hong Kong \* Equal contribution\.†\\daggerProject Lead\.‡\\ddaggerCorresponding author\. Email:yepeng@pjlab\.org\.cn
###### Abstract
Mixture\-of\-Experts \(MoE\) architectures have been widely adopted in large language models, yet parameter\-efficient fine\-tuning \(PEFT\) for MoE models remains underexplored\. Existing PEFT methods for MoE either ignore router priors with uniform adapters, reducing efficiency and risking forgetting, or rely on static expert selection, limiting per\-token capacity and cross\-expert feature learning\. In this paper, we make the first attempt to fine\-tune MoE models with MoE\-style low\-rank adaptation: our method, entitled MoE2\-LoRA, deeply couples the pretrained expert specialization with task\-specific adaptivity via a dual\-channel Routing\-Conditioned Projection \(RCP\) module, which reuses base router activations to inform LoRA routing\. We further introduce a single global LoRA expert pool shared across all layers, enabling model\-wide adaptation with emergent layer\-wise affinities and balanced expert utilization\. MoE2\-LoRA simultaneously benefits from the advantages of prior reuse, dynamic adapter routing, and model\-wide knowledge sharing\. Evaluated on multiple MoE backbones with varying scales and expert granularities, MoE2\-LoRA consistently achieves state\-of\-the\-art downstream accuracy while retaining stronger general capabilities\.
MoE2\-LoRA: When MoE Models Meet MoE\-style Low\-Rank Adaptation
Qingyu Yang1,2,∗Haonan He1,3,∗,†Minglei Li1,4Jingqi Ye1,3Tao Chen1Lei Bai1Peng Ye1,4,5,‡1Shanghai Artificial Intelligence Laboratory2KTH Royal Institute of Technology3University of Science and Technology of China4Fudan University5The Chinese University of Hong Kong\* Equal contribution\.†\\daggerProject Lead\.‡\\ddaggerCorresponding author\.Email:yepeng@pjlab\.org\.cn
## 1Introduction
Mixture\-of\-Experts \(MoE\)\(Jacobs et al\.,[1991](https://arxiv.org/html/2607.21978#bib.bib17)\)architectures, which activate only a small fraction of parameters for each input, have recently been widely adopted by large language models \(LLMs\) such as DeepSeek\-V3\(Liu et al\.,[2024a](https://arxiv.org/html/2607.21978#bib.bib21)\)and Qwen3Yang et al\. \([2025](https://arxiv.org/html/2607.21978#bib.bib35)\), owing to their advantages including reasoning efficiency and scaling performance comparable to that of dense models\. For dense architecture based LLMs, efficient and effective fine\-tuning methods have been thoroughly studied\. For example, LoRA\(Hu et al\.,[2022](https://arxiv.org/html/2607.21978#bib.bib16)\), one of the most widely used parameter\-efficient fine\-tuning \(PEFT\) methods designed for dense models, matches the performance of full fine\-tuning by decomposing weight updates into low\-rank subspaces\. In comparison, how to efficiently fine\-tune MoE models remains an underexplored problem, primarily due to their sparse activation nature, unstable gradient flow, and massive parameter scale, which highlights the pressing need for PEFT methods specifically tailored for MoE models\.
Existing PEFT methods for MoE models are predominantly built upon applying LoRA adapters to MoE experts \(Fig\.[1](https://arxiv.org/html/2607.21978#S1.F1)\(I\.a\)\), and further attempt to exploit the routing priors of the MoE model to guide expert\-level adaptation, that is, to decide which experts should receive adaptation or how much \(Fig\.[1](https://arxiv.org/html/2607.21978#S1.F1)\(I\.b\)\)\. For example, ESFT\(Wang et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib32)\)and DAS\-LoRA\(Tang et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib31)\)first collect offline router statistics to identify a subset of critical experts, and then apply LoRA only to those experts; DR\-LoRA\(Deng et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib7)\)adaptively assigns higher ranks to experts with higher routing frequency and gradient\-based importance\. While these methods improve parameter efficiency, they still suffer from two major limitations: \(i\) Insufficient adaptation of critical experts\. By selecting a subset of trainable experts or high\-rank experts, any expert who falls outside this subset but is activated at inference time receives insufficient task\-specific updates\. \(ii\) Lack of high\-level interactions: Adapting experts independently neglects both the cooperative dynamics among co\-activated experts and the cross\-layer interactions of MoE models \(Fig\.[5](https://arxiv.org/html/2607.21978#S5.F5)\), which could enhance performance and reduce parameter redundancy\.
Figure 1:\(I\) Comparison of PEFT\-on\-MoE methods: \(a\) Per\-expert LoRA, which equip each MoE expert with a LoRA adapter; \(b\) Static expert\-subset selection methods, which select a subset of experts offline to be trained with LoRA; \(c\) MoELoRA\-style methods, which adapt each MoE module with a MoE\-style low\-rank adapter; \(d\) MoE2\. \(II\) Annotation\. \(III\) LoRA Experts Preference: We present the preferred layers of LoRA experts in MoE2’s global expert pool, demonstrating the layer\-specific learning and cross\-layer learning capabilities of MoE2\.To address these limitations, we seek to fine\-tune MoE models with dedicated MoE mechanisms, shifting the adaptation horizon from individual experts to each MoE module and the entire MoE model\. This direction resonates with another line of research that integrates LoRA with an MoE structure \(Fig\.[1](https://arxiv.org/html/2607.21978#S1.F1)\(I\.c\)\), which yields MoE\-style LoRA methods such as MoELoRA\(Luo et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib25)\)and MoLA\(Gao et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib10)\)\. These methods, however, are not specifically designed for MoE models: their MoE mechanisms operate independently of those inherent to MoE models, thereby leaving the expert interactions and cross\-layer interactions encoded in the base MoE models unutilized\. In light of this, if it is possible to deeply bind the two MoE mechanisms of MoE models and MoE\-style LoRA, we could mitigate the aforementioned limitations\. These observations naturally raise a unifying question:When MoE models meet MoE\-style low\-rank adaptation, how can we deeply bind the two MoE mechanisms at both the MoE module level and the MoE model level, thereby enhancing task performance while preserving general capabilities?
To answer this question, we introduce MoE2\-LoRA \(Figure[1](https://arxiv.org/html/2607.21978#S1.F1)\(I\.d\)\), an MoE\-style low\-rank adaptation method tailored for MoE models\. At its core lies the Routing\-Conditioned Projection \(RCP\) module, which reuses the pretrained MoE routing activations to drive LoRA expert selection\. Concretely, RCP fuses two routing signals: a primary channelzzcarrying the base router activations, along with an auxiliary channelhch\_\{c\}, which injects a task\-specific correction\. A learnable matrixWlW\_\{l\}maps the concatenated input into the LoRA routing space, producing LoRA expert selection coupled with the base model’s expert selection\. Further, in order to enable model\-level knowledge sharing, we maintain aglobal poolof LoRA experts shared across all MoE layers\. As demonstrated in Figure[1](https://arxiv.org/html/2607.21978#S1.F1)\(III\), this design yields a striking emergent property: each LoRA expert develops a pronounced layer affinity, while still contributing to other layers to a lesser extent, and the overall assignment of experts across layers remains highly uniform\. This indicates that the global pool achieves model\-wide cross\-layer knowledge sharing with high parameter utilization, avoiding the isolation and redundancy of layer\-specific adapters\.
As a result, our MoE2\-LoRA unifies three key properties: pretrained routing prior reuse, MoE\-style LoRA expert selection, and model\-level cross\-layer learning through a global expert pool\. Equipped with this model\-level design, MoE2\-LoRA attains state\-of\-the\-art downstream accuracy on four MoE backbones of varying scale and expert granularity, OLMoE\-1B\-7BMuennighoff et al\. \([2024](https://arxiv.org/html/2607.21978#bib.bib27)\), DeepSeek\-V2\-LiteLiu et al\. \([2024a](https://arxiv.org/html/2607.21978#bib.bib21)\), Qwen3\-30B\-A3BYang et al\. \([2025](https://arxiv.org/html/2607.21978#bib.bib35)\), and Qwen3\.5\-35B\-A3BQwen Team \([2026](https://arxiv.org/html/2607.21978#bib.bib28)\), improving over the strongest PEFT baselines by up to \+2\.56 points in in\-domain average accuracy, while exhibiting superior general capability retention\. Our contributions are summarized as follows:
- •We diagnose two key limitations of existing PEFT methods for MoE models and introduce the principle of fine\-tuning MoE models with MoE\-style PEFT, which requires deeply binding the two MoE mechanisms and establishing both module level and model level interactions\. This principle unifies expert\-level, module\-level, and model\-level adaptation, overcoming the identified limitations and serving as the foundation for our method design\.
- •We propose MoE2\-LoRA, which realizes this insight through two key components: \(1\) a Routing\-Conditioned Projection \(RCP\) module that repurposes the base router’s activations to drive LoRA expert selection, and \(2\) a single global LoRA expert pool shared across all MoE layers\. As visualized in Figure[1](https://arxiv.org/html/2607.21978#S1.F1)\(I\), the global pool yields emergent layer affinities and balanced expert utilization, achieving efficient model\-wide knowledge sharing\.
- •Across four MoE backbones \(OLMoE\-1B\-7B, DeepSeek\-V2\-Lite, Qwen3\-30B\-A3B, Qwen3\.5\-35B\-A3B\) and various benchmarks \(math, code, general, multimodal\), MoE2\-LoRA achieves state\-of\-the\-art downstream performance and general capability retention, improving over the base model in both task and general performance\.
## 2Related Work
### 2\.1PEFT and LoRA\.
The scaling of LLMs to billions of parameters has rendered conventional full fine\-tuning computationally prohibitive\. This bottleneck has spurred the widespread adoption of PEFT techniques, epitomized by LoRA, owing to its plug\-and\-play deployability, strong downstream performance, and zero inference latency\. By restricting updates to a small subset of trainable parameters, these methods enable resource\-efficient model adaptation while preserving the model’s original inference cost\. Though LoRA matches the performance of full fine\-tuning in many scenarios, a gap persists on challenging tasks such as scientific reasoning and coding\. This gap has been extensively investigated and effectively narrowed by a recent body of follow\-up work that refines LoRA with improved algorithms\. For example, LoRA\+\(Hayou et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib12)\)decouples the learning rates for the down\- and up\-projection weights of a low\-rank adapter, effectively enhancing training stability over standard LoRA\. GoRA\(He et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib13)\)adaptively assigns ranks to adapters based on gradient information, thereby strengthening parameter utilization\. Similarly, methods such as MoELoRA\(Luo et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib25)\), MoLA\(Gao et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib10)\), and LoRAMoE\(Dou et al\.,[2023](https://arxiv.org/html/2607.21978#bib.bib8)\)improve LoRA’s performance by integrating MoE mechanisms\.
### 2\.2PEFT for MoE models\.
PEFT methods are predominantly implemented and evaluated on dense models, without being tailored to MoE architectures\. Existing studies on PEFT methods for MoE models remain sparse and insufficiently comparative\. Though PERFTLiu et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib23)\)has briefly investigated various strategies for training MoE models with LoRA, most PEFT methods for MoE models concentrate on selecting experts to be fine\-tuned or adaptively assigning ranks to the adapters of MoE experts based on routing statistics and expert importance information\. For example, ESFT\(Wang et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib32)\)analyzes the router’s output probability distribution on downstream task data and accordingly selects a subset of experts to be fine\-tuned; MoE\-SieveManzoni \([2026](https://arxiv.org/html/2607.21978#bib.bib26)\)counts how many tokens each expert receives per layer on a small calibration dataset, then applies LoRA adapters exclusively to the selected experts; CEFT\(Bai et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib2)\)identifies experts that progressively amplify attention to relevant contextual information and selectively fine\-tunes only these experts; DR\-LoRADeng et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib7)\)assigns heterogeneous LoRA ranks across experts based on an expert saliency score that jointly considers routing frequency and gradient\-based rank importance\.
## 3Method
### 3\.1MoE Routing and LoRA
An MoE layer contains a routerG\(⋅\)G\(\\cdot\)andNEN\_\{E\}feed\-forward experts\{Ei\}i=1NE\\\{E\_\{i\}\\\}\_\{i=1\}^\{N\_\{E\}\}\. Given a token representationh∈ℝDh\\in\\mathbb\{R\}^\{D\}, the router produces logitsz=G\(h\)z=G\(h\)and activates a top\-KKexpert set𝒮h\\mathcal\{S\}\_\{h\}\. The MoE output is
MoE\(h\)=∑i∈𝒮hgi\(h\)Ei\(h\),\\mathrm\{MoE\}\(h\)=\\sum\_\{i\\in\\mathcal\{S\}\_\{h\}\}g\_\{i\}\(h\)E\_\{i\}\(h\),wheregi\(h\)g\_\{i\}\(h\)is the normalized routing weight\.
LoRA parameterizes a low\-rank update to a pretrained weightW∈ℝdout×dinW\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}as
ΔW=αrBA,\\Delta W=\\frac\{\\alpha\}\{r\}BA,whereA∈ℝr×dinA\\in\\mathbb\{R\}^\{r\\times d\_\{\\mathrm\{in\}\}\},B∈ℝdout×rB\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times r\}, andr≪min\(dout,din\)r\\ll\\min\(d\_\{\\mathrm\{out\}\},d\_\{\\mathrm\{in\}\}\)\. In PEFT\-on\-MoE, the pretrained router and experts are frozen, and only the introduced adapter parameters are optimized\.
### 3\.2Routing\-Conditioned Projection \(WlW\_\{l\}\)
MoE2\-LoRA introduces an additional LoRA expert pool\{EjLoRA\}j=1NL\\\{E\_\{j\}^\{\\mathrm\{LoRA\}\}\\\}\_\{j=1\}^\{N\_\{L\}\}and learns per\-token routingpL∈ℝNLp\_\{L\}\\in\\mathbb\{R\}^\{N\_\{L\}\}over this pool\. The key design choice is the*routing input*: rather than learning routing from raw hidden statehhor hardwiring it to base routing, MoE2\-LoRA uses a*two\-channel input*: \(1\) Primary channel: the base router logitsz=G\(h\)∈ℝNEz=G\(h\)\\in\\mathbb\{R\}^\{N\_\{E\}\}\. This provides a directly pretrained routing signal for LoRA expert selection instead of learning routing entirely from scratch; \(2\) Auxiliary channel: a low\-rank projection of the hidden state,hc=Wh⋅h∈ℝdbh\_\{c\}=W\_\{h\}\\cdot h\\in\\mathbb\{R\}^\{d\_\{b\}\}withdb≪Dd\_\{b\}\\ll D\. This adds a learn from scratch signal that complements the base router channel\. The LoRA routing is computed as:
pL=softmax\(Wl⋅\[z;hc\]\),p\_\{L\}\\;=\\;\\mathrm\{softmax\}\\bigl\(W\_\{l\}\\cdot\[\\,z\\,;\\,h\_\{c\}\\,\]\\bigr\),\(1\)whereWl∈ℝNL×\(NE\+db\)W\_\{l\}\\in\\mathbb\{R\}^\{N\_\{L\}\\times\(N\_\{E\}\+d\_\{b\}\)\}is a learnable per\-layer projection matrix\. We then select the top\-KLK\_\{L\}LoRA experts,𝒮hLoRA=TopK\(pL\)\\mathcal\{S\}\_\{h\}^\{\\mathrm\{LoRA\}\}=\\mathrm\{TopK\}\(p\_\{L\}\), and renormalize their weights to sum to one:p~L,j=pL,j/∑j′∈𝒮hLoRApL,j′\\tilde\{p\}\_\{L,j\}=p\_\{L,j\}/\\sum\_\{j^\{\\prime\}\\in\\mathcal\{S\}\_\{h\}^\{\\mathrm\{LoRA\}\}\}p\_\{L,j^\{\\prime\}\}\. Fig\.[2](https://arxiv.org/html/2607.21978#S3.F2)illustrates the complete forward pass\.
Figure 2:Overview of the MoE2\-LoRA forward pass on a single MoE layer\.This design conditions LoRA routing on pretrained routing signals while preserving the flexibility of task\-specific routing adaptation through the learnable projectionWlW\_\{l\}\.
### 3\.3Global Shared LoRA Pool
MoE2\-LoRA’s LoRA expert pool is*shared across layers*: a set ofNLN\_\{L\}LoRA experts is reused by multiple MoE blocks, with each block accessing the pool through its own per\-layerWlW\_\{l\}\.
This contrasts with heuristic allocation\-based designs such as MoLA, which modifies layer heterogeneity by*statically pre\-specifying*each layer’s expert pool size and composition, an architectural choice that requires search over per\-layer capacity schemes\. MoE2\-LoRA instead handles layer heterogeneity*dynamically at the routing level*: each layer can flexibly select from the shared pool through its learnableWlW\_\{l\}, with no per\-layer capacity hyperparameter\. The same expert can serve multiple layers, and the same layer can route differently from its neighbors\.
An evident advantage of cross\-layer expert sharing is that the parameter count of the trainable LoRA experts pool do not scale linearly with model depth, in contrast to per\-layer pool designs where the parameter count grows asL×Nper\-layerL\\times N\_\{\\mathrm\{per\\text\{\-\}layer\}\}\. The choice of the total number of LoRA experts and the number of activated experts per input and per layer is clearly more flexible than that of the per\-layer pool designs\.
### 3\.4Forward Pass
Combining the RCP routing of §[3\.2](https://arxiv.org/html/2607.21978#S3.SS2)with the global shared pool of §[3\.3](https://arxiv.org/html/2607.21978#S3.SS3), the MoE2\-LoRA block output is
y=MoE\(h\)\+∑j∈𝒮hLoRAp~L,j⋅EjLoRA\(h\),y\\;=\\;\\mathrm\{MoE\}\(h\)\\;\+\\;\\sum\_\{j\\in\\mathcal\{S\}\_\{h\}^\{\\mathrm\{LoRA\}\}\}\\tilde\{p\}\_\{L,j\}\\cdot E\_\{j\}^\{\\mathrm\{LoRA\}\}\(h\),\(2\)whereEjLoRA\(h\)=\(α/r\)BjAjhE\_\{j\}^\{\\mathrm\{LoRA\}\}\(h\)=\(\\alpha/r\)\\,B\_\{j\}A\_\{j\}his the rank\-rrLoRA expert, andp~L,j\\tilde\{p\}\_\{L,j\}are the renormalized top\-KLK\_\{L\}routing weights defined in §[3\.2](https://arxiv.org/html/2607.21978#S3.SS2)\(or its hierarchical variant whenG\>1G\>1\)\.
## 4Experiments
In this section, we evaluate MoE2\-LoRA and baseline methods on multiple MoE backbones across various tasks \(math, code, general\-retention, and multimodal\)\.
### 4\.1Main Experiments
##### Base Models\.
We conduct the main experiments on three MoE backbones:OLMoE\-1B\-7B\(Muennighoff et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib27)\),DeepSeek\-V2\-Lite\(DeepSeek\-AI,[2024](https://arxiv.org/html/2607.21978#bib.bib6)\), andQwen3\-30B\-A3B\(Yang et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib35)\), covering different model scales and MoE architectures\. This wide set of models ensures a comprehensive evaluation across scales and model families\.
##### Baselines\.
We compare MoE2\-LoRA against five baselines covering the dominant families of PEFT\-for\-MoE: PERFT\-E\(Liu et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib23)\)\(per\-expert LoRA following the base routing\), MoELoRA\(Luo et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib25)\)\(per\-layer LoRA expert pool with an independent learned router\), MoLA\(Gao et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib10)\)\(heuristic expert allocation for MoE\-style LoRA\), DAS\-LoRA\(Tang et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib31)\)\(DAS\-guided static selection of a subset of base experts for adaptation\), and FFT \(full fine\-tuning\)\. All methods are applied only to the MoE modules while keeping all other modules frozen\.
##### Datasets\.
We organize datasets by domain\.*Math:*fine\-tune on MetaMathQA\-R1\(Yu et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib36)\); evaluate on GSM8K\(Cobbe et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib5)\)and MATH\-500\(Lightman et al\.,[2023](https://arxiv.org/html/2607.21978#bib.bib20)\)\.*Code:*fine\-tune on MagicCoder\-OSS\(Wei et al\.,[2023](https://arxiv.org/html/2607.21978#bib.bib33)\)\(7575K samples\); evaluate on MBPP\(Austin et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib1)\)and HumanEval\(Chen et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib3)\)\.*General retention:*evaluation on five general\-domain benchmarks \(MMLU\(Hendrycks et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib15)\), WinoGrande\(Sakaguchi et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib29)\), ARC\-Challenge\(Clark et al\.,[2018](https://arxiv.org/html/2607.21978#bib.bib4)\), StrategyQA\(Geva et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib11)\), and CommonsenseQA\(Talmor et al\.,[2019](https://arxiv.org/html/2607.21978#bib.bib30)\)\) using the code fine\-tuned checkpoints to measure general capability retention after domain\-specific adaptation\.
##### Fair comparison\.
We keep all essential evaluation criteria identical for all evaluated baselines: sharing the same training and testing recipe\. For all evaluated methods except FFT, the trainable parameter budgets vary by at most∼10%\{\\sim\}10\\%\. \(per\-method counts and configurations are listed in Appendix[A\.2](https://arxiv.org/html/2607.21978#A1.SS2), and training\-time and memory comparisons with PEFT baselines are reported in Appendix[A\.4](https://arxiv.org/html/2607.21978#A1.SS4)\)\.
### 4\.2Main Experimental Results
Tab\.[1](https://arxiv.org/html/2607.21978#S4.T1)reports in\-domain performance \(Math and Code\) together with out\-of\-domain general retention across three MoE backbones\.
Across all three backbones, MoE2\-LoRA achieves the highest in\-domain average among PEFT methods while maintaining strong general capability retention after domain\-specific fine\-tuning\. The gains remain consistent across models with different scales and expert granularities, ranging from OLMoE\-1B\-7B to Qwen3\-30B\-A3B\. In particular, MoE2\-LoRA improves over prior PEFT\-on\-MoE methods on both math and code benchmarks simultaneously, whereas several baselines show a stronger trade\-off between in\-domain adaptation and out\-of\-domain retention\.
In\-domain \(Math \+ Code\)General RetentionMethodGSM8KMATHMBPPHEAvgMMLUWGARC\-CSQACSQAAvgOLMoE\-1B\-7BMuennighoff et al\. \([2024](https://arxiv.org/html/2607.21978#bib.bib27)\)Base1\.060\.223\.0114\.239\.6345\.0851\.2252\.8456\.7445\.5450\.28FFT46\.7017\.228\.7921\.7528\.6145\.0150\.6752\.0553\.4248\.4849\.93per\-expert LoRALiu et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib23)\)29\.117\.628\.7914\.4319\.9848\.4450\.9956\.8352\.1147\.5851\.19MoLAGao et al\. \([2025](https://arxiv.org/html/2607.21978#bib.bib10)\)26\.387\.829\.1815\.2419\.6548\.5851\.2257\.4250\.8046\.6850\.94MoELoRALuo et al\. \([2024](https://arxiv.org/html/2607.21978#bib.bib25)\)29\.915\.629\.5715\.8520\.2347\.7151\.4656\.4853\.8645\.2150\.94DAS\-LoRATang et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib31)\)34\.878\.229\.9616\.4622\.3747\.8251\.5457\.6852\.8445\.3751\.05MoE2\-LoRA36\.3211\.231\.9116\.8724\.0748\.5051\.4658\.1953\.2846\.6051\.61DeepSeek\-V2\-LiteDeepSeek\-AI \([2024](https://arxiv.org/html/2607.21978#bib.bib6)\)Base3\.710\.026\.4618\.9012\.2746\.0949\.9655\.0349\.3442\.5948\.60FFT65\.0526\.047\.8640\.4544\.8451\.5552\.1764\.4249\.4950\.1253\.55per\-expert LoRALiu et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib23)\)49\.2814\.447\.8632\.3235\.9751\.5951\.3064\.8549\.2049\.8053\.35MoLAGao et al\. \([2025](https://arxiv.org/html/2607.21978#bib.bib10)\)46\.2513\.849\.0334\.1535\.8151\.6050\.1264\.3349\.2052\.4253\.53MoELoRALuo et al\. \([2024](https://arxiv.org/html/2607.21978#bib.bib25)\)48\.5214\.649\.8133\.1336\.5251\.2250\.9164\.4251\.5349\.9653\.61DAS\-LoRATang et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib31)\)46\.5514\.649\.0333\.5435\.9349\.6451\.5461\.7751\.0949\.8052\.77MoE2\-LoRA52\.2418\.452\.1433\.5439\.0851\.7251\.8565\.5350\.2252\.0154\.27Qwen3\-30B\-A3BYang et al\. \([2025](https://arxiv.org/html/2607.21978#bib.bib35)\)Base73\.6930\.669\.2663\.4159\.2477\.7771\.8292\.8372\.4985\.3480\.05FFT95\.3074\.4072\.7682\.5281\.2374\.4671\.3592\.7561\.8681\.9876\.48per\-expert LoRALiu et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib23)\)94\.4771\.6072\.3782\.1180\.1475\.0572\.3893\.3470\.1684\.0378\.99MoLAGao et al\. \([2025](https://arxiv.org/html/2607.21978#bib.bib10)\)95\.0772\.6071\.9882\.1180\.4473\.7972\.7792\.2471\.0382\.7278\.51MoELoRALuo et al\. \([2024](https://arxiv.org/html/2607.21978#bib.bib25)\)94\.8472\.0073\.5482\.5280\.7368\.1771\.0386\.0172\.0575\.4374\.54DAS\-LoRATang et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib31)\)94\.6973\.2073\.9381\.7180\.8869\.1871\.1187\.8073\.6584\.0377\.15MoE2\-LoRA95\.2274\.4074\.3282\.3281\.5776\.2171\.5193\.6070\.0185\.8379\.43
Table 1:Main results across three MoE backbones on in\-domain tasks \(Math, Code\) and out\-of\-domain general retention\. Avg denotes the combined in\-domain average over GSM8K, MATH, MBPP, HumanEval\. Bold indicates the best PEFT result excluding FFT\.
### 4\.3Multimodal Experiments
To evaluate whether MoE2\-LoRA remains effective in multimodal settings, we further conduct experiments on a vision\-language MoE model in the medical imaging domain\.
##### Setup\.
We evaluate MoE2\-LoRA and two representative baselines onQwen3\.5\-35B\-A3B\(Qwen Team,[2026](https://arxiv.org/html/2607.21978#bib.bib28)\)\. For in\-domain medical VQA \(Vision Question Answering\) we fine\-tune on a mixture of LLaVA\-Med\(Li et al\.,[2023](https://arxiv.org/html/2607.21978#bib.bib19)\)and the train splits of three medical VQA datasets, and evaluate on the test splits of VQA\-RAD\(Lau et al\.,[2018](https://arxiv.org/html/2607.21978#bib.bib18)\), SLAKE\(Liu et al\.,[2021](https://arxiv.org/html/2607.21978#bib.bib22)\), and PathVQA\(He et al\.,[2020](https://arxiv.org/html/2607.21978#bib.bib14)\)\. For*general VL retention*, we evaluated the checkpoints trained on medical VQA datasets on MMBench\(Liu et al\.,[2024b](https://arxiv.org/html/2607.21978#bib.bib24)\), MME\(Fu et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib9)\), and RealWorldQA\(xAI,[2024](https://arxiv.org/html/2607.21978#bib.bib34)\)\.
##### Results\.
Tab\.[2](https://arxiv.org/html/2607.21978#S4.T2)reports both axes\. MoE2\-LoRA achieves the best accuracy on every in\-domain benchmark and the best positive transfer on every general benchmark, demonstrating the effectiveness of MoE2\-LoRA on multimodal settings\.
Medical VQAGeneral RetentionMethodVQA\-RADSLAKEPathVQAAvgMMBenchMMERealWorldQAAvgQwen3\.5\-35B\-A3BQwen Team \([2026](https://arxiv.org/html/2607.21978#bib.bib28)\)Base model69\.3272\.2850\.4764\.0290\.1491\.0346\.4175\.86MoELoRALuo et al\. \([2024](https://arxiv.org/html/2607.21978#bib.bib25)\)68\.2087\.3064\.8273\.4491\.6192\.0446\.1476\.60DAS\-LoRATang et al\. \([2026](https://arxiv.org/html/2607.21978#bib.bib31)\)68\.6087\.4064\.6073\.5391\.6192\.3846\.5476\.84MoE2\-LoRA69\.2587\.6865\.6774\.2092\.1992\.4247\.1977\.27Table 2:Medical VQA accuracy and general VL retention on Qwen3\.5\-35B\-A3B\.
### 4\.4Experiments on Capacity Scaling
We study scaling behavior under different trainable parameter budgets \(0\.12%0\.12\\%,0\.24%0\.24\\%, and0\.48%0\.48\\%\) on OLMoE math \(Tab\.[3](https://arxiv.org/html/2607.21978#S4.T3)\)\. As the trainable budget increases, MoE2\-LoRA consistently improves on GSM8K, rising from0\.37070\.3707at0\.12%0\.12\\%to0\.39800\.3980at0\.48%0\.48\\%\. In contrast, MoELoRA shows limited gains under larger budgets, particularly on GSM8K, where performance remains around0\.300\.30across all settings\.
The results show that MoE2\-LoRA maintains stronger scaling behavior than MoELoRA under increased adaptation capacity\.
MethodParamsGSM8KMATH\-500MoELoRA0\.12%29\.915\.6MoELoRA0\.24%30\.108\.4MoELoRA0\.48%30\.029\.4MoE2\-LoRA0\.12%37\.078MoE2\-LoRA0\.24%39\.589\.2MoE2\-LoRA0\.48%39\.810\.6
Table 3:Capacity scaling on OLMoE math\. MoE2\-LoRA scales more effectively than MoELoRA\.
## 5Analysis and Discussion
In this section, we first implement ablation studies on MoE2\-LoRA to evaluate the effectiveness of its components\. We then conduct in\-depth analyses on the working mechanisms of the RCP and the global expert pool, uncovering the underlying advantages of our designs\.
### 5\.1Component Ablation
We isolate MoE2\-LoRA’s two architectural choices, routing\-conditioned projection \(RCP\) and the shared global pool,through a matched\-budget ablation chain on DeepSeek\-V2\-Lite \(math\)\. Starting from a MoELoRA\-style baseline, we first replace hidden\-state routing with RCP while keeping the per\-layer pool fixed, and then introduce the globally shared pool on top of RCP\.
VariantPoolRoutingGSM8KMATH\-500AvgBaselineper\-layerhidden\-state48\.5214\.631\.56\+ RCPper\-layerRCP49\.0515\.432\.23\+ Global Pool \(MoE2\-LoRA\)globalRCP52\.2418\.435\.32Global Onlyglobalhidden\-state48\.0716\.232\.14Router\-only \(per\-layer\)per\-layerzzonly45\.4916\.831\.15Router\-only \(global\)globalzzonly46\.1715\.830\.99Table 4:Component ablation on DeepSeek\-V2\-Lite math\.Tab\.[4](https://arxiv.org/html/2607.21978#S5.T4)shows the main ablation chain \(top three rows\) together with additional reference variants \(bottom\)\. Replacing hidden\-state routing with RCP under a fixed per\-layer pool improves Avg from31\.5631\.56to32\.2332\.23, showing that incorporating pretrained router structure improves LoRA expert selection even without cross\-layer sharing\. Introducing the globally shared pool on top of RCP further improves Avg to35\.3235\.32, indicating that global sharing and routing\-conditioned projection are complementary\.
We further ablate thehch\_\{c\}\(Auxiliary channel\) by routing only from pretrained router logitsz∈ℝNEz\\in\\mathbb\{R\}^\{N\_\{E\}\}\. Both router\-only variants underperform the full RCP design, indicating that pretrained routing structure alone is insufficient\. Effective LoRA routing requires combining inherited router priors with input\-dependent adaptation signals from the hidden state\.
### 5\.2Analysis on RCP
Figure 3:Per\-layer routing\-alignment F\-statistic on DeepSeek\-V2\-Lite: MoE2\-LoRA \(red\) versus the independent\-routing baseline \(gray\)\.To verify that RCP’s per\-layer projectionWℓW\_\{\\ell\}actually inherits information from the base MoE router rather than learning an independent routing function, we measure the F\-statistic of structural alignment between LoRA routing and base routing, with details provided in Appendix[B\.2](https://arxiv.org/html/2607.21978#A2.SS2)\. A larger F\-statistic indicates that LoRA expert selection varies more systematically with the base router’s expert assignments, and therefore better reflects the routing structure learned by the pretrained MoE model\.
On DeepSeek\-V2\-Lite, MoE2\-LoRA achievesF=75\.0F=75\.0, compared withF=26\.4F=26\.4for MoELoRA, yielding a2\.8×2\.8\\timesmargin\. Fig\.[3](https://arxiv.org/html/2607.21978#S5.F3)further breaks down this comparison across layers\. MoE2\-LoRA obtains a higher F\-statistic on all 26 layers\. This layer\-wise consistency suggests that the alignment does not come from a few isolated layers, but is a systematic effect of conditioning LoRA routing on pretrained router logits\.
These results show that RCP produces LoRA routing patterns that are substantially more aligned with the base MoE routing structure than MoELoRA’s independent router\. This provides direct evidence that the pretrained router logits are effectively reflected in adapter selection, supporting the intended role of RCP as a router\-conditioned projection rather than an independently learned LoRA router\.
### 5\.3Analysis on Global Pool
The shared pool architecture imposes no explicit depth constraint on expert usage\. We analyze the resulting organization of the trained pool on DeepSeek\-V2\-Lite, focusing on its layer\-aware depth organization \(§[5\.3\.1](https://arxiv.org/html/2607.21978#S5.SS3.SSS1)\), non\-uniform capacity allocation \(§[5\.3\.2](https://arxiv.org/html/2607.21978#S5.SS3.SSS2)\), and the cross\-layer representational overlap that may support such sharing \(§[5\.3\.3](https://arxiv.org/html/2607.21978#S5.SS3.SSS3)\)\.
For the trained\-model probes, we run forward passes of the trained MoE2\-LoRA on 300 GSM8K test prompts and collect the setTTof all token positions\. For each layerℓ\\elland LoRA expertee, we record the top\-KLK\_\{L\}selection frequency
fℓ,e=1\|T\|∑t∈T𝟙\[e∈𝒮htLoRA\],f\_\{\\ell,e\}\\;=\\;\\frac\{1\}\{\|T\|\}\\sum\_\{t\\in T\}\\mathbb\{1\}\\\!\\bigl\[e\\in\\mathcal\{S\}\_\{h\_\{t\}\}^\{\\mathrm\{LoRA\}\}\\bigr\],where𝒮htLoRA\\mathcal\{S\}\_\{h\_\{t\}\}^\{\\mathrm\{LoRA\}\}is the activated LoRA set for tokenttat layerℓ\\ell\(§[3\.2](https://arxiv.org/html/2607.21978#S3.SS2)\)\. We use the discrete top\-KLK\_\{L\}frequency rather than the soft routing weight to align with the experts that the model actually engages at inference\.
#### 5\.3\.1Global Pool Learns Layer\-aware Expert Assignment
A potential concern of a shared global LoRA pool is that experts may be used in an unstructured manner across layers, without reflecting layer\-specific adaptation behavior\. We therefore analyze whether the learned routing exhibits depth\-dependent organization\. For each expertee, we compute its preferred depth
ℓ¯e=∑ℓℓ⋅fℓ,e∑ℓfℓ,e,\\bar\{\\ell\}\_\{e\}=\\frac\{\\sum\_\{\\ell\}\\ell\\cdot f\_\{\\ell,e\}\}\{\\sum\_\{\\ell\}f\_\{\\ell,e\}\},wherefℓ,ef\_\{\\ell,e\}denotes the activation frequency of experteeat layerℓ\\ell\. Experts are then sorted byℓ¯e\\bar\{\\ell\}\_\{e\}\.
Fig\.[1](https://arxiv.org/html/2607.21978#S1.F1)\(III\) visualizes the activation matrixfℓ,ef\_\{\\ell,e\}defined above, with MoE layers on the x\-axis, LoRA experts sorted by preferred depthℓ¯e\\bar\{\\ell\}\_\{e\}on the y\-axis, and color intensity denoting activation frequency\. The resulting diagonal structure, quantified by a Spearman correlation ofρ=0\.92\\rho=0\.92betweenℓ¯e\\bar\{\\ell\}\_\{e\}and each expert’s most frequently activated layer, indicates that expert usage is depth\-structured rather than layer\-agnostic\.
Notably, the resulting structure is localized rather than strictly layer\-exclusive\. Individual experts typically concentrate their activation mass within a narrow band of neighboring layers, instead of being activated by only a single layer\. As a result, the global pool simultaneously preserves layer\-dependent specialization and enables cross\-layer expert reuse\.
Appendix[B\.1](https://arxiv.org/html/2607.21978#A2.SS1)shows that this depth\-organized structure remains consistent across different pool sizesNL∈\{128,256,512\}N\_\{L\}\\in\\\{128,256,512\\\}and across both math and code probes \(Figs\.[6](https://arxiv.org/html/2607.21978#A2.F6),[7](https://arxiv.org/html/2607.21978#A2.F7)\)\.
#### 5\.3\.2Per\-layer Expert Capacity Is Non\-uniform
Figure 4:Per\-layer effective expert count on math \(blue\) and code \(orange\) probes\. Dashed line: the fixed allocation used by MoELoRA\.For each layerℓ\\ell, we compute the effective expert count as the perplexity of its usage distribution,exp\(H\(fℓ\)\)\\exp\(H\(f\_\{\\ell\}\)\)\. Fig\.[4](https://arxiv.org/html/2607.21978#S5.F4)shows that the learned allocation is highly non\-uniform: the effective expert count ranges from≈15\\approx 15in the deepest layers to≈49\\approx 49in the shallowest layers, a3\.3×3\.3\\timesvariation\.
This result suggests that adaptation demand varies substantially across depth, which is difficult to capture with fixed per\-layer allocation\. MoELoRA assigns the same number of experts to each layer \(N=20N\{=\}20; dashed line\), while MoLA requires manually specifying a layer\-wise schedule such as bottom\-heavy, top\-heavy, or hourglass\. In contrast, MoE2\-LoRA exposes no explicit architectural knob for per\-layer capacity, yet learns a non\-uniform allocation automatically\.
#### 5\.3\.3Cross\-layer Representational Overlap Supports Pool Sharing
We further examine why sharing LoRA experts across layers is feasible\. Specifically, we measure linear CKA\(Appendix[B\.3](https://arxiv.org/html/2607.21978#A2.SS3)\) between MoE block outputs of every layer pair on the frozen DeepSeek\-V2\-Lite base model over200200GSM8K prompts\.
Figure 5:Linear CKA between MoE block outputs across the2626layers of DeepSeek\-V2\-Lite \(base model, no LoRA\)\.Fig\.[5](https://arxiv.org/html/2607.21978#S5.F5)shows clear cross\-layer similarity: off\-diagonal CKA averages0\.430\.43, with adjacent layers averaging0\.560\.56compared with0\.350\.35for distant pairs \(\|i−j\|≥10\|i\{\-\}j\|\{\\geq\}10\)\. This depth\-local representational overlap provides a plausible basis for sharing LoRA experts across nearby or related layers, consistent with the soft depth\-affinity pattern observed above\.
## 6Conclusion
We introduced MoE2\-LoRA, an MoE\-style PEFT method that deeply binds the MoE mechanisms of pretrained MoE models and MoE\-style low\-rank adaptation through routing\-conditioned projection \(RCP\) and a globally shared LoRA expert pool\. Across four MoE backbones spanning different scales and expert granularities, MoE2\-LoRA achieves strong downstream performance while preserving general capabilities under matched training budgets\. Further analysis shows that MoE2\-LoRA learns structured cross\-layer adaptation while preserving alignment with pretrained MoE routing, confirming the effectiveness of its model\-level design\. Together, these results show that MoE2\-LoRA is an effective approach to PEFT for MoE models\.
## Limitations
##### Adapter merging is not straightforward\.
Like other routing\-based PEFT methods such as MoELoRA, MoE2\-LoRA relies on dynamic token\-wise expert selection rather than a fixed low\-rank update\. As a result, the learned adaptation cannot be directly merged into the base weights in the same manner as standard LoRA\.
Developing mergeable approximations for dynamic routing\-based PEFT architectures remains an important direction for future work\.
##### Scaling the global pool increases projection parameter cost\.
The globally shared expert pool improves adaptation flexibility, but scaling the pool size also increases the parameter cost of the routing/projection interface, particularly the layer\-specific projection matrices used for expert selection\. A naive design that routes directly from hidden states would cause these projection parameters to grow rapidly with both hidden dimension and pool size\.
MoE2\-LoRA mitigates this issue by performing routing in the pretrained router space rather than the full hidden\-state space, substantially reducing projection parameter growth while preserving routing expressiveness\. Nevertheless, scaling to extremely large shared pools may still require more parameter\-efficient routing mechanisms\.
## References
- Austin et al\. \(2021\)Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and 1 others\. 2021\.Program synthesis with large language models\.*arXiv preprint arXiv:2108\.07732*\.
- Bai et al\. \(2025\)Jun Bai, Minghao Tong, Yang Liu, Zixia Jia, and Zilong Zheng\. 2025\.Understanding and leveraging the expert specialization of context faithfulness in mixture\-of\-experts llms\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*\.
- Chen et al\. \(2021\)Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, and 1 others\. 2021\.Evaluating large language models trained on code\.*arXiv preprint arXiv:2107\.03374*\.
- Clark et al\. \(2018\)Peter Clark and 1 others\. 2018\.Think you have solved question answering? try arc, the ai2 reasoning challenge\.*arXiv preprint arXiv:1803\.05457*\.
- Cobbe et al\. \(2021\)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others\. 2021\.Training verifiers to solve math word problems\.*arXiv preprint arXiv:2110\.14168*\.
- DeepSeek\-AI \(2024\)DeepSeek\-AI\. 2024\.Deepseek\-v2: A strong, economical, and efficient mixture\-of\-experts language model\.*arXiv preprint arXiv:2405\.04434*\.
- Deng et al\. \(2026\)Guanzhi Deng, Bo Li, Ronghao Chen, Huacan Wang, Lijie Wen, and Linqi Song\. 2026\.Dr\-lora: Dynamic rank lora for mixture\-of\-experts adaptation\.*arXiv preprint arXiv:2601\.04823*\.
- Dou et al\. \(2023\)Shihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Jun Zhao, Wei Shen, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Xiaoran Fan, and 1 others\. 2023\.Loramoe: Revolutionizing mixture of experts for maintaining world knowledge in language model alignment\.*arXiv preprint arXiv:2312\.09979*, 4\(7\)\.
- Fu et al\. \(2025\)Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, and Ran He\. 2025\.[Mme: A comprehensive evaluation benchmark for multimodal large language models](https://arxiv.org/abs/2306.13394)\.*Preprint*, arXiv:2306\.13394\.
- Gao et al\. \(2025\)Chongyang Gao, Kezhen Chen, Jinmeng Rao, Ruibo Liu, Baochen Sun, Yawen Zhang, Daiyi Peng, Xiaoyuan Guo, and VS Subrahmanian\. 2025\.Mola: Moe lora with layer\-wise expert allocation\.In*Findings of the Association for Computational Linguistics: NAACL 2025*, pages 5097–5112\.
- Geva et al\. \(2021\)Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant\. 2021\.Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies\.*Transactions of the Association for Computational Linguistics \(TACL\)*\.
- Hayou et al\. \(2024\)Soufiane Hayou, Nikhil Ghosh, and Bin Yu\. 2024\.Lora\+: Efficient low rank adaptation of large models\.*arXiv preprint arXiv:2402\.12354*\.
- He et al\. \(2025\)Haonan He, Peng Ye, Yuchen Ren, Yuan Yuan, Luyang Zhou, Shucun Ju, and Lei Chen\. 2025\.[Gora: Gradient\-driven adaptive low rank adaptation](https://arxiv.org/abs/2502.12171)\.*Preprint*, arXiv:2502\.12171\.
- He et al\. \(2020\)Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie\. 2020\.Pathvqa: 30000\+ questions for medical visual question answering\.*arXiv preprint arXiv:2003\.10286*\.
- Hendrycks et al\. \(2021\)Dan Hendrycks and 1 others\. 2021\.Measuring massive multitask language understanding\.In*ICLR*\.
- Hu et al\. \(2022\)Edward J\. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\. 2022\.Lora: Low\-rank adaptation of large language models\.In*International Conference on Learning Representations \(ICLR\)*\.
- Jacobs et al\. \(1991\)Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton\. 1991\.Adaptive mixtures of local experts\.*Neural computation*, 3\(1\):79–87\.
- Lau et al\. \(2018\)Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner\-Fushman\. 2018\.A dataset of clinically generated visual questions and answers about radiology images\.*Scientific data*, 5\(1\):180251\.
- Li et al\. \(2023\)Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao\. 2023\.[Llava\-med: Training a large language\-and\-vision assistant for biomedicine in one day](https://arxiv.org/abs/2306.00890)\.*Preprint*, arXiv:2306\.00890\.
- Lightman et al\. \(2023\)Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\. 2023\.Let’s verify step by step\.*arXiv preprint arXiv:2305\.20050*\.
- Liu et al\. \(2024a\)Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, and 1 others\. 2024a\.Deepseek\-v2: A strong, economical, and efficient mixture\-of\-experts language model\.*arXiv preprint arXiv:2405\.04434*\.
- Liu et al\. \(2021\)Bo Liu, Li\-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao\-Ming Wu\. 2021\.Slake: A semantically\-labeled knowledge\-enhanced dataset for medical visual question answering\.In*2021 IEEE 18th international symposium on biomedical imaging \(ISBI\)*, pages 1650–1654\. IEEE\.
- Liu et al\. \(2026\)Yilun Liu, Yunpu Ma, Yuetian Lu, Shuo Chen, Zifeng Ding, and Volker Tresp\. 2026\.[Parameter\-efficient routed fine\-tuning: Mixture\-of\-experts demands mixture of adaptation modules](https://doi.org/10.18653/v1/2026.findings-eacl.232)\.In*Findings of the Association for Computational Linguistics: EACL 2026*, pages 4439–4457, Rabat, Morocco\. Association for Computational Linguistics\.
- Liu et al\. \(2024b\)Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin\. 2024b\.[Mmbench: Is your multi\-modal model an all\-around player?](https://arxiv.org/abs/2307.06281)*Preprint*, arXiv:2307\.06281\.
- Luo et al\. \(2024\)Tongxu Luo, Jiahe Lei, Fangyu Lei, Weihao Liu, Shizhu He, Jun Zhao, and Kang Liu\. 2024\.[Moelora: Contrastive learning guided mixture of experts on parameter\-efficient fine\-tuning for large language models](https://arxiv.org/abs/2402.12851)\.*Preprint*, arXiv:2402\.12851\.
- Manzoni \(2026\)Andrea Manzoni\. 2026\.Moe\-sieve: Routing\-guided lora for efficient moe fine\-tuning\.*arXiv preprint arXiv:2603\.24044*\.
- Muennighoff et al\. \(2024\)Niklas Muennighoff, Luca Soldaini, and 1 others\. 2024\.Olmoe: Open mixture\-of\-experts language models\.*arXiv preprint arXiv:2409\.02060*\.
- Qwen Team \(2026\)Qwen Team\. 2026\.[Qwen3\.5: Towards native multimodal agents](https://qwen.ai/blog?id=qwen3.5)\.
- Sakaguchi et al\. \(2021\)Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi\. 2021\.Winogrande: An adversarial winograd schema challenge at scale\.*Communications of the ACM*, 64\(9\):99–106\.
- Talmor et al\. \(2019\)Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant\. 2019\.Commonsenseqa: A question answering challenge targeting commonsense knowledge\.In*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pages 4149–4158\.
- Tang et al\. \(2026\)Yiru Tang, Kun Zhou, Xin Zhao, Jing Sha, Zhichao Sheng, and Shijin Wang\. 2026\.[Exploring expert concentration for parameter\-efficient fine\-tuning of mixture\-of\-expert LLMs](https://openreview.net/forum?id=zBgjWTWgCh)\.
- Wang et al\. \(2024\)Zihan Wang and 1 others\. 2024\.Let the expert stick to his last: Expert\-specialized fine\-tuning for sparse architectural large language models\.*arXiv preprint arXiv:2407\.01906*\.
- Wei et al\. \(2023\)Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang\. 2023\.Magicoder: Empowering code generation with oss\-instruct\.*arXiv preprint arXiv:2312\.02120*\.
- xAI \(2024\)xAI\. 2024\.Grok\-1\.5 Vision Preview\.[https://x\.ai/blog/grok\-1\.5v](https://x.ai/blog/grok-1.5v)\.RealWorldQA dataset, available at[https://huggingface\.co/datasets/xai\-org/RealworldQA](https://huggingface.co/datasets/xai-org/RealworldQA)\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others\. 2025\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*\.
- Yu et al\. \(2024\)Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu\. 2024\.Metamath: Bootstrap your own mathematical questions for large language models\.In*International Conference on Learning Representations*, volume 2024, pages 45040–45061\.
## Appendix AReproducibility Details
### A\.1Evaluation Protocol
This appendix specifies the evaluation protocol used in all main\-text experiments\. The protocol is shared verbatim across MoE2\-LoRA and all baselines; only the trained adapter weights differ\.
##### Inference configuration\.
The decoding configuration is shared across all methods within each benchmark; only the trained adapter weights differ\. All benchmarks are evaluated in a strictly zero\-shot setting without in\-context examples\. GSM8K, MATH\-500, and MBPP use greedy decoding \(do\_sample=False,temperature=0\), while HumanEval uses stochastic decoding withtemperature=0\.2andtop\_p=0\.95\. We usemax\_new\_tokens=1024for all benchmarks\. MBPP is evaluated with pass@1, and for HumanEval we generaten=3n\{=\}3samples per problem and report the average pass@1 across runs to reduce variance on the small 164\-task test set\.
##### Prompt templates\.
We use a single fixed prompt template per benchmark, applied identically to every method:
- •GSM8K:"Solve the following math problem step by step\. End your answer with ‘\#\#\#\# <number\>’\.\\n\\nQuestion: \{question\}\\nAnswer:"
- •MATH\-500:"Solve the following math problem step by step\. Put your final answer in \\boxed\{\.\.\.\}\. \\n\\nProblem: \{problem\}\\nSolution:"
- •MBPP / HumanEval: standard function\-completion prompts \(problem description plus public test signature\); the model is asked to produce the function body, which is then executed against the held\-out unit tests\.
##### Answer extraction\.
Answer extraction is identical across all methods\. For GSM8K, we extract the final numeric answer using a regex match against\#\#\#\#\\s\*\(\-?\[0\-9\.,\]\+\)and evaluate after numeric normalization\. For MATH\-500, we extract the contents of the first\\boxed\{…\}expression \(including nested braces\) and evaluate after whitespace and LaTeX\-symbol normalization\. For MBPP and HumanEval, the generated function body is parsed from the completion and executed against the official unit tests, with pass@1 reported following the standard benchmark protocols\.
### A\.2Hyperparameters
##### MoE2\-LoRA\.
Table[5](https://arxiv.org/html/2607.21978#A1.T5)summarizes the key hyperparameters used for MoE2\-LoRA across different backbones\. Here,NLN\_\{L\}denotes the size of the shared LoRA expert pool, rank the per\-expert LoRA rank, top\-KKthe number of activated LoRA experts per token, anddbd\_\{b\}the hidden\-state bottleneck dimension used in routing\-conditioned projection\.
BackboneNLN\_\{L\}ranktop\-KKdbd\_\{b\}\#paramsOLMoE\-1B\-7B128164169\.1MDeepSeek\-V2\-Lite2561666421MQwen3\-30B\-A3B2563263239MTable 5:MoE2\-LoRA hyperparameters across backbones\. On Qwen3\-30B\-A3B the pool is further partitioned intoG=2G\{=\}2groups \(hierarchical routing\); the other two backbones use a single flat pool \(G=1G\{=\}1\)\.
##### MoELoRA\(Luo et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib25)\)\.
We use a per\-layer LoRA expert pool with dynamic routing from hidden states\. The configuration uses 16 experts \(rank 8, top\-22\) on OLMoE\-1B\-7B, 32 experts \(rank 6, top\-44\) on DeepSeek\-V2\-Lite, and 32 experts \(rank 6, top\-44\) on Qwen3\-30B\-A3B, corresponding to approximately8\.98\.9M,2121M, and4141M trainable parameters\.
##### MoLA\(Gao et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib10)\)\.
MoLA uses a per\-layer LoRA expert pool with a layer\-wise increasing allocation, motivated by the observation that deeper transformer layers benefit from more LoRA experts\. LoRA experts are routed independently from hidden states \(top\-22per token\)\. Following the original paper’s allocation scheme, we use ranks88,2222, and4040on OLMoE\-1B\-7B, DeepSeek\-V2\-Lite, and Qwen3\-30B\-A3B respectively, with per\-layer expert counts ranging from88to2424\(OLMoE, 4 stages\),44to1616\(DeepSeek, 26 layers\), and22to88\(Qwen3, 4 stages\)\. On DeepSeek\-V2\-Lite we scale MoLA’s rank from the original1616to2222to match the parameter budget of other baselines\. This yields approximately8\.98\.9M,21\.421\.4M, and3939M trainable parameters\.
##### DAS\-LoRA\(Tang et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib31)\)\.
For DAS\-LoRA, we retain the original CDAS\-based expert selection procedure but restrict training to the selected LoRA modules only, without additionally tuning router or dense backbone parameters, in order to maintain comparable PEFT settings across methods\. Following the original paper, we pre\-select a subset of base experts using Domain Advantage Score and attach LoRA modules only to the selected experts\. We use rank3232on OLMoE\-1B\-7B,5252on DeepSeek\-V2\-Lite, and4848on Qwen3\-30B\-A3B, yielding approximately8\.48\.4M,2222M, and3838M trainable parameters\.
##### Training\.
Within each backbone, all PEFT methods share the same optimizer \(AdamW\), cosine learning\-rate schedule with warmup ratio0\.10\.1, batch size, and number of epochs\. Method\-specific hyperparameters are tuned on the validation split under comparable trainable parameter budgets\.
### A\.3Training Compute
All experiments are conducted on a single 8×\\timesNVIDIA H800 80GB node using bfloat16 training\. DeepSeek\-V2\-Lite and OLMoE use standard DDP, while Qwen3\-30B\-A3B uses DeepSpeed ZeRO\-3\. Typical wall\-clock training time per run is approximately 35 minutes for OLMoE, 1\.5 hours for DeepSeek\-V2\-Lite, and 2\.5–3 hours for Qwen3\-30B\-A3B\. The total compute budget for all experiments is approximately 200 GPU\-hours\.
##### Reproducibility note\.
Each adapter is trained once per \(method, backbone\) configuration\. We report the per\-benchmark scores from that run directly, following the evaluation protocol in Appendix[A\.1](https://arxiv.org/html/2607.21978#A1.SS1): a single greedy sample for GSM8K, MATH\-500 and MBPP, andn=3n\{=\}3stochastic samples \(temperature0\.20\.2\) averaged for HumanEval\. Multi\-run training\-time statistics are not reported due to compute cost\.
### A\.4Training\-time and memory comparability with PEFT baselines
Table[6](https://arxiv.org/html/2607.21978#A1.T6)compares training time and peak training memory on DeepSeek\-V2\-Lite using 80k samples extracted from MetaMathQA for one training epoch\. Training time is measured on8×8\\timesH800 80GB GPUs with DDP, while peak memory is measured on a single H800 without ZeRO using batch size11and sequence length20482048\. All PEFT baselines are matched to roughly2222M trainable parameters; per\-expert LoRA is included for reference\.
MethodTrainableParamsTrainTime \(min\)Peak TrainMem \(GB\)MoE2\-LoRA21\.7M16\.644\.65MoELoRA\(Luo et al\.,[2024](https://arxiv.org/html/2607.21978#bib.bib25)\)22\.2M19\.843\.6MoLA\(Gao et al\.,[2025](https://arxiv.org/html/2607.21978#bib.bib10)\)21\.4M17\.242\.7DAS\-LoRA\(Tang et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib31)\)22\.2M17\.841\.8per\-expert LoRA\(Liu et al\.,[2026](https://arxiv.org/html/2607.21978#bib.bib23)\)†27\.3M23\.244\.3FFT \(DeepSeek\-V2\-Lite reference\)‡16B143120Table 6:Training\-time efficiency on DeepSeek\-V2\-Lite math\.†Per\-expert LoRA cannot be configured to a 22M budget without collapsing to rank\-1; we report its native 27M configuration\.All PEFT methods exhibit broadly comparable training\-time and memory cost, with training time falling within a narrow1717–2020minute range and peak memory between41\.841\.8–46\.446\.4GB\. MoE2\-LoRA lies slightly above the lightest baselines in both metrics, largely because it activates more LoRA experts per token\. Compared with per\-expert LoRA, the globally shared pool also avoids the larger training overhead associated with materializing independent LoRA modules for every expert\.
## Appendix BAdditional Analysis Details
### B\.1Robustness of the Depth\-Organized Expert Allocation
To verify that the depth\-organized diagonal observed in Section[5\.3\.1](https://arxiv.org/html/2607.21978#S5.SS3.SSS1)is not specific to a particular configuration, we reproduce the analysis across \(i\) different global pool sizes and \(ii\) different probe domains\.
##### Across pool sizes\.
Fig\.[6](https://arxiv.org/html/2607.21978#A2.F6)shows the expert\-affinity matrices on GSM8K forNL∈\{128,256,512\}N\_\{L\}\\in\\\{128,256,512\\\}\. The same depth\-organized diagonal appears in all three settings, indicating that experts consistently specialize to contiguous depth regions\. AsNLN\_\{L\}increases, the diagonal becomes thinner because the 26\-layer depth axis is partitioned among more experts; smaller pools produce correspondingly broader depth bands\.
Figure 6:Expert\-affinity matrices on GSM8K for MoE2\-LoRA with different global pool sizesNL∈\{128,256,512\}N\_\{L\}\\in\\\{128,256,512\\\}on DeepSeek\-V2\-Lite\. Rows are experts sorted by preferred depth; columns are MoE layers\. A consistent depth\-organized diagonal emerges across all pool sizes\.
##### Across domains\.
Fig\.[7](https://arxiv.org/html/2607.21978#A2.F7)compares GSM8K and MBPP using the same trained adapter \(NL=512N\_\{L\}=512\)\. Both domains exhibit the same depth\-organized structure, despite the code probe appearing slightly noisier due to the smaller MBPP evaluation set\.
Figure 7:Expert\-affinity matrices for the same trained adapter evaluated on GSM8K \(left\) and MBPP \(right\)\. Each panel is independently sorted by preferred depth\. The depth\-organized structure is preserved across domains\.Across both pool sizes and probe domains, the same depth\-organized allocation consistently emerges, suggesting that it is an intrinsic property of the shared global pool with routing\-conditioned projection rather than an artifact of a particular setting\.
### B\.2F\-statistic for routing alignment
The routing\-alignment F\-statistic reported in Section[5\.2](https://arxiv.org/html/2607.21978#S5.SS2)is a one\-way ANOVA statistic that measures how much of the variance in LoRA routing can be explained by the base MoE router’s expert assignment—higher values indicate stronger structural inheritance of the pretrained routing prior\.
Formally, letTTdenote the token set,𝐩L\(t\)∈ℝNL\\mathbf\{p\}\_\{L\}\(t\)\\in\\mathbb\{R\}^\{N\_\{L\}\}the LoRA routing distribution for tokentt, andeb\(t\)=argmaxe𝐩B\(t\)e\_\{b\}\(t\)=\\arg\\max\_\{e\}\\,\\mathbf\{p\}\_\{B\}\(t\)the base router’s top\-1 expert fortt\. We partitionTTintoKKgroups indexed by the base expert,Ge=\{t∈T:eb\(t\)=e\}G\_\{e\}=\\\{t\\in T:e\_\{b\}\(t\)=e\\\}, with sizesne=\|Ge\|n\_\{e\}=\|G\_\{e\}\|\(groups withne<5n\_\{e\}<5are dropped to avoid noisy estimates\)\. Let𝝁e=1ne∑t∈Ge𝐩L\(t\)\\boldsymbol\{\\mu\}\_\{e\}=\\tfrac\{1\}\{n\_\{e\}\}\\sum\_\{t\\in G\_\{e\}\}\\mathbf\{p\}\_\{L\}\(t\)denote the group\-mean LoRA distribution and𝝁¯=1\|T\|∑t𝐩L\(t\)\\bar\{\\boldsymbol\{\\mu\}\}=\\tfrac\{1\}\{\|T\|\}\\sum\_\{t\}\\mathbf\{p\}\_\{L\}\(t\)the grand mean\. The between\- and within\-group sums of squares are
SSbetween\\displaystyle\\mathrm\{SS\}\_\{\\text\{between\}\}=∑e=1Kne‖𝝁e−𝝁¯‖22,\\displaystyle=\\sum\_\{e=1\}^\{K\}n\_\{e\}\\,\\\|\\boldsymbol\{\\mu\}\_\{e\}\-\\bar\{\\boldsymbol\{\\mu\}\}\\\|\_\{2\}^\{2\},\(3\)SSwithin\\displaystyle\\mathrm\{SS\}\_\{\\text\{within\}\}=∑e=1K∑t∈Ge‖𝐩L\(t\)−𝝁e‖22,\\displaystyle=\\sum\_\{e=1\}^\{K\}\\sum\_\{t\\in G\_\{e\}\}\\\|\\mathbf\{p\}\_\{L\}\(t\)\-\\boldsymbol\{\\mu\}\_\{e\}\\\|\_\{2\}^\{2\},\(4\)and the F\-statistic is the ratio of mean squares
F=SSbetween/\(K−1\)SSwithin/\(\|T\|−K\)\.F\\;=\\;\\frac\{\\mathrm\{SS\}\_\{\\text\{between\}\}/\(K\-1\)\}\{\\mathrm\{SS\}\_\{\\text\{within\}\}/\(\|T\|\-K\)\}\.\(5\)A largerFFindicates that LoRA routing distributions are tightly clustered within each base\-expert group and well\-separated across groups—i\.e\., the LoRA router behaves predictably as a function of the base router\. We treatFFas a structural\-fit diagnostic for the RCP projectionWℓW\_\{\\ell\}rather than a quality metric; passive coupling schemes \(e\.g\., per\-expert LoRA\) attain trivially highFFby construction\.
### B\.3Linear CKA
Given centered layer representationsXi∈ℝn×diX\_\{i\}\\in\\mathbb\{R\}^\{n\\times d\_\{i\}\}andXj∈ℝn×djX\_\{j\}\\in\\mathbb\{R\}^\{n\\times d\_\{j\}\}, we compute linear centered kernel alignment \(CKA\) as
CKA\(Xi,Xj\)=‖Xi⊤Xj‖F2‖Xi⊤Xi‖F‖Xj⊤Xj‖F\.\\mathrm\{CKA\}\(X\_\{i\},X\_\{j\}\)=\\frac\{\\\|X\_\{i\}^\{\\top\}X\_\{j\}\\\|\_\{F\}^\{2\}\}\{\\\|X\_\{i\}^\{\\top\}X\_\{i\}\\\|\_\{F\}\\,\\\|X\_\{j\}^\{\\top\}X\_\{j\}\\\|\_\{F\}\}\.
Here∥⋅∥F\\\|\\cdot\\\|\_\{F\}denotes the Frobenius norm\. For each layer pair,XiX\_\{i\}andXjX\_\{j\}are constructed from the corresponding MoE block outputs collected over the evaluation prompts\.Similar Articles
Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation
This paper proposes a Mixture of LoRA and Full (MoLF) fine-tuning framework that uses gradient-guided optimizer routing to adaptively switch between LoRA and full fine-tuning. It aims to overcome the structural limitations of relying solely on static adaptation methods by combining the plasticity of full tuning with the regularization of LoRA.
HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
HELLoRA introduces activation-aware adapter placement for MoE models, attaching LoRA only to hot experts to reduce parameters and FLOPs while improving performance on reasoning, code, and safety tasks.
Hybrid-LoRA: Bridging Full Fine-Tuning and Low-Rank Adaptation for Post-Training
Hybrid-LoRA proposes a framework that selectively applies full fine-tuning to a small subset of modules while using LoRA for the rest, achieving performance near full fine-tuning with significantly lower computational cost. Experiments show improvements of up to 5.65% over existing parameter-efficient baselines.
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
This paper introduces LoRA-GA2, a fine-tuning algorithm that leverages multi-step gradient information to improve the performance of Low-Rank Adaptation for large language models, achieving better results on benchmarks while preserving efficiency.
@jbhuang0604: LoRA, low-rank adaptation, is arguably the most popular parameter-efficient fine-tuning method for LLMs. But how does i…
LoRA (low-rank adaptation) is the most popular parameter-efficient fine-tuning method for LLMs. This video introduces how LoRA and its variants (LoRA+, QLoRA, VeRA, DoRA) work.