Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
Summary
This paper investigates whether routing drift in merged Mixture-of-Experts Large Language Models indicates failure, and proposes a routing analysis toolkit and Selective Router Repair method for assessment.
View Cached Full Text
Cached at: 09/29/26, 04:08 AM
Paper page - Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
Source: https://huggingface.co/papers/2609.32821
Abstract
Modelmergingefficientlycombinesspecializedlargelanguagemodels(LLMs)withoutjointretraining,butcansubstantiallyalterexpertroutinginMixture-of-Experts(MoE)models.Suchroutingdriftisofteninterpretedasroutingfailure,raisingafundamentalquestionthatremainsunclear:doesroutingdriftafterMoEmergingactuallyindicateroutingfailure,andwhatevidenceshouldjustifyrepair?WeinvestigatethesequestionsacrossDeepSeekMoE,OLMoE,andQwen3-MoEproposingaroutinganalysistoolkitforcontrolledcounterfactualinterventionsandtoken-levelanalysis.Bycrossingsourceandmergedrouterinputsandparameters,weattributemostexpertreassignmentstoinputshiftsratherthanparameterchangesatthesamelayer.However,source-relativeroutingdifferencespoorlypredictnext-tokenlikelihoodgainsfromsource-routerestoration,anddifferentexpertselectionscanproducedirectionallysimilarmixtureoutputs.Wethereforeoperationalizeroutingfailureastasklossrecoverableunderaspecifiedroutingintervention,withnon-routingparametersfixed.Thesetestsdetectrecoverablelossunderdeliberateroutercorruption,whereassource-routerestorationdoesnotestablishreliabletaskbenefitsintheevaluatedmergedmodels.Motivatedbythese,weproposeSelectiveRouterRepair(SRR)asacasestudy,andfindthatsource-specialisttoken-likelihoodadvantagesdonotreliablyidentifybeneficiallocalcorrections.Together,thesefindingsshowthatroutingdriftaloneisinsufficientevidenceofroutingfailure:source-informedcorrectionsmustbejudgedbytheirtask-levelinterventioneffects.TheanalysistoolkitandSRRcodearereleased.
View arXiv pageView PDFGitHubAdd to collection
Get this paper in your agent:
hf papers read 2609\.32821
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.32821 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.32821 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.32821 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Routing-Aware Expert Calibration for Machine Unlearning in Mixture-of-Experts Language Models
The paper proposes TRACE, a method for machine unlearning in Mixture-of-Experts language models that calibrates retain regularization by reweighting token-level retain losses to address forget-retain routing mismatch. Experiments show improved forget-utility trade-off across multiple MoE LLMs.
A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models
This paper investigates whether bilingual Mixture-of-Experts (MoE) language models develop linguistically structured expert routing. It finds that interpretable linguistic organization emerges within MoE routing patterns, and that curriculum training influences specialization in language balance.
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
This paper proposes using router weight sensitivity under lightweight fine-tuning (e.g., LoRA) to identify and prune experts in Mixture-of-Experts models, enabling significant memory and latency reductions with minimal accuracy loss.
Attention-Aware Routing: Coupling Routing and Attention in MoEs
This paper introduces Attention-Aware Routing (AAR), a method that enhances Mixture-of-Experts language models by incorporating attention weights into the router, improving mathematical reasoning performance and revealing coupled dynamics between routing and attention.
Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing
This paper investigates the predictive value of earlier expert selections in sparse mixture-of-experts models beyond the most recent layer, finding significant gains in routing prediction across layers.