Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Hugging Face Daily Papers Papers

Summary

This paper investigates whether routing drift in merged Mixture-of-Experts Large Language Models indicates failure, and proposes a routing analysis toolkit and Selective Router Repair method for assessment.

Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such routing drift is often interpreted as routing failure, raising a fundamental question that remains unclear: does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair? We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as task loss recoverable under a specified routing intervention, with non-routing parameters fixed. These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose Selective Router Repair (SRR) as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that routing drift alone is insufficient evidence of routing failure: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:08 AM

Paper page - Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Source: https://huggingface.co/papers/2609.32821

Abstract

Modelmergingefficientlycombinesspecializedlargelanguagemodels(LLMs)withoutjointretraining,butcansubstantiallyalterexpertroutinginMixture-of-Experts(MoE)models.Suchroutingdriftisofteninterpretedasroutingfailure,raisingafundamentalquestionthatremainsunclear:doesroutingdriftafterMoEmergingactuallyindicateroutingfailure,andwhatevidenceshouldjustifyrepair?WeinvestigatethesequestionsacrossDeepSeekMoE,OLMoE,andQwen3-MoEproposingaroutinganalysistoolkitforcontrolledcounterfactualinterventionsandtoken-levelanalysis.Bycrossingsourceandmergedrouterinputsandparameters,weattributemostexpertreassignmentstoinputshiftsratherthanparameterchangesatthesamelayer.However,source-relativeroutingdifferencespoorlypredictnext-tokenlikelihoodgainsfromsource-routerestoration,anddifferentexpertselectionscanproducedirectionallysimilarmixtureoutputs.Wethereforeoperationalizeroutingfailureastasklossrecoverableunderaspecifiedroutingintervention,withnon-routingparametersfixed.Thesetestsdetectrecoverablelossunderdeliberateroutercorruption,whereassource-routerestorationdoesnotestablishreliabletaskbenefitsintheevaluatedmergedmodels.Motivatedbythese,weproposeSelectiveRouterRepair(SRR)asacasestudy,andfindthatsource-specialisttoken-likelihoodadvantagesdonotreliablyidentifybeneficiallocalcorrections.Together,thesefindingsshowthatroutingdriftaloneisinsufficientevidenceofroutingfailure:source-informedcorrectionsmustbejudgedbytheirtask-levelinterventioneffects.TheanalysistoolkitandSRRcodearereleased.

View arXiv pageView PDFGitHubAdd to collection

Get this paper in your agent:

hf papers read 2609\.32821

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.32821 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.32821 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.32821 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Attention-Aware Routing: Coupling Routing and Attention in MoEs

arXiv cs.AI

This paper introduces Attention-Aware Routing (AAR), a method that enhances Mixture-of-Experts language models by incorporating attention weights into the router, improving mathematical reasoning performance and revealing coupled dynamics between routing and attention.