RAZOR:剪枝大语言模型中的可替换专家

arXiv cs.LG 论文

摘要

RAZOR是一种无需训练的方法,用于在混合专家大语言模型中通过评估功能可替换性来剪枝可替换的专家,在推理任务上相比现有方法能取得更优性能。

arXiv:2609.30465v2 Announce Type: new Abstract: Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Common scores rank experts by routing frequency or output magnitude, which measures isolated contribution rather than deletion damage. What decides the damage is functional replaceability, whether the surviving computation can reproduce what is removed. A large contribution may be replaceable by the remaining mixture, whereas a small one may carry a direction the survivors cannot recover. We introduce RAZOR, a training-free method that scores replaceability from consensus residuals, the deviations of individual expert outputs from their original weighted mixture. Holding the layer input fixed, these residuals yield the exact output change from deleting one expert, including survivor reweighting and the replacement expert promoted by router refill. RAZOR aggregates this change over calibration tokens and prunes to a layerwise budget using forward passes alone, without gradients, subset search, or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25% and 50% expert removal, RAZOR attains the highest macro average over nine reasoning-centered tasks among the evaluated pruning methods in all eight model-budget settings. Against REAP on GLM-4.7-Flash and Qwen3.6-35B-A3B, it gains 2.12-5.59 points on this average and lowers reverse KL in all four comparisons. Retained accuracy is not the whole picture, as pruned Qwen3.6-35B-A3B still shifts in response diversity, formatting, and termination.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:37

# RAZOR: Pruning Replaceable Experts in LLMs
Source: [https://arxiv.org/html/2609.30465](https://arxiv.org/html/2609.30465)
Mao ZhengAffiliation:Foundation Model Department, Tencent, ChinaEmail:[nickmysong@tencent\.com](mailto:)Affiliation:[Code](https://github.com/nick7nlp/Razor)[Models](https://huggingface.co/collections/Nickyang/razor)

###### Abstract

Mixture\-of\-experts \(MoE\) models activate only a few experts per token yet store the entire expert pool\. Whole\-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability\. Common scores rank experts by routing frequency or output magnitude, which measures isolated contribution rather than deletion damage\. What decides the damage is functional replaceability, whether the surviving computation can reproduce what is removed\. A large contribution may be replaceable by the remaining mixture, whereas a small one may carry a direction the survivors cannot recover\. We introduceRazor, a training\-free method that scores replaceability from consensus residuals, the deviations of individual expert outputs from their original weighted mixture\. Holding the layer input fixed, these residuals yield the exact output change from deleting one expert, including survivor reweighting and the replacement expert promoted by router refill\.Razoraggregates this change over calibration tokens and prunes to a layerwise budget using forward passes alone, without gradients, subset search, or recovery training\. On GLM\-4\.7\-Flash, Qwen3\.6\-35B\-A3B, DeepSeek\-V4\-Flash\-0731, and Hy3 at 25% and 50% expert removal,Razorattains the highest macro average over nine reasoning\-centered tasks among the evaluated pruning methods in all eight model–budget settings\. Against REAP on GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B, it gains 2\.12–5\.59 points on this average and lowers reverse KL in all four comparisons\. Retained accuracy is not the whole picture, as pruned Qwen3\.6\-35B\-A3B still shifts in response diversity, formatting, and termination\.

## 1Introduction

Mixture\-of\-experts \(MoE\) models increase capacity while activating only a small subset of experts per token, yet sparse computation does not reduce the cost of storing the full expert pool\([Shazeer et al\., 2017](https://arxiv.org/html/2609.30465#bib.bib35);[Lepikhin et al\., 2021](https://arxiv.org/html/2609.30465#bib.bib17);[Fedus et al\., 2022](https://arxiv.org/html/2609.30465#bib.bib6)\)\. Whole\-expert pruning addresses this burden by reducing the pool itself\([Lu et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib26);[Muzio et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib30);[Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15);[Liu et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib22)\)\. Our focus is pruning reasoning models while retaining their reasoning ability\. Reasoning is the demanding case for pruning, since a long chain of dependent steps carries early errors forward, so damage that a single\-step benchmark would absorb can compound over a trajectory\. At a fixed budget, we therefore seek to preserve the output distribution by assessing*functional replaceability*, not simply contribution magnitude\.

Common pruning scores characterize experts by selection frequency or output magnitude\([Muzio et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib30);[Jaiswal et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib12);[Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15)\)\. These statistics do not capture how expert outputs combine, even when the output norm is weighted by routing mass\. A large contribution may be replaceable by the surviving mixture, whereas a small one may supply a component that the survivors cannot recover\. Isolated importance therefore differs from deletion damage, which depends on the computation available after removal\.

Figure 1:Replaceability depends on output geometry and router refill\.A constructed single\-token example without \(a\) and with \(b\) refill\. Damage is L2 distance to the original mixture\. Bold marks row minima\. Appendix[A](https://arxiv.org/html/2609.30465#A1)gives the construction and plotting details\.Deletion changes a routed mixture by renormalizing the survivors’ weights and, when the router refills the vacated slot, promoting an unselected expert\. Figure[1](https://arxiv.org/html/2609.30465#S1.F1)isolates both effects in a constructed single\-token example\. Without refill, panel \(a\) shows that removing the expert closest to the original mixture is less harmful than removing the one with the smallest output norm\. With refill, panel \(b\) shows that a promoted expert can compensate for a removed balancing contribution and change the preferred deletion\. Output magnitude, fixed\-support damage, and refill\-aware damage can therefore select different experts\. More generally, identical gates and output norms can correspond to different least\-harmful deletions \(Proposition[1](https://arxiv.org/html/2609.30465#Thmproposition1)\)\.

We score functional replaceability with three consensus\-residual criteria, each adding one step of fidelity to the fixed\-input deletion counterfactual\. The reference is the original weighted MoE output, which we call the*consensus*without assuming expert agreement\. RCS measures each expert’s routing\-weighted deviation from this mixture\. RCS\-LOO accounts for survivor renormalization, and RCS\-Refill additionally includes router\-selected replacement, the criterion we callRazor\. The latter two recover exact single\-deletion output changes under their respective fixed\-input assumptions\. Comparing them tests whether each refinement improves the retained checkpoint\. We aggregate these scores over calibration tokens and retain the highest\-scoring experts within each layer’s budget, using only forward computation, without gradients, subset search, or recovery training\. Local change remains a surrogate for the model\-level objective because joint deletions interact and altered activations propagate across layers\.

Our benchmark suite scores multi\-step trajectories in competition mathematics, program synthesis, software engineering, and long\-context inference, so it tests reasoning ability rather than isolated recall\. On GLM\-4\.7\-Flash, Qwen3\.6\-35B\-A3B, DeepSeek\-V4\-Flash\-0731, and Hy3 at 25% and 50% expert removal,Razorhas the highest nine\-task macro average among the evaluated pruning methods in all eight settings, exceeding REAP by 2\.12–5\.59 points on the two backbones where REAP was also benchmarked\. On matched GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B checkpoints, it also lowers reverse KL relative to REAP in all four model–budget settings, so the replaceability signal improves the retained checkpoint under both views\.

Two findings bound how far that improvement goes\. Pruning this aggressively is not lossless, since all four backbones fall below their originals in macro score at 50% removal\. More importantly, the gains do not reach generation behavior\. Response analyses on Qwen3\.6\-35B\-A3B show that diversity does not follow the task ranking, and thatRazorat 50% removal emits fewer stray closing delimiters than REAP while reaching the output cap more often than the unpruned model\. Retained accuracy alone therefore does not certify a pruned reasoning model, which is why we report these properties rather than the benchmark suite alone\.

## 2Methodology

### 2\.1Pruning objective and routed mixture

We retainBBofEErouted experts per layer, withk≤B<Ek\\leq B<E, without recovery training\. Preserving the original output distribution at this budget is the goal, and local Euclidean output change is our surrogate for it, not a global guarantee\. We fix notation for that surrogate below and then show that output magnitude cannot supply it\.

For token representationxx, the router selectskkexpertsS⁡\(x\)S\(x\)\. Expertiiproducesfi​\(x\)f\_\{i\}\(x\)with normalized weightwi​\(x\)≥0w\_\{i\}\(x\)\\geq 0, where∑i∈S⁡\(x\)wi​\(x\)=1\\sum\_\{i\\in S\(x\)\}w\_\{i\}\(x\)=1\. The routed output is

c⁡\(x\)=∑i∈S⁡\(x\)wi​\(x\)​fi​\(x\)\.c\(x\)=\\sum\_\{i\\in S\(x\)\}w\_\{i\}\(x\)f\_\{i\}\(x\)\.\(1\)We call this mixture the*consensus*, without assuming that its experts agree\. Unchanged residual\-connection, dense, and shared\-expert branches cancel in the local comparison\. A common routed\-output scale does not change within\-layer rankings \(Appendix[B\.3](https://arxiv.org/html/2609.30465#A2.SS3)\)\.

The statisticwi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}measures the magnitude of expertii’s weighted output contribution\. Deletion removes that contribution and renormalizes the survivors, so its local effect depends on information that even the complete set of gates and output norms cannot supply\.

###### Proposition 1\(Magnitude does not identify the least harmful deletion\)\.

Even with equal known gates, the expert\-output norms do not in general determine which single\-expert deletion minimizes the local output shift after survivor renormalization without refill\.

###### Proof\.

Consider equally weighted scalar outputs\(8,3,10\)\(8,3,10\)and\(8,3,−10\)\(8,3,\-10\)\. Both have the same output norms and gates\. Their original mixtures are77and1/31/3\. The absolute shifts after deleting the first, second, or third expert are\(1/2,2,3/2\)\(1/2,2,3/2\)in the first case and\(23/6,4/3,31/6\)\(23/6,4/3,31/6\)in the second\. The unique least harmful deletion is therefore the first expert in one case and the second in the other\. A rule observing only the gates and norms cannot distinguish them\. ∎

Reweighting output norms cannot recover the missing directional information\. Givenfif\_\{i\}andwiw\_\{i\}, however, the shared vectorccdetermines the survivor output\.

### 2\.2From consensus residual to exact deletion effect

To isolate one deletion, we hold the layer input and expert outputs fixed and remove a selected expert withwi​\(x\)<1w\_\{i\}\(x\)<1\. We first model router refill, then recover deletion without replacement as a special case\.

###### Assumption 1\(Refill\)\.

Deletingi∈S⁡\(x\)i\\in S\(x\)promotes the highest\-ranked unselected expertrrunder the router’s selection rule\. Its nonnegative mixture score, divided by the original selected\-score sum, defines the pseudo\-weightwr​\(x\)w\_\{r\}\(x\)\. The surviving scores and this promoted score are then renormalized over their combined mass1−wi​\(x\)\+wr​\(x\)1\-w\_\{i\}\(x\)\+w\_\{r\}\(x\)\.

###### Assumption 2\(Fixed support\)\.

The surviving weights are renormalized overS⁡\(x\)∖\{i\}S\(x\)\\setminus\\\{i\\\}without promoting an unselected expert\. Equivalently,wr​\(x\)≡0w\_\{r\}\(x\)\\equiv 0in Assumption[1](https://arxiv.org/html/2609.30465#Thmassumption1)\.

We writec~−i\\tilde\{c\}^\{\-i\}for the refilled mixture andc−ic^\{\-i\}for the fixed\-support one\. The first uses the routed outputs and the promoted expert’s output at the same layer input, whereas the second needs only the routed outputs\. Neither requires evaluating a separate end\-to\-end pruned model for each expert\.

###### Proposition 2\(Single\-expert deletion under refill\)\.

Under Assumption[1](https://arxiv.org/html/2609.30465#Thmassumption1), the token\-level deletion damage is

δi​\(x\)≡‖c⁡\(x\)−c~−i​\(x\)‖2=‖wi​\(x\)​ri​\(x\)−wr​\(x\)​rr​\(x\)‖2Di​\(x\),\\delta\_\{i\}\(x\)\\equiv\\big\\\|c\(x\)\-\\tilde\{c\}^\{\-i\}\(x\)\\big\\\|\_\{2\}=\\frac\{\\big\\\|w\_\{i\}\(x\)r\_\{i\}\(x\)\-w\_\{r\}\(x\)r\_\{r\}\(x\)\\big\\\|\_\{2\}\}\{D\_\{i\}\(x\)\},\(2\)whererj​\(x\)=fj​\(x\)−c⁡\(x\)r\_\{j\}\(x\)=f\_\{j\}\(x\)\-c\(x\)is the consensus residual of expertjj\.

###### Proof\.

Suppressingxx, the promoted slot carries masswrw\_\{r\}while the surviving selected mass is1−wi1\-w\_\{i\}, soc~−i=\(c−wi​fi\+wr​fr\)/Di\\tilde\{c\}^\{\-i\}=\(c\-w\_\{i\}f\_\{i\}\+w\_\{r\}f\_\{r\}\)/D\_\{i\}\. UsingDi−1=−wi\+wrD\_\{i\}\-1=\-w\_\{i\}\+w\_\{r\},

c−c~−i=Di​c−c\+wi​fi−wr​frDi=wi​\(fi−c\)−wr​\(fr−c\)Di\.c\-\\tilde\{c\}^\{\-i\}=\\frac\{D\_\{i\}c\-c\+w\_\{i\}f\_\{i\}\-w\_\{r\}f\_\{r\}\}\{D\_\{i\}\}=\\frac\{w\_\{i\}\(f\_\{i\}\-c\)\-w\_\{r\}\(f\_\{r\}\-c\)\}\{D\_\{i\}\}\.Taking norms yields the result\. ∎

Settingwr=0w\_\{r\}=0recovers the fixed\-support case in closed form\.

###### Proposition 3\(Single\-expert leave\-one\-out impact\)\.

Under Assumption[2](https://arxiv.org/html/2609.30465#Thmassumption2), the token\-level deletion damage is

δiloo​\(x\)≡‖c⁡\(x\)−c−i​\(x\)‖2=wi​\(x\)1−wi​\(x\)​‖fi​\(x\)−c⁡\(x\)‖2\.\\delta\_\{i\}^\{\\mathrm\{loo\}\}\(x\)\\equiv\\big\\\|c\(x\)\-c^\{\-i\}\(x\)\\big\\\|\_\{2\}=\\frac\{w\_\{i\}\(x\)\}\{1\-w\_\{i\}\(x\)\}\\,\\\|f\_\{i\}\(x\)\-c\(x\)\\\|\_\{2\}\.\(3\)

###### Proof\.

Settingwr=0w\_\{r\}=0in Proposition[2](https://arxiv.org/html/2609.30465#Thmproposition2)gives denominator1−wi1\-w\_\{i\}and numeratorwi​‖fi−c‖2w\_\{i\}\\\|f\_\{i\}\-c\\\|\_\{2\}\. ∎

These identities give replaceability a concrete form\. Fixed\-support damage is equivalentlywi​‖fi−c−i‖2w\_\{i\}\\\|f\_\{i\}\-c^\{\-i\}\\\|\_\{2\}, so distance to the survivor output measures how far the remaining computation is from reproducing the removed expert, and routing mass converts that distance into mixture change\. Small damage is therefore ambiguous by construction, arising from either a small distance or a small routing mass\. The LOO factor1/\(1−wi\)1/\(1\-w\_\{i\}\)corrects for self\-inclusion inccand follows from the deletion counterfactual rather than serving as a tunable importance weight \(Appendix[A\.2](https://arxiv.org/html/2609.30465#A1.SS2)\)\. Refill then adds an opposing weighted residual and changes the normalizer, so the promoted expert can compensate for the removed contribution, though by an amount set by residual magnitudes and routing weights as well as direction\. An exact matchwi​ri=wr​rrw\_\{i\}r\_\{i\}=w\_\{r\}r\_\{r\}leaves the local mixture unchanged\.

Exact single\-deletion effects are not additive\. Without refill, removing a routed subset𝒜⊆S⁡\(x\)\\mathcal\{A\}\\subseteq S\(x\)with total weightW𝒜=∑i∈𝒜wi<1W\_\{\\mathcal\{A\}\}=\\sum\_\{i\\in\\mathcal\{A\}\}w\_\{i\}<1gives

c−c−𝒜=∑i∈𝒜wi​\(fi−c\)1−W𝒜\.c\-c^\{\-\\mathcal\{A\}\}=\\frac\{\\sum\_\{i\\in\\mathcal\{A\}\}w\_\{i\}\(f\_\{i\}\-c\)\}\{1\-W\_\{\\mathcal\{A\}\}\}\.\(4\)Residuals can reinforce or cancel, while removed mass changes the denominator \(Appendix[A\.3](https://arxiv.org/html/2609.30465#A1.SS3)\)\. Scalar ranking does not optimize this joint effect, and refill does not restore additivity\. Promoting several ranks at once couples the removed experts through both the numerator and the set\-dependent normalizer\. IfW𝒜=1W\_\{\\mathcal\{A\}\}=1, no routing mass survives and the fixed\-support counterfactual is undefined, even ifB≥kB\\geq kexperts remain in the layer\. Deployment reselects top\-kkfrom the retained pool and propagates changed activations across layers, so matching its refill rule at a single token still does not guarantee checkpoint fidelity\.

### 2\.3Expert scoring and pruning

A criterion becomes a pruning rule once the token\-level quantity is reduced to one score per expert\. RCS useswi​‖ri‖2w\_\{i\}\\\|r\_\{i\}\\\|\_\{2\}, RCS\-LOO usesδiloo\\delta\_\{i\}^\{\\mathrm\{loo\}\}, and RCS\-Refill usesδi\\delta\_\{i\}\. Unless explicitly varied, all three aggregate their token quantities by root mean square \(RMS\) over calibration tokens routed to each expert, which weights variable damage above steady damage of the same average size\.*We writeRazorfor RCS\-Refill under this aggregation*, the instance evaluated throughout unless another is named\. Because the reduction does not follow from the deletion identity, Appendices[C\.2](https://arxiv.org/html/2609.30465#A3.SS2)and[C\.3](https://arxiv.org/html/2609.30465#A3.SS3)treat it as a separate choice and test it directly\. Table[5](https://arxiv.org/html/2609.30465#A3.T5)maps the criteria to their experimental coverage\.

We retain the highest\-scoring experts within each layer’s budget and deploy top\-kkrouting over the retained pools\. Appendix[A\.4](https://arxiv.org/html/2609.30465#A1.SS4)gives the implementation and pseudocode\.

Figure 2:Domain\-wise reverse KL at 25% \(top\) and 50% \(bottom\) removal\.Values are in nats, with lower \(inward\) better\. Compare methods along each spoke, not by polygon area\. Whiskers show pointwise 95% paired\-batch intervals\. Shading distinguishes ID/OOD domains\.

## 3Experiments

Four research questions narrow from what a pruned model predicts to why the scoring choices behind it matter\.

1. RQ1\.How well does pruning preserve output distributions across domains and routing conditions?
2. RQ2\.Does that preserved distribution translate into retained downstream performance across four backbones and two removal budgets?
3. RQ3\.Which response properties shift even where reasoning accuracy is retained?
4. RQ4\.Do refinements to the local deletion score improve pruning, and how do scoring and aggregation choices affect expert selection and deletion\-set interactions?

Each answer below qualifies the previous one rather than confirming it\.

### 3\.1Experimental Setup

##### Models and pruning methods\.

We compare RCS, RCS\-LOO, andRazoron GLM\-4\.7\-Flash, Qwen3\.6\-35B\-A3B, DeepSeek\-V4\-Flash\-0731, and Hy3 at 25% and 50% expert removal, without recovery training\. All three use conditional RMS unless an aggregation ablation is specified, and each retains the highest\-scoring experts per layer\. Matched checkpoint and component studies use GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B, whose routers select 4/64 and 8/256 experts, respectively\.

##### Baselines\.

We compare against routing frequency \(Frequency\), a conditional\-mean adaptation of expert activation norm \(EAN\)\([Jaiswal et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib12)\), and REAP\([Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15)\), which averages routing\-weighted output norms conditional on use\. Downstream benchmark comparisons against these three baselines use GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B, while separate mask\-based domain\-wise and routing\-stratified diagnostics compare against REAP on all four backbones \(Appendices[E\.2](https://arxiv.org/html/2609.30465#A5.SS2)and[E\.3](https://arxiv.org/html/2609.30465#A5.SS3)\)\. REAP⋆denotes an external checkpoint calibrated with 24,576 samples, distinct from the matched REAP baseline\.

Table 1:Nine\-task benchmark results on GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B at 25% and 50% removal\. Scores are scaled by 100\. Bold marks the best pruned score in each column within a model–budget block\.ModelMethodMathInstruction FollowingKnow\.ToolLCCodingOverallAvgAIME’26IFEvalIFBenchAvgSuperGPQABFCL v4LongB\. v2HE\+LCBSWEAvg\[\]\[\]0%—31\.779\.740\.059\.938\.565\.325\.379\.937\.45\.240\.844\.8\[\]\[\]Frequency26\.377\.337\.357\.329\.562\.821\.974\.429\.35\.236\.340\.4\[\]\[\]EAN30\.476\.339\.357\.831\.564\.727\.275\.033\.84\.837\.942\.6\[\]\[\]REAP⋆30\.074\.936\.755\.832\.664\.927\.275\.030\.63\.636\.441\.7\[\]\[\]REAP30\.874\.937\.356\.133\.765\.027\.075\.034\.33\.637\.642\.4\[\]\[\]RCS29\.677\.136\.056\.633\.365\.726\.073\.233\.84\.237\.142\.1\[\]\[\]RCS\-LOO31\.376\.739\.758\.234\.364\.827\.475\.634\.84\.238\.243\.2\[\]\[\]25%Razor33\.378\.240\.359\.334\.265\.628\.576\.239\.84\.640\.244\.5\[\]\[\]Frequency0\.062\.830\.346\.615\.751\.118\.754\.917\.80\.224\.327\.9\[\]\[\]EAN15\.870\.629\.049\.825\.454\.823\.759\.233\.92\.631\.935\.0\[\]\[\]REAP23\.361\.025\.343\.227\.856\.022\.167\.126\.71\.231\.734\.5\[\]\[\]RCS29\.663\.827\.345\.624\.959\.119\.571\.327\.63\.034\.036\.2\[\]\[\]RCS\-LOO26\.367\.530\.749\.129\.058\.620\.972\.029\.43\.234\.937\.5\[\]\[\]GLM\-4\.7\-Flash50%Razor30\.872\.032\.352\.233\.759\.125\.973\.229\.84\.035\.740\.1\[\]\[\]0%—76\.383\.933\.758\.862\.867\.650\.191\.571\.917\.060\.161\.6\[\]\[\]Frequency75\.079\.134\.356\.751\.467\.245\.990\.970\.112\.057\.758\.4\[\]\[\]EAN75\.481\.732\.357\.057\.465\.848\.792\.764\.816\.658\.059\.5\[\]\[\]REAP77\.179\.533\.756\.657\.065\.848\.890\.970\.914\.658\.859\.8\[\]\[\]RCS76\.780\.038\.059\.057\.366\.348\.592\.171\.918\.060\.761\.0\[\]\[\]RCS\-LOO76\.380\.634\.757\.757\.166\.449\.792\.772\.617\.460\.960\.8\[\]\[\]25%Razor77\.581\.938\.360\.157\.867\.049\.993\.372\.620\.262\.062\.1\[\]\[\]Frequency68\.871\.031\.751\.441\.143\.142\.589\.655\.07\.250\.650\.0\[\]\[\]EAN72\.174\.130\.352\.246\.158\.746\.188\.447\.18\.448\.052\.4\[\]\[\]REAP71\.773\.231\.052\.146\.363\.947\.588\.466\.412\.655\.855\.7\[\]\[\]RCS73\.374\.336\.355\.348\.364\.148\.391\.567\.612\.857\.357\.4\[\]\[\]RCS\-LOO75\.478\.232\.055\.147\.064\.547\.992\.167\.913\.657\.957\.6\[\]\[\]Qwen3\.6\-35B\-A3B50%Razor75\.079\.334\.056\.749\.265\.048\.392\.169\.415\.859\.158\.7

##### Data and metrics\.

Calibration uses the 2,048\-example, seven\-domainRazorCalpool with model\-specific chat templates\. Predictive fidelity is reverse KL,DKL\(q∥p\)D\_\{\\mathrm\{KL\}\}\(q\\\|p\), from original \(pp\) and pruned \(qq\) predictions on shared reference prefixes\. Assistant\-token NLL givesΔ​NLL=NLLq−NLLp\\Delta\\mathrm\{NLL\}=\\mathrm\{NLL\}\_\{q\}\-\\mathrm\{NLL\}\_\{p\}and relative excess PPL,exp⁡\(Δ​NLL\)−1\\exp\(\\Delta\\mathrm\{NLL\}\)\-1\. Fidelity is reported over eight axes, four covered byRazorCal\(mathematics, code, instruction following, and tools\) and four held out from it \(chat, creative writing, safety, and SQL\), so calibration\-covered and held\-out performance stay separable\. The matched component sweep pools tokens within axes and weights axes equally, whereas domain\-wise and routing\-stratified studies use grouped estimators\.

The nine downstream tasks are AIME’26\([Mathematical Association of America, n\.d\.](https://arxiv.org/html/2609.30465#bib.bib28)\), IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2609.30465#bib.bib44)\), IFBench\([Pyatkin et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib33)\), SuperGPQA\([Du et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib5)\), BFCL v4\([Patil et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib32);[Mao et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib27)\), LongBench v2\([Bai et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib1)\), HumanEval\+\([Liu et al\., 2023](https://arxiv.org/html/2609.30465#bib.bib23)\), LiveCodeBench \(2026 latest\)\([Jain et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib11)\), and SWE\-bench Verified\([Jimenez et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib14);[Chowdhury et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib3)\)\. Each scores a derivation or a trajectory rather than a single prediction, spanning competition mathematics, coding from function level to repository level, long\-context inference, graduate\-level knowledge, verifiable constraint following, and tool use\. Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)details task versions and score aggregation\. Response analyses measure diversity, stray closing thinking delimiters, code\-fence structure, length, and length\-limit finishes, which are observable properties of the generated text rather than measures of reasoning quality, and qualify the task results rather than replacing them\.

##### Hyperparameters and evaluation protocol\.

Every scoring\-rule comparison matches calibration inputs and layerwise budgets\. For downstream benchmarks, we use temperature00, except for SWE\-bench Verified, where we use0\.70\.7\. AIME’26 reports avg@8, and the other eight benchmarks report avg@3\. Diversity and response\-form diagnostics follow the separate protocol of Section[3\.4](https://arxiv.org/html/2609.30465#S3.SS4)\. Context and output limits, the thinking setting, and cross\-study matching are detailed in Appendices[B\.1](https://arxiv.org/html/2609.30465#A2.SS1)–[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)\. Diagnostic sampling and uncertainty are described in Appendices[E\.2](https://arxiv.org/html/2609.30465#A5.SS2)–[E\.6](https://arxiv.org/html/2609.30465#A5.SS6)\.

### 3\.2Predictive Fidelity \(RQ1\)

The output distribution is what our criterion targets directly, so it is where a replaceability signal should show its clearest advantage\. We test that at three resolutions\. Figure[2](https://arxiv.org/html/2609.30465#S2.F2)compares REAP and the three residual criteria per domain, Table[6](https://arxiv.org/html/2609.30465#A3.T6)gives matched equal\-axis GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B estimates, and Figure[8](https://arxiv.org/html/2609.30465#A5.F8)resolves the same comparison by routing concentration under both fidelity measures, separating calibration\-covered from held\-out domains\. All scored criteria use conditional RMS throughout, so the contrasts isolate the token quantity\. Lower reverse KL indicates less drift, and radar spokes are scaled independently from mostly nonzero origins, so methods are comparable on the same spoke and not by polygon area\.

Obs 1\.Razorreduces predictive drift, but KL and likelihood need not agree\.Across the four matched GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B settings,Razorhas lower reverse KL than REAP in all four and than RCS\-LOO in three \(Table[6](https://arxiv.org/html/2609.30465#A3.T6)\)\. A separate routing\-stratified comparison across all four backbones illustrates the distinction between fidelity measures\. At 50% removal,Razorlowers KL relative to REAP in 73 of 80 model–domain\-group–decile estimates, but keeps PPL closer to the original in only 47\. The corresponding RCS\-LOO counts are 72 and 47 \(Appendix[E\.4](https://arxiv.org/html/2609.30465#A5.SS4)\)\. Reduced distributional drift therefore need not imply better preservation of reference\-token likelihood\.

Table 2:Nine\-task benchmark results on DeepSeek\-V4\-Flash\-0731 and Hy3 at 25% and 50% removal\. Conventions follow Table[1](https://arxiv.org/html/2609.30465#S3.T1)\.Obs 2\. KL improvements are not uniform across domains\.Of eight domains, the counts whereRazorlowers KL relative to REAP at 25% and 50% removal are, respectively, 5 and 7 for GLM\-4\.7\-Flash, 6 and 5 for Qwen3\.6\-35B\-A3B, 6 and 8 for DeepSeek\-V4\-Flash\-0731, and 6 and 7 for Hy3\. Yet Qwen3\.6\-35B\-A3B’s equal\-domain held\-out mean at 50% is 0\.2% higher than REAP’s, and the relative\-change interval includes zero\. Aggregate gains thus coexist with domain\-level regressions\. Appendix[E\.2](https://arxiv.org/html/2609.30465#A5.SS2)gives the corresponding RCS\-LOO results\. These domain\-wise and routing\-stratified views reweight overlapping predictions rather than provide independent replications\.

Figure 3:Qwen3\.6\-35B\-A3B response diversity and response form\.The left panel shows Distinct\-nnchanges from Original\. The right panels show stray</think\>per 10,000 LiveCodeBench completions, median token length, and loop share among length\-limited completions\. Dashed lines mark Original, with 95% bands in the form panels\. Whiskers show pointwise 95% bootstrap intervals\.
### 3\.3Downstream Performance \(RQ2\)

Lower distributional drift is only useful if it survives decoding into multi\-step trajectories, which the fidelity measures of RQ1 do not test\. Tables[1](https://arxiv.org/html/2609.30465#S3.T1)and[2](https://arxiv.org/html/2609.30465#S3.T2)therefore report the nine\-task suite for all four backbones, with unpruned references at 0%\. Overall Avg weights all nine tasks equally, while the twoAvgcolumns average instruction\-following and coding tasks\. Macro gaps quoted below use unrounded task\-score averages\.

Obs 3\.Razorleads the macro comparison across all eight model–budget settings\.Among the evaluated pruning methods,Razorhas the highest macro score in every setting, exceeding RCS\-LOO by 0\.74–2\.58 points across all eight\. Against RCS\-LOO, it wins 61 of 72 paired tasks, ties two and loses nine, so the macro advantage does not imply a uniform taskwise ordering\. The REAP comparison spans GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B across four model–budget settings, whereRazorexceeds REAP by 2\.12–5\.59 points and wins all 36 paired task comparisons\. In comparison, RCS\-LOO exceeds REAP by 0\.80–3\.01 points while losing three of the same 36 tasks\. These results favor refill for macro task utility within the residual family\. Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)gives the per\-task breakdown and the scope of each comparison\.

Obs 4\. Near\-original performance at 25% gives way to consistent losses at 50%\.The 25%Razorcheckpoints remain close to or above the originals in macro score, whereas 50% removal is lossy on all four backbones\. For example, GLM\-4\.7\-Flash falls from 44\.5 at 25% to 40\.1 at 50%, against 44\.8 unpruned\. Qwen3\.6\-35B\-A3B falls from 62\.1 to 58\.7, against 61\.6\. These small gains above the original macro scores do not establish that pruning improves the original models\.

Figure 4:RMS aggregation and expert selection forRazor\.Left to right, panels show relative KL change from mean \(negative favors RMS\), expert CV, RMS\-only retained experts per layer, and deletion\-set Gram ratios\. Shading marks 50% removal, used in the last three panels\.
### 3\.4Generation Behavior \(RQ3\)

A benchmark score records whether a reasoning trajectory reached the right answer, not how it was written or when it stopped, so the gains of RQ2 leave generation behavior untested\. Four properties of Qwen3\.6\-35B\-A3B generations make that behavior observable without judging content \(Figure[3](https://arxiv.org/html/2609.30465#S3.F3)\)\. Diversity compares all four criteria by Distinct\-nnchange from Original \(n=2,3,4n=2,3,4\) over 16 fixed responses per question, weighting AIME’26, HumanEval\+, and LiveCodeBench equally\. Response form compares REAP andRazoron LiveCodeBench over 256 shared sample positions per question \(44,800 completions per variant\)\. Sampling uses temperature0\.70\.7, top\-p=0\.95p=0\.95, top\-k=20k=20, an 8,192\-token cap, and thinking disabled, so these results probe response behavior rather than explicit reasoning traces\. Diversity intervals use paired\-question resampling and form intervals resample questions within each variant, both conditioning on the fixed generations \(Appendices[E\.5](https://arxiv.org/html/2609.30465#A5.SS5.SSS0.Px1)and[E\.6](https://arxiv.org/html/2609.30465#A5.SS6)\)\.

Obs 5\. Higher diversity does not identify the best pruning criterion\.RCS\-LOO has the highest Distinct\-nnin five of the six combinations of removal budget and n\-gram order and exceeds REAP in all six, with a median gap of 3\.71 percentage points in relative change from Original, despiteRazor’s stronger macro task performance\. The exception is Distinct\-2 at 25%, whereRazorchanges by\+1\.73%\+1\.73\\%from Original versus\+1\.53%\+1\.53\\%for RCS\-LOO\. Changes are smaller at 25% than at 50%\. In a separate four\-task analysis, Qwen3\.6\-35B\-A3B’s RCS\-LOO–REAP Distinct\-4 gap reverses under a 128\-token prefix control that also changes question eligibility, so the reversal cannot be attributed to length alone \(Appendix[E\.5](https://arxiv.org/html/2609.30465#A5.SS5)\)\. Higher diversity therefore establishes neither correctness nor a ranking robust to this joint control\.

Obs 6\. Closer response form does not imply preserved termination rates\.At 50% removal, stray</think\>emissions fall from REAP’s 13\.4 toRazor’s 8\.5 per 10,000 completions, whereas RCS\-LOO reaches 68\.5 \(Table[10](https://arxiv.org/html/2609.30465#A5.T10)\)\. Median token length is also closer to Original, at 2\.37k forRazorversus 2\.43k unpruned and 3\.94k for REAP\. These response\-form comparisons draw on a different sample pool from the diversity result above and are not adjusted for length\.

Termination is the one property where closer form buys nothing, since length\-limited finishes reach 25\.8% forRazorand 24\.2% for REAP against 16\.0% unpruned\. Degenerate repetition does not explain those caps\. Among length\-limited completions, loop shares under the rightmost panel’s criterion \(fewer than one quarter of token 4\-grams unique\) are 0\.27%/0\.43% forRazorand 0\.30%/0\.91% for REAP at 25%/50% removal, so loops account for under 1% of capped responses and the remaining failures stay unexplained \(Appendix[E\.6](https://arxiv.org/html/2609.30465#A5.SS6)\)\.

### 3\.5Discussion \(RQ4\)

The preceding results compare finished checkpoints without showing which scoring choice produced them\. We therefore ask what each choice contributes, first to checkpoint fidelity and then to the retained set itself\.

#### 3\.5\.1Scoring components

Our criteria form a ladder in which each rung represents more of one local deletion counterfactual, so the ladder is testable\. RCS scores the routing\-weighted consensus residual, RCS\-LOO adds survivor renormalization, andRazoradds router\-selected replacement\. Table[6](https://arxiv.org/html/2609.30465#A3.T6)compares them against REAP on matched GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B checkpoints at both budgets, and the component contrasts of Appendix[C\.6](https://arxiv.org/html/2609.30465#A3.SS6)hold the remaining choices fixed\. The factorial study \(Appendix[C\.5](https://arxiv.org/html/2609.30465#A3.SS5)\) comes from a separate evaluation collection and is therefore read on its own\.

Obs 7\. Consensus subtraction helps, while counterfactual refinements have mixed effects\.Consensus subtraction lowers KL in all four model–budget settings when routing weights and mean aggregation are held fixed\. This matched contrast isolates the reference change, unlike REAP\-to\-RCS, which also changes aggregation\. Subsequent refinements are model\-dependent\. Under conditional RMS, adding the LOO factor lowers KL by 0\.6%/5\.6% on GLM\-4\.7\-Flash but raises it by 7\.1%/11\.8% on Qwen3\.6\-35B\-A3B at 25%/50% removal\. Refill then improves on RCS\-LOO in three of four settings, with GLM\-4\.7\-Flash at 25% the exception \(\+0\.5%\+0\.5\\%KL\)\. Thus a more complete single\-deletion counterfactual need not produce a better jointly pruned checkpoint\. The choice ofRazorrests on the benchmark comparisons as well as these fidelity results, rather than local exactness alone\.

#### 3\.5\.2Expert selection

Fidelity differences between scoring rules say nothing about which experts those rules actually keep\. Figure[4](https://arxiv.org/html/2609.30465#S3.F4)closes that gap forRazorby tracing one aggregation choice from its fidelity effect to its selection effect, through RMS\-versus\-mean KL across four budgets, per\-expert coefficient of variation \(CV\), per\-layer keep\-set swaps, and deletion\-set Gram ratios, with Table[7](https://arxiv.org/html/2609.30465#A3.T7)placingRazorand RCS\-LOO statistics side by side\. Hollow markers in the KL panel indicate disagreement in sign across its four evaluation shards, and the latter three panels use 50% removal\.

Obs 8\. RMS changes retained sets, with model\-dependent fidelity benefits\.RMS yields lower KL than mean aggregation at all four budgets on GLM\-4\.7\-Flash, but higher KL above 25% removal on Qwen3\.6\-35B\-A3B\. This differs from RCS\-LOO, whose RMS\-over\-mean gains are small on Qwen3\.6\-35B\-A3B \(Appendix[C\.6](https://arxiv.org/html/2609.30465#A3.SS6)\)\. Variation in the RMS premium1\+CV2\\sqrt\{1\+\\mathrm\{CV\}^\{2\}\}reweights experts rather than uniformly rescaling them\. At 50% removal, RMS retains an average of 11\.0 experts per layer that mean would discard on Qwen3\.6\-35B\-A3B and 1\.85 on GLM\-4\.7\-Flash, from retention budgets of 128 and 32, respectively\. The per\-layer ranges are 5–30 and 0–7\. This establishes changed selection, not that the swaps cause benchmark gains\. Appendix[C\.3](https://arxiv.org/html/2609.30465#A3.SS3)gives further aggregation contrasts\.

Obs 9\. Local diagnostics expose interactions that scalar rankings do not resolve\.The Gram\-ratio panel shows net reinforcement in four GLM\-4\.7\-Flash layers forRazor, whereas RCS\-LOO\-selected sets show net cancellation in all 46 GLM\-4\.7\-Flash layers and 29 of 40 Qwen3\.6\-35B\-A3B layers \(Appendix[C\.7\.2](https://arxiv.org/html/2609.30465#A3.SS7.SSS2)\)\. These interactions concern the fixed\-support numerator, constrained by∑iwi​\(fi−c\)=0\\sum\_\{i\}w\_\{i\}\(f\_\{i\}\-c\)=0\. The denominator and refill remain outside this probe\. Single\-deletion rankings also show no consistent advantage over REAP, despite favorable RCS\-LOO comparisons in concentrated\-routing bins \(Appendices[C\.7\.1](https://arxiv.org/html/2609.30465#A3.SS7.SSS1)and[C\.7\.3](https://arxiv.org/html/2609.30465#A3.SS7.SSS3)\)\. Thus the probes expose interactions without establishing optimal joint selection or the cause of the benchmark gains\.

## 4Related Work

MoE compression methods differ in the removal unit, the importance signal, and the selection procedure applied afterwards\.Razorvaries only the second, removing whole experts under a fixed layerwise budget with retained weights and routers untouched\. The closest work is therefore what scores whole experts independently, what models relations among them, and what searches over retained sets, ordered below by distance from that setting\. Appendix[F](https://arxiv.org/html/2609.30465#A6)treats the families that change the removal unit or the surviving computation, including merging, router adaptation, budget allocation, and finer\-grained compression\.

##### Expert scoring\.

At the individual\-expert level, pruning methods use routing, activation, or weight statistics\([Muzio et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib30);[Jaiswal et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib12);[Liu et al\., 2026a](https://arxiv.org/html/2609.30465#bib.bib24)\)\. Deletion\-aware methods also model the computation after removal\. REAP includes promoted substitution and survivor renormalization but retains only the removed expert’s routing\-weighted output norm\([Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15)\), while a unified framework derives exact damage yet omits the rerouting residual from practical scores\([Liu et al\., 2026b](https://arxiv.org/html/2609.30465#bib.bib25)\)\.Razorinstead scores the full fixed\-input consensus\-relative deletion–refill change\.

##### Expert relations\.

Independent scores can miss structure shared across experts\. STUN and HC\-SMoE cluster router behavior or expert outputs, while SHAPE, ConMoE, and MAESTRO use routing co\-occurrence, parameter\-space replaceability, or cross\-layer transitions\([Lee et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib16);[Chen et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib2);[Zhang, 2026](https://arxiv.org/html/2609.30465#bib.bib42);[Yao et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib41);[Goel et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib8)\)\. These relations support selection, merging, or path modeling, whereasRazorscores deletion with survivor renormalization and refill, without recovery\.

##### Candidate\-set search\.

Beyond relational ranking, other methods evaluate retained sets through reconstruction, joint pruning–merging, or layerwise and blockwise search\([Lu et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib26);[Liu et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib22);[Yang et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib40)\)\.Razorperforms no subset search\. Its exactness for one fixed\-input deletion does not solve joint pruning, because deletions interact and a refill candidate may itself be removed\.

## 5Conclusion

We presentedRazorfor training\-free expert pruning of reasoning MoE models\. Its consensus residuals score functional replaceability under survivor renormalization and router refill\. Across four post\-trained MoE models and two removal budgets, it achieves the highest nine\-task macro average among the evaluated methods in all eight settings and lowers reverse KL relative to REAP in all four matched GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B comparisons\. Scoring replaceability rather than contribution therefore retains more reasoning ability at the same budget, and it does so from forward passes alone, which puts the gain within reach of any pipeline that can already run calibration\. What our results do not establish is that a more faithful local counterfactual is always the better ranking, since the refinements that improve benchmark scores do not consistently improve checkpoint fidelity, and selection under these criteria is less stable across calibration budgets than under REAP\. Response diversity, formatting, and termination also shift independently of the reported gains, so retained accuracy alone does not certify a pruned reasoning model\.

## AI use statement

Generative AI tools assisted with the refinement of academic prose, focusing on clarity, coherence, and precision of expression\. They also supported literature search and coding\.

## Ethics statement

This work studies expert pruning of post\-trained MoE language models, aiming to reduce model storage requirements while retaining their capabilities\. Removing experts can also alter behaviors shaped by post\-training, including safety\-related responses, even when aggregate task performance is preserved\. Agreement with the original model on safety\-related inputs measures predictive fidelity, but does not establish safety or preservation of alignment\. We therefore emphasize independent safety and alignment evaluation before deployment, particularly in safety\-sensitive applications\. The study uses previously released calibration data \(Appendix[B\.1](https://arxiv.org/html/2609.30465#A2.SS1)\) and involves no human\-subject research\.

## Reproducibility statement

Section[2](https://arxiv.org/html/2609.30465#S2)and Appendix[A](https://arxiv.org/html/2609.30465#A1)provide the scoring criterion and derivations\. Appendix[A\.4](https://arxiv.org/html/2609.30465#A1.SS4)gives the algorithm and implementation details\. Appendix[B](https://arxiv.org/html/2609.30465#A2)documents calibration data, model configurations, evaluation protocols, score aggregation, and uncertainty estimates\. Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)states the evaluator settings and the scope of each cross\-study comparison\. All backbones and calibration sources are publicly released, so the reported checkpoints can be reconstructed from the scoring criterion and layerwise budgets given here\.

## References

- Bai et al\. \(2025\)Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Yuxiao Dong, Jie Tang, Lei Hou, and Juanzi Li\.LongBench v2: Towards deeper understanding and reasoning on realistic long\-context multitasks\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 3639–3664\. Association for Computational Linguistics, 2025\.doi:10\.18653/v1/2025\.acl\-long\.183\.URL[https://aclanthology\.org/2025\.acl\-long\.183/](https://aclanthology.org/2025.acl-long.183/)\.
- Chen et al\. \(2025\)I\-Chun Chen, Hsu\-Shen Liu, Wei\-Fang Sun, Chen\-Hao Chao, Yen\-Chang Hsu, and Chun\-Yi Lee\.Retraining\-free merging of sparse MoE via hierarchical clustering\.In*Proceedings of the 42nd International Conference on Machine Learning*, pp\. 8594–8620\. PMLR, 2025\.URL[https://proceedings\.mlr\.press/v267/chen25aq\.html](https://proceedings.mlr.press/v267/chen25aq.html)\.
- Chowdhury et al\. \(2024\)Neil Chowdhury, James Aung, Chan Jun Shern, Oliver Jaffe, Dane Sherburn, Giulio Starace, Evan Mays, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E\. Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry\.Introducing SWE\-bench Verified\.OpenAI, August 2024\.URL[https://openai\.com/index/introducing\-swe\-bench\-verified/](https://openai.com/index/introducing-swe-bench-verified/)\.
- DeepSeek\-AI \(2026\)DeepSeek\-AI\.DeepSeek\-V4: Towards highly efficient million\-token context intelligence\.*arXiv preprint arXiv:2606\.19348*, 2026\.URL[https://arxiv\.org/abs/2606\.19348](https://arxiv.org/abs/2606.19348)\.
- Du et al\. \(2025\)Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, King Zhu, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, Kaixin Deng, Shawn Gavin, Shian Jia, Sichao Jiang, Qinrui Li, Rui Li, Sirun Li, Yizhi Li, Yunwen Li, Yiyan Liao, David Ma, Yuansheng Ni, Haoran Que, Qiyao Wang, Zekun Moore Wang, Zhoufutu Wen, Siwei Wu, Tyshawn Hsing, Ming Xu, Zhenzhu Yang, Junting Zhou, Yuelin Bai, Xingyuan Bu, Chenglin Cai, Liang Chen, Yifan Chen, Chengtuo Cheng, Tianhao Cheng, Keyi Ding, Siming Huang, Yun Huang, Yaoru Li, Yizhe Li, Zhaoqun Li, Tianhao Liang, Chengdong Lin, Hongquan Lin, Yinghao Ma, Tianyang Pang, Zhongyuan Peng, Zifan Peng, Qige Qi, Shi Qiu, Xingwei Qu, Shanghaoran Quan, Yizhou Tan, Chenqing Wang, Hao Wang, Yiya Wang, Yubo Wang, Zili Wang, Jiajun Xu, Kexin Yang, Ruibin Yuan, Yuanhao Yue, Tianyang Zhan, Chun Zhang, Jinyang Zhang, Xingjian Zhang, Xiyue Zhang, Yue Zhang, Yongchi Zhao, Xiangyu Zheng, Chenghua Zhong, Meng Cao, Yang Gao, Zhoujun Li, Dayiheng Liu, Qian Liu, Tianyu Liu, Shiwen Ni, Junran Peng, Yujia Qin, Wenbo Su, Guoyin Wang, Shi Wang, Jian Yang, Min Yang, Xiang Yue, Zhaoxiang Zhang, Wangchunshu Zhou, Jiaheng Liu, Qunshu Lin, Wenhao Huang, and Ge Zhang\.SuperGPQA: Scaling LLM evaluation across 285 graduate disciplines\.In*Advances in Neural Information Processing Systems*, volume 38, Main Conference\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-3766\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/a3c5af1f56fc73eef1ba0f442739f5ca\-Paper\-Datasets\_and\_Benchmarks\_Track\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/a3c5af1f56fc73eef1ba0f442739f5ca-Paper-Datasets_and_Benchmarks_Track.pdf)\.
- Fedus et al\. \(2022\)William Fedus, Barret Zoph, and Noam Shazeer\.Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity\.*Journal of Machine Learning Research*, 23\(120\):1–39, 2022\.URL[https://jmlr\.org/papers/v23/21\-0998\.html](https://jmlr.org/papers/v23/21-0998.html)\.
- GLM\-4\.5 Team \(2025\)GLM\-4\.5 Team\.GLM\-4\.5: Agentic, reasoning, and coding \(ARC\) foundation models\.*arXiv preprint arXiv:2508\.06471*, 2025\.URL[https://arxiv\.org/abs/2508\.06471](https://arxiv.org/abs/2508.06471)\.
- Goel et al\. \(2026\)Palaash Goel, Ayush Maheshwari, and Tanmoy Chakraborty\.It takes a MAESTRO to prune bad experts\.*arXiv preprint arXiv:2607\.08601*, 2026\.URL[https://arxiv\.org/abs/2607\.08601](https://arxiv.org/abs/2607.08601)\.
- Hugging Face Smol Models Research \(2025\)Hugging Face Smol Models Research\.SmolTalk2\.Hugging Face dataset card, 2025\.URL[https://huggingface\.co/datasets/HuggingFaceTB/smoltalk2](https://huggingface.co/datasets/HuggingFaceTB/smoltalk2)\.Accessed 2026\-09\-14\.
- Hyeon & Do \(2026\)Sieun Hyeon and Jaeyoung Do\.Is retraining\-free enough? the necessity of router calibration for efficient MoE compression\.*arXiv preprint arXiv:2603\.02217*, 2026\.URL[https://arxiv\.org/abs/2603\.02217](https://arxiv.org/abs/2603.02217)\.
- Jain et al\. \(2025\)Naman Jain, King Han, Alex Gu, Wen\-Ding Li, Fanjia Yan, Tianjun Zhang, Sida I\. Wang, Armando Solar\-Lezama, Koushik Sen, and Ion Stoica\.LiveCodeBench: Holistic and contamination free evaluation of large language models for code\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/hash/94074dd5a072d28ff75a76dabed43767\-Abstract\-Conference\.html](https://proceedings.iclr.cc/paper_files/paper/2025/hash/94074dd5a072d28ff75a76dabed43767-Abstract-Conference.html)\.
- Jaiswal et al\. \(2025\)Ajay Jaiswal, Jianyu Wang, Yixiao Li, Pingzhi Li, Tianlong Chen, Zhangyang Wang, Chong Wang, Ruoming Pang, and Xianzhi Du\.Finding fantastic experts in MoEs: A unified study for expert dropping strategies and observations\.*arXiv preprint arXiv:2504\.05586*, 2025\.URL[https://arxiv\.org/abs/2504\.05586](https://arxiv.org/abs/2504.05586)\.
- Jha et al\. \(2026\)Saurav Jha, Maryam Hashemzadeh, Ali Saheb Pasand, Ali Parviz, Min\-Joong Lee, and Boris Knyazev\.REAM: Merging improves pruning of experts in LLMs\.*arXiv preprint arXiv:2604\.04356*, 2026\.URL[https://arxiv\.org/abs/2604\.04356](https://arxiv.org/abs/2604.04356)\.
- Jimenez et al\. \(2024\)Carlos E\. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R\. Narasimhan\.SWE\-bench: Can language models resolve real\-world GitHub issues?In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=VTF8yNQM66](https://openreview.net/forum?id=VTF8yNQM66)\.
- Lasby et al\. \(2026\)Mike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie, Yani Ioannou, and Vithursan Thangarasa\.REAP the experts: Why pruning prevails for one\-shot MoE compression\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=ukGxWd2aDG](https://openreview.net/forum?id=ukGxWd2aDG)\.
- Lee et al\. \(2025\)Jaeseong Lee, Seung\-won Hwang, Aurick Qiao, Daniel F Campos, Zhewei Yao, and Yuxiong He\.STUN: Structured\-then\-unstructured pruning for scalable MoE pruning\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 13660–13676\. Association for Computational Linguistics, July 2025\.doi:10\.18653/v1/2025\.acl\-long\.671\.URL[https://aclanthology\.org/2025\.acl\-long\.671/](https://aclanthology.org/2025.acl-long.671/)\.
- Lepikhin et al\. \(2021\)Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen\.GShard: Scaling giant models with conditional computation and automatic sharding\.In*International Conference on Learning Representations*, 2021\.URL[https://openreview\.net/forum?id=qrwe7XHTmYb](https://openreview.net/forum?id=qrwe7XHTmYb)\.
- Lewkowycz et al\. \(2022\)Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman\-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur\-Ari, and Vedant Misra\.Solving quantitative reasoning problems with language models\.In*Advances in Neural Information Processing Systems*, volume 35, pp\. 3843–3857\. Curran Associates, Inc\., 2022\.doi:10\.52202/068431\-0278\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/file/18abbeef8cfe9203fdf9053c9c4fe191\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2022/file/18abbeef8cfe9203fdf9053c9c4fe191-Paper-Conference.pdf)\.
- Li et al\. \(2016\)Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan\.A diversity\-promoting objective function for neural conversation models\.In*Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pp\. 110–119\. Association for Computational Linguistics, 2016\.doi:10\.18653/v1/N16\-1014\.URL[https://aclanthology\.org/N16\-1014/](https://aclanthology.org/N16-1014/)\.
- Li et al\. \(2026\)Ke Li, Zheng Yang, Zhongbin Zhou, Feng Xue, Zhonglin Jiang, and Wenxiao Wang\.HEAPr: Hessian\-based efficient atomic expert pruning in output space\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=JAbMgS7gl6](https://openreview.net/forum?id=JAbMgS7gl6)\.
- Li et al\. \(2024\)Pingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi\-Lin Sung, Yu Cheng, Mohit Bansal, and Tianlong Chen\.Merge, then compress: Demystify efficient SMoE with hints from its routing policy\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=eFWG9Cy3WK](https://openreview.net/forum?id=eFWG9Cy3WK)\.
- Liu et al\. \(2024\)Enshu Liu, Junyi Zhu, Zinan Lin, Xuefei Ning, Matthew B\. Blaschko, Shengen Yan, Guohao Dai, Huazhong Yang, and Yu Wang\.Efficient expert pruning for sparse mixture\-of\-experts language models: Enhancing performance and reducing inference costs\.*arXiv preprint arXiv:2407\.00945*, 2024\.URL[https://arxiv\.org/abs/2407\.00945](https://arxiv.org/abs/2407.00945)\.
- Liu et al\. \(2023\)Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang\.Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation\.In*Advances in Neural Information Processing Systems*, volume 36, pp\. 21558–21572\. Curran Associates, Inc\., 2023\.doi:10\.52202/075280\-0943\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/43e9d647ccd3e4b7b5baab53f0368686\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/43e9d647ccd3e4b7b5baab53f0368686-Paper-Conference.pdf)\.
- Liu et al\. \(2026a\)Zongfang Liu, Guangyi Chen, Shengkun Tang, Yifan Shen, Huan Wang, and Xin Yuan\.AIMER: Calibration\-free task\-agnostic MoE expert pruning\.*arXiv preprint arXiv:2603\.18492*, 2026a\.URL[https://arxiv\.org/abs/2603\.18492](https://arxiv.org/abs/2603.18492)\.
- Liu et al\. \(2026b\)Zongfang Liu, Jinghui Zhang, Zijian Ma, Guangyi Chen, and Xin Yuan\.How to score experts for one\-shot MoE expert pruning: A unified formulation and selection principle\.*arXiv preprint arXiv:2606\.15716*, 2026b\.URL[https://arxiv\.org/abs/2606\.15716](https://arxiv.org/abs/2606.15716)\.
- Lu et al\. \(2024\)Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, and Hongsheng Li\.Not all experts are equal: Efficient expert pruning and skipping for mixture\-of\-experts large language models\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 6159–6172\. Association for Computational Linguistics, August 2024\.doi:10\.18653/v1/2024\.acl\-long\.334\.URL[https://aclanthology\.org/2024\.acl\-long\.334/](https://aclanthology.org/2024.acl-long.334/)\.
- Mao et al\. \(2025\)Huanzhi Mao, Raymond Tsao, Jingzhuo Zhou, Shishir G\. Patil, and Joseph E\. Gonzalez\.BFCL V4: Agentic part 1: Web search, July 2025\.URL[https://gorilla\.cs\.berkeley\.edu/blogs/15\_bfcl\_v4\_web\_search\.html](https://gorilla.cs.berkeley.edu/blogs/15_bfcl_v4_web_search.html)\.
- Mathematical Association of America \(n\.d\.\)Mathematical Association of America\.MAA invitational competitions\.Official competition website, n\.d\.URL[https://maa\.org/maa\-invitational\-competitions/](https://maa.org/maa-invitational-competitions/)\.
- Merity et al\. \(2017\)Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher\.Pointer sentinel mixture models\.In*International Conference on Learning Representations*, 2017\.URL[https://openreview\.net/forum?id=Byj72udxe](https://openreview.net/forum?id=Byj72udxe)\.
- Muzio et al\. \(2024\)Alexandre Muzio, Alex Sun, and Churan He\.SEER\-MoE: Sparse expert efficiency through regularization for mixture\-of\-experts\.*arXiv preprint arXiv:2404\.05089*, 2024\.URL[https://arxiv\.org/abs/2404\.05089](https://arxiv.org/abs/2404.05089)\.
- NVIDIA \(2026\)NVIDIA\.Nemotron 3 Super: Open, efficient mixture\-of\-experts hybrid Mamba\-Transformer model for agentic reasoning\.*arXiv preprint arXiv:2604\.12374*, 2026\.URL[https://arxiv\.org/abs/2604\.12374](https://arxiv.org/abs/2604.12374)\.
- Patil et al\. \(2025\)Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng\-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E\. Gonzalez\.The berkeley function calling leaderboard \(BFCL\): From tool use to agentic evaluation of large language models\.In*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, pp\. 48371–48392\. PMLR, 2025\.URL[https://proceedings\.mlr\.press/v267/patil25a\.html](https://proceedings.mlr.press/v267/patil25a.html)\.
- Pyatkin et al\. \(2025\)Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert, and Hannaneh Hajishirzi\.Generalizing verifiable instruction following\.In*Advances in Neural Information Processing Systems*, volume 38, Main Conference\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-1645\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/46499a0622ecf568b72d17b61e45dbd5\-Paper\-Datasets\_and\_Benchmarks\_Track\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/46499a0622ecf568b72d17b61e45dbd5-Paper-Datasets_and_Benchmarks_Track.pdf)\.
- Qwen Team \(2026\)Qwen Team\.Qwen3\.5: Towards native multimodal agents, February 2026\.URL[https://qwen\.ai/blog?id=qwen3\.5](https://qwen.ai/blog?id=qwen3.5)\.
- Shazeer et al\. \(2017\)Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean\.Outrageously large neural networks: The sparsely\-gated mixture\-of\-experts layer\.In*International Conference on Learning Representations*, 2017\.URL[https://openreview\.net/forum?id=B1ckMDqlg](https://openreview.net/forum?id=B1ckMDqlg)\.
- Team Olmo et al\. \(2025\)Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng, Shane Arora, Shashank Gupta, Taira Anderson, Teng Xiao, Tyler Murray, Tyler Romero, Victoria Graf, Akari Asai, Akshita Bhagia, Alexander Wettig, Alisa Liu, Aman Rangapur, Chloe Anastasiades, Costa Huang, Dustin Schwenk, Harsh Trivedi, Ian Magnusson, Jaron Lochner, Jiacheng Liu, Lester James V\. Miranda, Maarten Sap, Malia Morgan, Michael Schmitz, Michal Guerquin, Michael Wilson, Regan Huff, Ronan Le Bras, Rui Xin, Rulin Shao, Sam Skjonsberg, Shannon Zejiang Shen, Shuyue Stella Li, Tucker Wilde, Valentina Pyatkin, Will Merrill, Yapei Chang, Yuling Gu, Zhiyuan Zeng, Ashish Sabharwal, Luke Zettlemoyer, Pang Wei Koh, Ali Farhadi, Noah A\. Smith, and Hannaneh Hajishirzi\.Olmo 3, 2025\.URL[https://arxiv\.org/abs/2512\.13961](https://arxiv.org/abs/2512.13961)\.Dataset card:[https://huggingface\.co/datasets/allenai/Dolci\-Instruct\-SFT](https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT)\.
- Teknium \(2023\)Teknium\.OpenHermes 2\.5: An open dataset of synthetic data for generalist LLM assistants\.Hugging Face dataset card, 2023\.URL[https://huggingface\.co/datasets/teknium/OpenHermes\-2\.5](https://huggingface.co/datasets/teknium/OpenHermes-2.5)\.
- Tencent Hy Team \(2026\)Tencent Hy Team\.Hy3\.Hugging Face model card, 2026\.URL[https://huggingface\.co/tencent/Hy3](https://huggingface.co/tencent/Hy3)\.
- Xie et al\. \(2024\)Yanyue Xie, Zhi Zhang, Ding Zhou, Cong Xie, Ziang Song, Xin Liu, Yanzhi Wang, Xue Lin, and An Xu\.MoE\-Pruner: Pruning mixture\-of\-experts large language model using the hints from its router\.*arXiv preprint arXiv:2410\.12013*, 2024\.URL[https://arxiv\.org/abs/2410\.12013](https://arxiv.org/abs/2410.12013)\.
- Yang et al\. \(2024\)Cheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang, Yu Gong, Yuanlin Duan, Wenqi Jia, Miao Yin, Yu Cheng, and Bo Yuan\.MoE\-I2: Compressing mixture of experts models through inter\-expert pruning and intra\-expert low\-rank decomposition\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pp\. 10456–10466\. Association for Computational Linguistics, 2024\.doi:10\.18653/v1/2024\.findings\-emnlp\.612\.URL[https://aclanthology\.org/2024\.findings\-emnlp\.612/](https://aclanthology.org/2024.findings-emnlp.612/)\.
- Yao et al\. \(2026\)Yilun Yao, Jiaming Pan, Elsie Dai, Peizhuang Cong, Yaoming Li, and Tong Yang\.ConMoE: Expert\-pool consolidation via prototype reassignment for MoE compression\.*arXiv preprint arXiv:2605\.29350*, 2026\.URL[https://arxiv\.org/abs/2605\.29350](https://arxiv.org/abs/2605.29350)\.
- Zhang \(2026\)Yuhao Zhang\.SHAPE: Coalition\-aware expert pruning for sparse mixture\-of\-experts LLMs\.*arXiv preprint arXiv:2606\.09886*, 2026\.URL[https://arxiv\.org/abs/2606\.09886](https://arxiv.org/abs/2606.09886)\.
- Zhang et al\. \(2026\)Zeliang Zhang, Nikhil Ghosh, Jiani Liu, Bin Yu, and Xiaodong Liu\.Does a global perspective help prune sparse MoEs elegantly?*arXiv preprint arXiv:2604\.06542*, 2026\.URL[https://arxiv\.org/abs/2604\.06542](https://arxiv.org/abs/2604.06542)\.
- Zhou et al\. \(2023\)Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou\.Instruction\-following evaluation for large language models\.*arXiv preprint arXiv:2311\.07911*, 2023\.URL[https://arxiv\.org/abs/2311\.07911](https://arxiv.org/abs/2311.07911)\.

Roadmap\.The appendix supports three claims made in the main text and bounds what each one rests on\. Appendices[A](https://arxiv.org/html/2609.30465#A1)and[B](https://arxiv.org/html/2609.30465#A2)establish that the reported scores are the exact local deletion quantities and that the calibration, fidelity, and benchmark protocols are kept apart\. Appendices[C](https://arxiv.org/html/2609.30465#A3)and[D](https://arxiv.org/html/2609.30465#A4)establish which scoring choices actually move the retained set, finding that aggregation can matter as much as the residual reference and that selection agreement is less stable for residual criteria than for REAP\. Appendix[E](https://arxiv.org/html/2609.30465#A5)establishes that the fidelity gains are uneven across evaluation axes, routing regimes, and metrics, and that they do not extend to generation behavior\. Appendix[F](https://arxiv.org/html/2609.30465#A6)places the method among compression families that change the expert pool or the router\.

Scope of the reported estimates\.Two limits apply throughout and are not repeated in each subsection\. First, every study evaluates fixed checkpoints, so no reported interval captures variability over calibration draws, pruning runs, or independent training runs\. Second, matching a scoring rule across two studies does not imply that they share calibration inputs or retained expert sets, so fidelity and benchmark results are compared within a study rather than across studies\.

## Appendix AAdditional Derivations and Implementation

This appendix follows the scoring quantity from its geometry to its implementation\. The residual decomposition below shows what the consensus residual measures, and Appendix[A\.2](https://arxiv.org/html/2609.30465#A1.SS2)shows how the leave\-one\-out and refill factors convert it into exact single\-deletion damage\. Appendix[A\.3](https://arxiv.org/html/2609.30465#A1.SS3)then shows why those exact single\-deletion quantities do not compose into an exact objective for a deletion set, which is what makes our selection a ranking heuristic rather than a solved optimization\. Appendix[A\.4](https://arxiv.org/html/2609.30465#A1.SS4)describes how the scores are accumulated in one calibration pass and how the retained sets are realized as smaller checkpoints\.

### A\.1Residual decomposition

The consensus residual depends on the angle between an expert output and the mixture, which is the information that gates and output norms omit\. Expanding it gives

‖fi−c‖22\\displaystyle\\\|f\_\{i\}\-c\\\|\_\{2\}^\{2\}=‖fi‖22−2​⟨fi,c⟩\+‖c‖22,\\displaystyle=\\\|f\_\{i\}\\\|\_\{2\}^\{2\}\-2\\langle f\_\{i\},c\\rangle\+\\\|c\\\|\_\{2\}^\{2\},\(5\)⟨fi,c⟩\\displaystyle\\langle f\_\{i\},c\\rangle=∥fi∥2∥c∥2cosθi\.\\displaystyle=\\\|f\_\{i\}\\\|\_\{2\}\\,\\\|c\\\|\_\{2\}\\cos\\theta\_\{i\}\.\(6\)At fixed norms, a positive inner product reduces the residual and a negative one increases it\. Even without collinear or near\-duplicate outputs, an aligned expert may therefore be replaceable, while one balancing the others may be costly to lose\. The mixture still includes expertii, so a further self\-inclusion correction is needed to obtain distance to the survivors\. Because mixture norms, gates, and routed sets vary across calibration tokens, this geometric identity constrains each token’s damage without determining which retained set performs better\.

##### Constructed example in Figure[1](https://arxiv.org/html/2609.30465#S1.F1)\.

Letu,vu,vbe orthonormal and letEA,EB,ECE\_\{A\},E\_\{B\},E\_\{C\}have outputs8​u8u,3​u\+4​v3u\+4v, and10​u−4​v10u\-4v, with equal initial weights\. Their consensus isc=7​uc=7u, the routed\-expert mixture at one layer, not the model’s final output\. Residuals are‖fi−c‖2\\\|f\_\{i\}\-c\\\|\_\{2\}\. Without refill, deletingEA,EB,ECE\_\{A\},E\_\{B\},E\_\{C\}and renormalizing the survivors gives damage0\.50\.5,8≈2\.83\\sqrt\{8\}\\approx 2\.83, and2\.52\.5, respectively\. With refill, the same initially unselectedED=10​u−5​vE\_\{D\}=10u\-5venters after every deletion, with survivor and replacement masses1,1,0\.81,1,0\.8\. For deletion ofECE\_\{C\}, the bars show−\(fC−c\)\-\(f\_\{C\}\-c\),0\.8​\(fD−c\)0\.8\(f\_\{D\}\-c\), and their unnormalized sum\. Thevvterms cancel, givingΔc=−0\.6u/2\.8\\Delta c=\-0\.6u/2\.8and damage3/14≈0\.213/14\\approx 0\.21, versus1\.511\.51for deletingEAE\_\{A\}and3\.663\.66for deletingEBE\_\{B\}\. Thus magnitude, fixed\-support damage, and refill damage favorEB,EA,ECE\_\{B\},E\_\{A\},E\_\{C\}, respectively\.

In the panels, curves span output coordinates, with mixture curves using2\.5×2\.5\\timesthe expert vertical scale\. Dotted curves isolateuucomponents, and shading marks coordinate\-wise differences\. The construction illustrates the role of output geometry and refill rather than a measured checkpoint\-level gain\.

### A\.2Survivor\-distance identity

Forwi<1w\_\{i\}<1, starting from the fixed\-support post\-removal output gives

c−i\\displaystyle c^\{\-i\}=c−wi​fi1−wi,\\displaystyle=\\frac\{c\-w\_\{i\}f\_\{i\}\}\{1\-w\_\{i\}\},\(7\)fi−c−i\\displaystyle f\_\{i\}\-c^\{\-i\}=fi−c−wi​fi1−wi=fi−c1−wi\.\\displaystyle=f\_\{i\}\-\\frac\{c\-w\_\{i\}f\_\{i\}\}\{1\-w\_\{i\}\}=\\frac\{f\_\{i\}\-c\}\{1\-w\_\{i\}\}\.\(8\)Taking norms gives the survivor\-distance identity used in the main text,

‖fi−c‖2=\(1−wi\)​‖fi−c−i‖2,\\\|f\_\{i\}\-c\\\|\_\{2\}=\(1\-w\_\{i\}\)\\,\\\|f\_\{i\}\-c^\{\-i\}\\\|\_\{2\},\(9\)and combining equation[9](https://arxiv.org/html/2609.30465#A1.E9)with Proposition[3](https://arxiv.org/html/2609.30465#Thmproposition3)yields

‖c−c−i‖2=wi​‖fi−c−i‖2,\\\|c\-c^\{\-i\}\\\|\_\{2\}=w\_\{i\}\\\|f\_\{i\}\-c^\{\-i\}\\\|\_\{2\},\(10\)which expresses damage as routing mass times distance to the survivors\. The distance describes local functional replaceability at the same input, while the weight converts that distance into mixture change\. Neither factor alone measures redundancy in the final retained model\.

The factor1/\(1−wi\)1/\(1\-w\_\{i\}\)is what converts the observed consensus residual into this survivor distance\. Omitting it yields\(1−wi\)​δiloo\(1\-w\_\{i\}\)\\delta\_\{i\}^\{\\mathrm\{loo\}\}at each token, which is a different proxy rather than the same deletion damage under a rescaling, becausewiw\_\{i\}varies across tokens\. The identity also makes the local comparison translation invariant\. With inputs and gates fixed, adding a common vectorggto all routed outputs shifts both mixtures byggand leaves deletion damage unchanged, although the output norms that magnitude scores rely on do change\. This invariance holds for the local damage, not for model predictions\.

Forδiloo\>0\\delta\_\{i\}^\{\\mathrm\{loo\}\}\>0, expanding the refill numerator gives

δiδiloo=1−wi1−wi\+wr​1−2ρicosθi\+ρi2,ρi=wr​‖rr‖2wi​‖ri‖2,\\frac\{\\delta\_\{i\}\}\{\\delta\_\{i\}^\{\\mathrm\{loo\}\}\}=\\frac\{1\-w\_\{i\}\}\{1\-w\_\{i\}\+w\_\{r\}\}\\,\\sqrt\{1\-2\\rho\_\{i\}\\cos\\theta\_\{i\}\+\\rho\_\{i\}^\{2\}\},\\qquad\\rho\_\{i\}=\\frac\{w\_\{r\}\\\|r\_\{r\}\\\|\_\{2\}\}\{w\_\{i\}\\\|r\_\{i\}\\\|\_\{2\}\},\(11\)whereθi\\theta\_\{i\}is the angle between nonzero residuals\. The two factors pull in opposite directions\. The denominator factor is at most one and strictly smaller whenwr\>0w\_\{r\}\>0, whereas the numerator factor exceeds one exactly whenρi\>2cosθi\\rho\_\{i\}\>2\\cos\\theta\_\{i\}, and equals one when the promoted weighted residual vanishes\. Orthogonal residuals thus already raise the numerator factor for anyρi\>0\\rho\_\{i\}\>0, yet routing mass, residual magnitudes, and alignment jointly decide whether refill increases the damage ratio itself\. Appendix[C\.4](https://arxiv.org/html/2609.30465#A3.SS4)reports score ratios and partial geometric summaries, which measure the observed effect without decomposing it into these three factors\.

### A\.3Exact perturbation for a selected set

Joint removal couples the experts it deletes, so the exact single\-deletion identities do not extend additively to a selected set\. For a removed routed subset𝒜\\mathcal\{A\}withW𝒜<1W\_\{\\mathcal\{A\}\}<1, the survivor mixture isc−𝒜=\(c−∑j∈𝒜wj​fj\)/\(1−W𝒜\)c^\{\-\\mathcal\{A\}\}=\(c\-\\sum\_\{j\\in\\mathcal\{A\}\}w\_\{j\}f\_\{j\}\)/\(1\-W\_\{\\mathcal\{A\}\}\)\. Subtracting it fromccgives Eq\.[4](https://arxiv.org/html/2609.30465#S2.E4)\. All residualsrj=fj−cr\_\{j\}=f\_\{j\}\-cshare the original consensus, so the squared numerator expands as

‖∑j∈𝒜wj​rj‖22=∑j∈𝒜wj2​‖rj‖22⏟diagonal\+∑j≠l∈𝒜wj​wl​⟨rj,rl⟩⏟interference\.\\Big\\\|\\sum\_\{j\\in\\mathcal\{A\}\}w\_\{j\}r\_\{j\}\\Big\\\|\_\{2\}^\{2\}=\\underbrace\{\\sum\_\{j\\in\\mathcal\{A\}\}w\_\{j\}^\{2\}\\\|r\_\{j\}\\\|\_\{2\}^\{2\}\}\_\{\\text\{diagonal\}\}\+\\underbrace\{\\sum\_\{j\\neq l\\in\\mathcal\{A\}\}w\_\{j\}w\_\{l\}\\langle r\_\{j\},r\_\{l\}\\rangle\}\_\{\\text\{interference\}\}\.\(12\)Interference cross terms and the set\-dependent denominator block an additive decomposition into independent removal costs\. The diagonal term also sums experts within a token, whereas RMS aggregates tokens within an expert, so the expansion neither derives RMS nor gives an exact joint objective obtained by summing expert scores\.

Refill does not repair this\. Removing𝒜\\mathcal\{A\}promotes the\|𝒜\|\|\\mathcal\{A\}\|highest\-ranked unselected expertsℛ\\mathcal\{R\}, so the numerator becomes‖∑j∈𝒜wj​rj−∑r∈ℛwr​rr‖2\\\|\\sum\_\{j\\in\\mathcal\{A\}\}w\_\{j\}r\_\{j\}\-\\sum\_\{r\\in\\mathcal\{R\}\}w\_\{r\}r\_\{r\}\\\|\_\{2\}and the denominator becomes1−W𝒜\+Wℛ1\-W\_\{\\mathcal\{A\}\}\+W\_\{\\mathcal\{R\}\}\. Both now depend on the whole set, and the expansion acquires cross terms withinℛ\\mathcal\{R\}and between𝒜\\mathcal\{A\}andℛ\\mathcal\{R\}on top of those in Eq\.[12](https://arxiv.org/html/2609.30465#A1.E12)\. Summing scalar single\-expert scores captures none of these interactions, which is why our selection remains a ranking heuristic under a fixed budget\.

### A\.4Implementation and complexity

Algorithm 1Chunked scoring and routing\-aware physical pruning for the RCS family0:Model, calibration data

𝒟\\mathcal\{D\}with valid\-position masks, budgets

BℓB\_\{\\ell\}, chunk size

CC, criterion

a∈\{RCS,RCS​\-​LOO,RCS​\-​Refill\}a\\in\\\{\\mathrm\{RCS\},\\mathrm\{RCS\\mbox\{\-\}LOO\},\\mathrm\{RCS\\mbox\{\-\}Refill\}\\\}, clamp

ϵ\\epsilon
1:Initialize per\-shard

qℓ​i←0q\_\{\\ell i\}\\leftarrow 0,

nℓ​i←0n\_\{\\ell i\}\\leftarrow 0for learned\-router layers

2:forlearned\-router layer input from an unpruned forward over

𝒟\\mathcal\{D\}do

3:Obtain routed sets

StS\_\{t\}and normalized selected weights

wtw\_\{t\}
4:if

a=RCS​\-​Refilla=\\mathrm\{RCS\\mbox\{\-\}Refill\}then

5:Obtain the rank\-

\(k\+1\)\(k\{\+\}1\)expert

rtr\_\{t\}and its pseudo\-weight

wr​\(t\)w\_\{r\}\(t\)
6:endif

7:fortoken chunk

XXof size at most

CCdo

8:Reconstruct routed outputs

fi​\(t\)f\_\{i\}\(t\)and form

ct=∑i∈Stwi​\(t\)​fi​\(t\)c\_\{t\}=\\sum\_\{i\\in S\_\{t\}\}w\_\{i\}\(t\)f\_\{i\}\(t\)and

ri​\(t\)=fi​\(t\)−ctr\_\{i\}\(t\)=f\_\{i\}\(t\)\-c\_\{t\}
9:if

a=RCSa=\\mathrm\{RCS\}then

10:

di​\(t\)←wi​\(t\)​‖ri​\(t\)‖2d\_\{i\}\(t\)\\leftarrow w\_\{i\}\(t\)\\\|r\_\{i\}\(t\)\\\|\_\{2\}
11:elseif

a=RCS​\-​LOOa=\\mathrm\{RCS\\mbox\{\-\}LOO\}then

12:

di​\(t\)←wi​\(t\)​‖ri​\(t\)‖2/max⁡\(1−wi​\(t\),ϵ\)d\_\{i\}\(t\)\\leftarrow w\_\{i\}\(t\)\\\|r\_\{i\}\(t\)\\\|\_\{2\}/\\max\(1\-w\_\{i\}\(t\),\\epsilon\)
13:else

14:Reconstruct

frt​\(t\)f\_\{r\_\{t\}\}\(t\), set

rrt​\(t\)=frt​\(t\)−ctr\_\{r\_\{t\}\}\(t\)=f\_\{r\_\{t\}\}\(t\)\-c\_\{t\}
15:

di​\(t\)←‖wi​\(t\)​ri​\(t\)−wr​\(t\)​rrt​\(t\)‖2/max⁡\(1−wi​\(t\)\+wr​\(t\),ϵ\)d\_\{i\}\(t\)\\leftarrow\\\|w\_\{i\}\(t\)r\_\{i\}\(t\)\-w\_\{r\}\(t\)r\_\{r\_\{t\}\}\(t\)\\\|\_\{2\}/\\max\(1\-w\_\{i\}\(t\)\+w\_\{r\}\(t\),\\epsilon\)
16:endif

17:For valid routed pairs, accumulate

qℓ​i\+=di​\(t\)2q\_\{\\ell i\}\\mathrel\{\+\}=d\_\{i\}\(t\)^\{2\},

nℓ​i\+=1n\_\{\\ell i\}\\mathrel\{\+\}=1; discard chunk\-local outputs

18:endfor

19:endfor

20:Sum

q,nq,nacross disjoint shards; set

sℓ​i=qℓ​i/max⁡\(nℓ​i,1\)s\_\{\\ell i\}=\\sqrt\{q\_\{\\ell i\}/\\max\(n\_\{\\ell i\},1\)\}
21:Retain the

BℓB\_\{\\ell\}highest\-scoring experts in each learned\-router layer

22:For fixed\-hash layers, select by table counts and remap to distinct surviving routes

23:returnModel with consistently sliced expert tensors, router rows, and updated route indices

#### A\.4\.1Online statistics and bounded working buffers

Algorithm[1](https://arxiv.org/html/2609.30465#alg1)changes only the token\-level quantity across the three criteria, which share chunking, conditional RMS aggregation, ranking, and physical pruning\. The RCS and RCS\-LOO branches reconstruct only the routed expert outputs, whereas RCS\-Refill additionally reconstructs the rank\-\(k\+1\)\(k\{\+\}1\)promoted expert\. Selecting RCS\-Refill with conditional RMS givesRazor\. For each learned\-router layer, the collector obtains the routing context for the batch, partitions its token positions, and locally reconstructs the required expert outputs for each chunk\. These outputs form the consensus and residuals before the collector advances to the next chunk\. The cost is therefore additional local computation alongside one original forward pass, rather than a separate end\-to\-end model evaluation for every candidate deletion\.

Chunking bounds the working buffers\. WithTTpacked token positions, top\-kkrouting, and expert\-output widthdd, storing the routed outputs together with the promoted one requiresO⁡\(T⁡\(k\+1\)​d\)O\(T\(k\+1\)d\)space\. Restricting the working set toCCpositions replaces this withO⁡\(C⁡\(k\+1\)​d\)O\(C\(k\+1\)d\), plusO⁡\(C​d\)O\(Cd\)for the consensus\. Onlyqi=∑tδi​\(t\)2q\_\{i\}=\\sum\_\{t\}\\delta\_\{i\}\(t\)^\{2\}and the valid routed\-token countnin\_\{i\}persist for RMS scoring, requiringO⁡\(E\)O\(E\)state per layer\. Disjoint calibration shards accumulate these quantities independently\. Summing them before takingqi/max⁡\(ni,1\)\\sqrt\{q\_\{i\}/\\max\(n\_\{i\},1\)\}gives the pooled conditional RMS, unlike averaging shard\-level RMS scores\. Unobserved experts receive zero score, and we remove the lowestE−BE\-Bscores without a secondary criterion for exact ties\. These bounds cover only the storage controlled by chunking, excluding model weights and ordinary forward activations\.

One implementation choice affects the computed numbers\. Each denominator is clamped from below by10−610^\{\-6\}, givingmax⁡\(1−wi\+wr,10−6\)\\max\(1\-w\_\{i\}\+w\_\{r\},10^\{\-6\}\)for the refill score andmax⁡\(1−wi,10−6\)\\max\(1\-w\_\{i\},10^\{\-6\}\)for its fixed\-support special case\. Because the identities describe the unclamped quantity, an active clamp makes the computed score differ from exact local damage\.

#### A\.4\.2Layer\-wise execution

Output chunking bounds activation buffers but leaves the memory occupied by model weights untouched\. For models that exceed the available device capacity, a layer\-wise schedule loads one decoder layer, propagates the current calibration shard through it, collects its statistics, and releases its weights before loading the next layer\. This is what makes the largest backbones here scorable on a fixed device budget\. Every such schedule realizes the same score definition, so the resulting checkpoint does not depend on which one is used\.

#### A\.4\.3Architecture\-specific routing

The derivation requires normalized selected weights, not a particular router parameterization\. We preserve each architecture’s native selection rule and score function, using softmax probabilities for Qwen3\.6\-35B\-A3B, sigmoid scores for GLM\-4\.7\-Flash, DeepSeek\-V4\-Flash\-0731’s native score transformation, and sigmoid scores for Hy3\. Where a correction bias determines selection, it is not added to the mixture weights\. Selected unmodified scores are renormalized before evaluating the damage formula\. The implementation retains the routed\-output scaleλ\\lambda, although omitting this common positive factor from the displayed score leaves within\-layer rankings unchanged\.

DeepSeek\-V4\-Flash\-0731 includes three fixed\-hash layers whose expert identities come from a frozen token\-to\-expert table, so they fall outside learned\-router saliency and need a separate rule\. To retainBBexperts, we select the most frequent experts by counting occurrences over the entire table, without weighting by calibration\-token occurrence\. Each table row then preserves its surviving expert identities, and removed entries are replaced, in slot order, by the cosine\-nearest surviving router\-weight vector not already used in that row\. WithB≥kB\\geq k, the result haskkdistinct routes, and reindexing follows the same survivor order as parameter slicing\. The procedure repairs routing rather than merging expert weights or solving an assignment problem\. The routing\-stratified study shares this hash\-layer selection across criteria, so its comparisons isolate learned\-router saliency \(Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)\)\.

#### A\.4\.4Physical pruning

We slice expert\-indexed tensors and corresponding router rows with the same keep indices, leaving dense blocks and shared experts unchanged\. Under a common retained width, the result is a smaller MoE checkpoint requiring neither runtime masks nor custom sparse kernels\. For Qwen3\.6\-35B\-A3B, reducing 256 to 128 routed experts per layer yields an approximately 19B\-parameter model\.

Removing experts reduces stored parameters, but preserving top\-kkdoes not proportionally reduce active expert computation\. We do not measure serving latency, energy use, or end\-to\-end acceleration\.

## Appendix BData and Evaluation Protocol

A pruning method that selects experts from data can appear to generalize merely because the same data reappears at evaluation\. Four collections therefore serve four purposes and are kept disjoint in role throughout\. The calibration poolRazorCalscores experts \(Appendix[B\.1](https://arxiv.org/html/2609.30465#A2.SS1)\), the eight\-axis reference pool measures predictive fidelity \(Appendix[B\.2](https://arxiv.org/html/2609.30465#A2.SS2)\), the nine\-task suite measures downstream performance \(Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)\), and a separate repeated\-generation collection measures response behavior \(Appendix[E\.5](https://arxiv.org/html/2609.30465#A5.SS5)\)\. Only the first influences which experts are retained\. The fidelity and diversity collections also split their axes into four thatRazorCalcovers and four that it does not, so calibration\-covered and held\-out performance can be reported separately throughout Appendix[E](https://arxiv.org/html/2609.30465#A5)\. All four draw on open Nemotron post\-training datasets, whose development and reuse across the model family are described in the Nemotron 3 Super technical report\([NVIDIA, 2026](https://arxiv.org/html/2609.30465#bib.bib31)\)\. The specific dataset versions are listed below\.

### B\.1Calibration data

RazorCalcontains 2,048 chat\-format examples across seven domains \(Table[3](https://arxiv.org/html/2609.30465#A2.T3)\)\. It is a calibration mixture assembled from existing instruction datasets, not a new independently collected corpus\. Coding receives twice the allocation of each other domain to cover function completion, competitive programming, class\-based solutions, software engineering, and tool use\. Source\-specific sampling applies quotas, length filters, and deduplication\. Preprocessing removes dataset\-specific harness prefixes where appropriate and caps examples at 32K characters, not tokens\. Each checkpoint’s chat template renders the preserved conversation roles before tokenization\.

Table 3:RazorCaldomain allocations and calibration sources\. Counts denote source examples, not tokens\.##### Sources and composition\.

Coding includes 205 Python\-algorithm examples from Dolci\-Instruct\-SFT, the instruction\-tuning mixture accompanying Olmo 3\([Team Olmo et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib36)\)\. It adds 141 Python and 115 C\+\+ examples from Nemotron\-SFT\-Competitive\-Programming\-v2, 40 agentless software\-engineering examples from Nemotron\-SFT\-SWE\-v2, and 11 general tool\-use examples from Nemotron\-SFT\-OpenCode\-v1\. Nemotron\-SFT\-SWE\-v2 covers code localization, repair, and test generation, whereas Nemotron\-SFT\-OpenCode\-v1 covers agent interactions\.

Mathematics uses 256 AoPS examples from thehigh\_part00split of Nemotron\-Math\-v2, which pairs problems from Art of Problem Solving with model\-generated solution traces\. The allocation retains 160 examples from the earlier calibration pool and adds 96 separately selected examples from the same source\. High denotes the teacher’s reasoning mode, not problem difficulty\. Science combines 176 chemistry reasoning questions and 80 multiple\-choice questions from Nemotron\-Science\-v1\. Chinese\-STEM uses translated mathematics, code, and STEM examples from Nemotron\-SFT\-Multilingual\-v1 \(86/85/85\), which translates prompts and final answers rather than all reasoning traces\. Instruction following draws 256 examples from a Dolci constraint\-following pool containing Instruct, Think, and Persona examples\([Team Olmo et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib36)\), rather than from the Instruct mixture alone\. Tool calling uses 256 multi\-turn agent–tool conversations from Nemotron\-Agentic\-v1\. The world\-knowledge allocation contains 256 OpenHermes\-2\.5 examples\([Teknium, 2023](https://arxiv.org/html/2609.30465#bib.bib37)\), drawn from the cleaned, non\-thinking subset in SmolTalk2\([Hugging Face Smol Models Research, 2025](https://arxiv.org/html/2609.30465#bib.bib9)\)\. OpenHermes is a general\-purpose synthetic instruction corpus rather than a dedicated knowledge test, and all quotas balance source examples rather than token counts\.

Figure 5:Overview of theRazorCalcalibration pool\.\(a\) Cosine t\-SNE projection of Qwen3\-Embedding\-8B representations of the rendered conversations, with color and marker shape encoding domain and labels marking domain medians\. \(b\) Per\-domain character\-length coverage over the full observed range\. The curve is a kernel\-density estimate in log\-character space, the bar is the interquartile range with the median marked, the thin line spans minimum to maximum, and each dot is one of the 2,048 examples\. Domain counts are given in Table[3](https://arxiv.org/html/2609.30465#A2.T3)\.
##### Scored positions and coverage\.

For the diagnostic collection, calibration examples are rendered with thinking disabled and packed into 4,096\-token rows\. Calibration statistics include all nonpadding conversation positions, including prompts and role headers, unlike the assistant\-only fidelity targets in Appendix[B\.2](https://arxiv.org/html/2609.30465#A2.SS2)\. Qwen3\.6\-35B\-A3B scores 785,703 valid tokens from 921 source examples, while GLM\-4\.7\-Flash scores 785,632 tokens from 992\. We call these the*diagnostic calibration collections*below, and all score\-distribution, Gram\-matrix, and single\-expert ablation statistics in Appendix[C](https://arxiv.org/html/2609.30465#A3)come from them\. Because the two tokenizers give different token counts and source coverage at the same row capacity, these statistics are compared within a backbone rather than across backbones\.

### B\.2Fidelity evaluation data

Reverse KL and reference\-token perplexity use a pool of 8,192 conversations, with 1,024 examples on each of eight task axes\. These compare predictions on shared reference responses rather than scoring benchmark answers\. Four axes cover calibration task types\.

Mathematics\.Mathematical problem\-solving responses from themediumsplit of Nemotron\-Math\-v2\. Medium denotes the teacher’s reasoning mode, not problem difficulty\.

Code\.Python solutions from thecompetitive\_coding\_pythonsplit of Nemotron\-SFT\-Competitive\-Programming\-v2\.

Instruction following\.Constraint\-following conversations from thechat\_ifsplit of Nemotron\-Instruction\-Following\-Chat\-v1, selected by its instruction\-following capability label\.

Tools\.Agent–tool interaction trajectories from thetool\_callingsplit of Nemotron\-Agentic\-v1\.

The other four axes hold out task types\.

Chat\.Examples carrying the chat capability label in the same Chat\-v1chat\_ifsplit\.

Creative writing\.Examples selected by prompt keywords from thereasoning\_offsplit of Nemotron\-SFT\-Instruction\-Following\-Chat\-v2\. Creative writing is our selection, not an official split\.

Safety\.Reference safety\-aligned responses from thetrainsplit of Nemotron\-SFT\-Safety\-v1\.

SQL\.Examples from thetext\_to\_sqlsplit of the competitive\-programming source, separate from its Python subset, pairing natural\-language tasks and database schemas with SQL\. ID/OOD denotes calibration task\-type coverage, not disjoint source datasets or a formal distribution shift, and both groups contain synthetic data\.

Each axis retains 1,024 examples that fit the row capacity and carry enough assistant content to score\. Conversations end on an assistant response, and plain\-text assistant turns may be truncated while structured tool calls remain intact\.

Evaluation packs examples into 4,096\-token rows and scores assistant positions, excluding role headers and turn terminators\. Chat rendering disables thinking and omits separately supplied reasoning fields, so fidelity is measured on fixed reference prefixes rather than on generated reasoning trajectories\. The routing study uses six batches per axis \(Appendix[E\.3](https://arxiv.org/html/2609.30465#A5.SS3)\), and its evaluated batches form a subset of this pool\. The fidelity pool has no exact\-message overlap with calibration, although near\-duplicates and benchmark contamination have not been ruled out\.

### B\.3Evaluation configuration

Table[4](https://arxiv.org/html/2609.30465#A2.T4)summarizes the four main backbones\. Qwen3\.6\-35B\-A3B uses the hybrid attention–MoE architecture described for Qwen3\.5\([Qwen Team, 2026](https://arxiv.org/html/2609.30465#bib.bib34)\)\. The GLM and DeepSeek model families are described in their technical reports\([GLM\-4\.5 Team, 2025](https://arxiv.org/html/2609.30465#bib.bib7);[DeepSeek\-AI, 2026](https://arxiv.org/html/2609.30465#bib.bib4)\)\. The configurations below refer to the evaluated checkpoints\. Hy3 is the public release\([Tencent Hy Team, 2026](https://arxiv.org/html/2609.30465#bib.bib38)\)\. The architecture and routing settings describe the language\-model decoder\. Layer counts exclude auxiliary prediction modules\. GLM\-4\.7\-Flash and Hy3 each begin with one dense feed\-forward layer\. All 43 DeepSeek\-V4\-Flash\-0731 layers are MoE layers, of which three use fixed\-hash routing and the other 40 use learned\-router selection\. Its score transformation issoftplus⁡\(z\)\\sqrt\{\\operatorname\{softplus\}\(z\)\}\. The nominal parameter counts exclude auxiliary MTP and DSpark modules\.

Both removal ratios retain the original top\-kkand shared experts\. Calibration uses the 2,048\-example source pool, while the scored coverage in Appendix[B\.1](https://arxiv.org/html/2609.30465#A2.SS1)refers specifically to the Qwen3\.6\-35B\-A3B and GLM\-4\.7\-Flash diagnostic collection\. For downstream benchmarks, we use temperature00, except for SWE\-bench Verified, where we use0\.70\.7\. AIME’26 reports avg@8, and the other eight benchmarks report avg@3\. For each benchmark, sampling parameters are shared across all four backbones\. The maximum input length is 224K tokens for GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B, and 128K tokens for DeepSeek\-V4\-Flash\-0731 and Hy3\. All four use a 32K\-token output limit\. These are evaluation limits, not the models’ maximum context capacities\. The Qwen3\.6\-35B\-A3B and GLM\-4\.7\-Flash benchmark settings enable thinking\. Fidelity and response diversity use their own protocols, with generation analysis using temperature0\.70\.7\. Cross\-study comparability is discussed in Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)\.

Within a study, every scoring\-rule comparison matches calibration inputs and layerwise budgets\. The two limits stated at the start of this appendix apply to all of these settings\.

Table 4:Architectures and pruning settings of the four main backbones\. Parameter counts are nominal backbone sizes, and expert counts are per MoE layer\.
### B\.4Metric definitions

Letptp\_\{t\}andqtq\_\{t\}be the unpruned and pruned next\-token distributions on the same reference prefix\. Reverse KL is

DKL\(qt∥pt\)=∑vqt\(v\)logqt​\(v\)pt​\(v\)\.D\_\{\\mathrm\{KL\}\}\(q\_\{t\}\\\|p\_\{t\}\)=\\sum\_\{v\}q\_\{t\}\(v\)\\log\\frac\{q\_\{t\}\(v\)\}\{p\_\{t\}\(v\)\}\.\(13\)The main fidelity sweep weights assistant tokens within each evaluation axis and balances the eight axes equally\.

For reference tokenyty\_\{t\}, defineℓp​\(t\)=−log⁡pt​\(yt\)\\ell\_\{p\}\(t\)=\-\\log p\_\{t\}\(y\_\{t\}\)andℓq​\(t\)=−log⁡qt​\(yt\)\\ell\_\{q\}\(t\)=\-\\log q\_\{t\}\(y\_\{t\}\)\. Averaging these losses under the specified evaluation weights gives reference\-token NLL and

PPLp\\displaystyle\\mathrm\{PPL\}\_\{p\}=exp\(NLLp\),PPLq=exp\(NLLq\),\\displaystyle=\\exp\(\\mathrm\{NLL\}\_\{p\}\),\\qquad\\mathrm\{PPL\}\_\{q\}=\\exp\(\\mathrm\{NLL\}\_\{q\}\),\(14\)Δ​NLL\\displaystyle\\Delta\\mathrm\{NLL\}=NLLq−NLLp,PPLqPPLp−1=exp\(ΔNLL\)−1\.\\displaystyle=\\mathrm\{NLL\}\_\{q\}\-\\mathrm\{NLL\}\_\{p\},\\qquad\\frac\{\\mathrm\{PPL\}\_\{q\}\}\{\\mathrm\{PPL\}\_\{p\}\}\-1=\\exp\(\\Delta\\mathrm\{NLL\}\)\-1\.The final quantity is excess perplexity\. These likelihood metrics use observed reference tokens, not samples fromptp\_\{t\}orqtq\_\{t\}\. NegativeΔ​NLL\\Delta\\mathrm\{NLL\}indicates higher reference\-token likelihood under the pruned model, not necessarily better downstream capability\.

For normalized selected routing weights andk\>1k\>1, routing entropy is

Hr​\(x\)=−∑j∈S⁡\(x\)wj​\(x\)​log⁡wj​\(x\)log⁡k\.H\_\{r\}\(x\)=\-\\frac\{\\sum\_\{j\\in S\(x\)\}w\_\{j\}\(x\)\\log w\_\{j\}\(x\)\}\{\\log k\}\.\(15\)SmallerHrH\_\{r\}means more concentrated routing\. Predictive entropy instead summarizes the vocabulary distribution,

H\(pt\)=−∑vpt\(v\)logpt\(v\),ΔHt=H\(qt\)−H\(pt\)\.H\(p\_\{t\}\)=\-\\sum\_\{v\}p\_\{t\}\(v\)\\log p\_\{t\}\(v\),\\qquad\\Delta H\_\{t\}=H\(q\_\{t\}\)\-H\(p\_\{t\}\)\.\(16\)Absolute entropy drift averages predictive entropies within routing cells and across valid layers, takes the absolute pruned–original difference in each batch–decile cell, and then averages these differences\. Taking the absolute value after cell aggregation distinguishes this measure from the mean per\-token\|Δ​Ht\|\|\\Delta H\_\{t\}\|\. Predictive entropy also differs from reference\-token NLL, so its exponential is not the PPL reported here\. Appendix[E](https://arxiv.org/html/2609.30465#A5)specifies the sampling units and uncertainty estimates for each output analysis\.

### B\.5Downstream benchmark protocol and scope

The benchmark evaluates four backbones at two removal ratios each\. Table[1](https://arxiv.org/html/2609.30465#S3.T1)covers GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B, and Table[2](https://arxiv.org/html/2609.30465#S3.T2)covers DeepSeek\-V4\-Flash\-0731 and Hy3\. The full method comparison includes Frequency, EAN, REAP, and all three residual scores on GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B\. RCS, RCS\-LOO, andRazoruse conditional RMS\. DeepSeek\-V4\-Flash\-0731 and Hy3 compare the three residual scores against their unpruned references only, so they enter the RCS\-LOO\-relative and unpruned\-relative benchmark comparisons rather than the REAP\-relative ranges and win counts\.

##### Baseline scores\.

Forni=\|𝒟i\|\>0n\_\{i\}=\|\\mathcal\{D\}\_\{i\}\|\>0routed calibration tokens, the baseline scores are, up to common positive layer scales that do not affect ranking,

siFrequency=ni,siEAN=1ni​∑x∈𝒟i‖fi​\(x\)‖2,siREAP=1ni​∑x∈𝒟iwi​\(x\)​‖fi​\(x\)‖2\.s\_\{i\}^\{\\mathrm\{Frequency\}\}=n\_\{i\},\\qquad s\_\{i\}^\{\\mathrm\{EAN\}\}=\\frac\{1\}\{n\_\{i\}\}\\sum\_\{x\\in\\mathcal\{D\}\_\{i\}\}\\\|f\_\{i\}\(x\)\\\|\_\{2\},\\qquad s\_\{i\}^\{\\mathrm\{REAP\}\}=\\frac\{1\}\{n\_\{i\}\}\\sum\_\{x\\in\\mathcal\{D\}\_\{i\}\}w\_\{i\}\(x\)\\\|f\_\{i\}\(x\)\\\|\_\{2\}\.\(17\)Frequency counts hard routing selections, as in frequency\-based pruning\([Muzio et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib30);[Jaiswal et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib12)\)\. EAN is our conditional\-mean adaptation of activation\-norm pruning\([Jaiswal et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib12);[Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15)\)\. We divide by each expert’s routed\-token count, whereas the accumulated\-norm EAN in REAP’s baseline definition does not\. This expert\-specific normalization is not a common rescaling and can change the ranking\. REAP follows the routed\-token conditional mean in its original definition\([Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15)\)\. All three scores are computed without recovery training, and zero\-observation experts receive zero score\.

##### Mathematics and instruction following\.

The nine tasks cover six capability categories, each scoring a multi\-step solution rather than a single prediction\. AIME’26 uses the 2026 American Invitational Mathematics Examination\([Mathematical Association of America, n\.d\.](https://arxiv.org/html/2609.30465#bib.bib28)\)to assess competition mathematics\. IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2609.30465#bib.bib44)\)checks compliance with programmatically verifiable instructions\. IFBench\([Pyatkin et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib33)\)extends this evaluation to additional verifiable constraints, testing instruction\-following generalization beyond IFEval\.

##### Knowledge, tools, and long context\.

SuperGPQA\([Du et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib5)\)evaluates graduate\-level knowledge across 285 disciplines\. The Berkeley Function Calling Leaderboard, version 4 \(BFCL v4,[Patil et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib32);[Mao et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib27)\), evaluates function calling and agentic tool use\. LongBench v2\([Bai et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib1)\)tests comprehension and reasoning over long contexts\.

##### Coding tasks and versions\.

The three coding tasks cover function\-level generation, competition programming, and repository\-level issue resolution\. HumanEval\+\([Liu et al\., 2023](https://arxiv.org/html/2609.30465#bib.bib23)\)uses EvalPlus’s expanded tests to assess the functional correctness of generated Python functions\. LiveCodeBench\([Jain et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib11)\)draws from time\-indexed programming competitions\. The main benchmark uses its 2026 evaluation\-time latest version, abbreviated LCB in the tables, which is distinct from the 175\-question subset used for the generation diagnostics \(Appendix[E\.5](https://arxiv.org/html/2609.30465#A5.SS5.SSS0.Px1)\)\. SWE\-bench Verified\([Jimenez et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib14);[Chowdhury et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib3)\), abbreviated SWE, is the 500\-issue subset of SWE\-bench screened by software developers for well\-specified issues and appropriately scoped tests\. The reported benchmark designation is Verified for all four backbones, but the retained score summaries do not independently establish the evaluated issue sets\. Scores report the percentage of resolved issues averaged over three evaluations\.

##### Score aggregation\.

Instruction\-following and coding averages give equal weight to their two and three tasks, respectively, while Overall weights all nine equally\. Averages and paired macro gaps are computed from task scores rather than rounded displayed means\. The recorded task\-score table yields a 33/0/3 win/tie/loss count for RCS\-LOO against REAP and 36/0/0 forRazoracross GLM\-4\.7\-Flash and Qwen3\.6\-35B\-A3B at two budgets\. Three properties qualify these counts\. They are arithmetic summaries of recorded task scores, and the current artifact holds no raw evaluator reports for an independent rederivation\. Both sets of 36 comparisons reuse one task suite rather than replicate it, and neither extends to the two backbones without a REAP benchmark run\. The benchmark comparisons report no task\-level uncertainty\. AIME’26 reports avg@8 and the other eight benchmarks report avg@3\. The three residual criteria and REAP use matchedRazorCalinputs, whereas REAP⋆is an external GLM\-4\.7\-Flash checkpoint calibrated with 24,576 samples\.

##### Additional task contrasts\.

On Hy3,Razorscores 63\.4 and 59\.5 at 25% and 50% removal, respectively, versus 63\.7 unpruned\. On DeepSeek\-V4\-Flash\-0731, the corresponding scores are 61\.0 and 55\.6 versus 59\.2 unpruned \(Table[2](https://arxiv.org/html/2609.30465#S3.T2)\)\. The macro ordering is neither uniform across tasks nor monotonic across scoring refinements\. At 25% on Qwen3\.6\-35B\-A3B, RCS slightly exceeds RCS\-LOO in macro average \(61\.0 versus 60\.8\)\. At 50% on GLM\-4\.7\-Flash, EAN exceedsRazoron LCB \(33\.9 versus 29\.8\)\.

For RCS\-LOO versus REAP, both backbones carrying that comparison show larger macro gaps at 50% than at 25% \(0\.80 to 3\.01 on GLM\-4\.7\-Flash and 1\.02 to 1\.96 on Qwen3\.6\-35B\-A3B\), leaving open whether the widening is general\. On DeepSeek\-V4\-Flash\-0731 at 25%, RCS\-LOO gains on SWE \(12\.4 to 18\.8\) but loses on HE\+ \(91\.5 to 89\.6\) relative to RCS, the nearest measured criterion on that backbone\. At 50%, SWE suffers the largest proportional loss from the original for RCS\-LOO on both Hy3 and DeepSeek\-V4\-Flash\-0731\. Thus even favorable macro comparisons coexist with capability\-specific trade\-offs\.

##### Comparability\.

Benchmark and diagnostic results are compared within a study rather than across studies, since the two use separate checkpoints and calibration draws even when they share a scoring criterion\. On DeepSeek\-V4\-Flash\-0731, where the routing study matches fixed\-hash selections across criteria, its benchmark gap therefore reflects the full pruned model rather than learned\-router saliency alone\. The “2026 latest” LiveCodeBench designation tracks a rolling contest window rather than an immutable snapshot, so absolute LiveCodeBench scores are comparable across the methods evaluated here but not against externally reported numbers\. Input and output limits and shared sampling settings appear in Appendix[B\.3](https://arxiv.org/html/2609.30465#A2.SS3)\.

## Appendix CScoring Proxies and Expert Selection

A scoring rule makes two choices that the main text reports jointly, namely which quantity to measure at a routed token and how to reduce that quantity across tokens\. This appendix separates them\. Appendices[C\.2](https://arxiv.org/html/2609.30465#A3.SS2)through[C\.5](https://arxiv.org/html/2609.30465#A3.SS5)vary one choice at a time and find that neither is dominant, since conditional RMS beats the conditional mean for most token quantities but not all, and the residual reference beats the output\-norm reference under some aggregations but not others\. Appendix[C\.7](https://arxiv.org/html/2609.30465#A3.SS7)then asks what the winning scores change in the retained set itself, through three probes that measure selection consequences rather than score values\.

The results below come from four evaluation collections that are never pooled, so comparisons are read within a collection rather than across them\. Each result states which one it comes from\.

### C\.1Scoring criteria and budget comparisons

Table[5](https://arxiv.org/html/2609.30465#A3.T5)lists the three consensus\-residual scoring rules and the magnitude baseline REAP\-RMS, all with conditional RMS, alongside their experimental coverage\.

Table 5:Scoring criteria and experimental coverage\.Each criterion fixes a token quantity, and every criterion here conditions its aggregation on tokens routed to the expert\. RCS uses the routing\-weighted consensus residual, RCS\-LOO adds survivor renormalization, and RCS\-Refill additionally models the promoted expert\. The token quantities of RCS\-LOO and RCS\-Refill are exact under their respective deletion assumptions, yet the fixed\-support identity fixes only the combination of consensus residual and LOO factor\. It leaves both the aggregation rule and the performance of the selected checkpoint open, so the three rules need not form a performance hierarchy\.

The magnitude comparators differ in reference rather than in aggregation\. REAP\-RMS uses the output norm and differs from REAP solely in replacing the conditional mean with RMS, isolating aggregation from the reference change\. Our EAN adaptation averages‖fi‖2\\\|f\_\{i\}\\\|\_\{2\}and REAP averageswi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}, with source definitions and the EAN normalization difference in Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)\. The REAP\-RMS results belong to the factorial collection of Appendix[C\.5](https://arxiv.org/html/2609.30465#A3.SS5), whose Table[9](https://arxiv.org/html/2609.30465#A3.T9)crossesfif\_\{i\}againstfi−cf\_\{i\}\-candwiw\_\{i\}againstwi/\(1−wi\)w\_\{i\}/\(1\-w\_\{i\}\)over the available sum, mean, and RMS reductions, with refill damage as an additional quantity\.

Table[6](https://arxiv.org/html/2609.30465#A3.T6)compares REAP, RCS, RCS\-LOO, andRazorunder the same checkpoint\-evaluation protocol\. Moving from REAP to RCS lowers KL on both models, but changes both the reference and mean\-to\-RMS aggregation, so it does not isolate consensus subtraction\. The matched mean\-aggregation contrast appears in Appendix[C\.6](https://arxiv.org/html/2609.30465#A3.SS6)\. Adding the LOO denominator raises KL on Qwen3\.6\-35B\-A3B, while adding refill lowers it in three of four settings\. These estimates pool tokens within each axis and weight the eight axes equally\. TheΔ\\Deltavs REAP column gives RCS\-LOO’s relative change from REAP, while refill gain givesRazor’s relative reduction from RCS\-LOO\. Negative values favor RCS\-LOO in both columns, and percentages use unrounded KL\.

Table 6:Reverse KL across scoring rules in nats \(lower is better\)\.
### C\.2RMS aggregation and score distributions

For routed calibration tokens𝒟i=\{x∈𝒟:i∈S⁡\(x\)\}\\mathcal\{D\}\_\{i\}=\\\{x\\in\\mathcal\{D\}:i\\in S\(x\)\\\},Razoruses the conditional root mean square of refill damage,

si=1\|𝒟i\|​∑x∈𝒟i\(‖wi​\(x\)​ri​\(x\)−wr​\(x\)​rr​\(x\)‖21−wi​\(x\)\+wr​\(x\)\)2\.s\_\{i\}=\\sqrt\{\\frac\{1\}\{\|\\mathcal\{D\}\_\{i\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\_\{i\}\}\\left\(\\frac\{\\\|w\_\{i\}\(x\)r\_\{i\}\(x\)\-w\_\{r\}\(x\)r\_\{r\}\(x\)\\\|\_\{2\}\}\{1\-w\_\{i\}\(x\)\+w\_\{r\}\(x\)\}\\right\)^\{2\}\}\.\(18\)This definition applies when\|𝒟i\|\>0\|\\mathcal\{D\}\_\{i\}\|\>0, with unseen experts handled as specified in Appendix[A\.4](https://arxiv.org/html/2609.30465#A1.SS4)\. The token damage follows from Proposition[2](https://arxiv.org/html/2609.30465#Thmproposition2), but conditioning and RMS are aggregation choices, not consequences of that identity\. The same aggregation is applied towi​‖ri‖2w\_\{i\}\\\|r\_\{i\}\\\|\_\{2\}for RCS andδiloo\\delta\_\{i\}^\{\\mathrm\{loo\}\}for RCS\-LOO\.

RMS differs from the conditional mean by an explicit variance term, which is what gives larger observed shifts more weight\. With empirical conditional meanμi\\mu\_\{i\}and population standard deviationσi\\sigma\_\{i\}of token\-level damage over𝒟i\\mathcal\{D\}\_\{i\},

si=μi2\+σi2=μi1\+CVi2,CVi=σi/μi\(μi\>0\)\.s\_\{i\}=\\sqrt\{\\mu\_\{i\}^\{2\}\+\\sigma\_\{i\}^\{2\}\}=\\mu\_\{i\}\\sqrt\{1\+\\mathrm\{CV\}\_\{i\}^\{2\}\},\\qquad\\mathrm\{CV\}\_\{i\}=\\sigma\_\{i\}/\\mu\_\{i\}\\quad\(\\mu\_\{i\}\>0\)\.\(19\)Forμi<μj\\mu\_\{i\}<\\mu\_\{j\}, RMS ranksiiabovejjexactly whenσi2−σj2\>μj2−μi2\\sigma\_\{i\}^\{2\}\-\\sigma\_\{j\}^\{2\}\>\\mu\_\{j\}^\{2\}\-\\mu\_\{i\}^\{2\}\. Such reversals change the retained set only across the budget boundary\. This preference for variable damage motivates RMS as an empirical design choice rather than a guarantee of better pruning\. Appendix[C\.6](https://arxiv.org/html/2609.30465#A3.SS6)tests that choice by comparing mean and RMS while holding the token quantity fixed\.

Whether such reversals occur in practice is measurable from the accumulators already collected during scoring, at no extra storage cost\. We recover each expert’s mean and standard deviation from its routed\-token count and the accumulated first and second moments of token\-level damage\. Both the fixed\-support and refill instances carry these accumulators, which lets Table[7](https://arxiv.org/html/2609.30465#A3.T7)compare RCS\-LOO andRazor\. Equal\-expert quantiles cover all observed experts without a count threshold, comprising 10,240 on Qwen3\.6\-35B\-A3B and 2,944 on GLM\-4\.7\-Flash, and the CV panel of Figure[4](https://arxiv.org/html/2609.30465#S3.F4)uses the same experts\. Within\-layer dispersion ismedianℓ⁡\[stdi⁡\(CVℓ​i\)/meani⁡\(CVℓ​i\)\]\\operatorname\{median\}\_\{\\ell\}\[\\operatorname\{std\}\_\{i\}\(\\mathrm\{CV\}\_\{\\ell i\}\)/\\operatorname\{mean\}\_\{i\}\(\\mathrm\{CV\}\_\{\\ell i\}\)\], using population standard deviations\.

What changes a ranking is the premium’s variation across experts rather than its absolute size\. The median fixed\-support premium is 1\.189 on Qwen3\.6\-35B\-A3B and 1\.155 on GLM\-4\.7\-Flash\. We quantify the resulting boundary swaps by comparing direct top\-BBmean and RMS rankings on the same experts and averaging swap counts equally across layers, with each entry reporting the RMS\-only retained count over the retained\-set size\. Qwen3\.6\-35B\-A3B has more exchanges at both budgets\. Under refill, the median premium is 1\.166 on Qwen3\.6\-35B\-A3B and 1\.153 on GLM\-4\.7\-Flash, and the swap counts stay within one expert per layer of the fixed\-support counts\. The two counterfactuals therefore show similar relative variability and similar mean–RMS selection differences, which is compatible with unequal damage magnitudes and unequal second moments between them\.

Table 7:Damage variability and mean–RMS selection differences for RCS\-LOO andRazor\.One boundary crossing shows the mechanism concretely\. We take Qwen3\.6\-35B\-A3B layer 0, the first layer with a mean–RMS disagreement at 25% removal in the 4,096\-token\-row diagnostic, and within it the opposite\-status pair with the closest conditional means, selected without reference to downstream outcomes\. Expert 208 has mean/CV/RMS 0\.03323/0\.741/0\.04136 over 32,041 routed tokens, against 0\.03359/0\.637/0\.03982 over 23,766 tokens for expert 38\. Their mean\-to\-RMS ranks move from 194 to 190 and from 189 to 201, respectively, crossing the 192\-expert retention boundary\. Higher variability thus flips selection across a conditional\-mean margin of0\.000360\.00036, for one pair whose utility and presence in the benchmark checkpoints are not established here\.

### C\.3Conditional RMS, mean, and corpus\-sum reductions

The previous subsection compared conditional RMS against the conditional mean, which isolates the second moment\. A second and independent choice remains, namely whether to condition on expert use at all, and the two are easily conflated because both appear to make RMS “weight large damage more\.” Separating them requires a third reduction that drops the conditioning\. Letgig\_\{i\}be any non\-negative token quantity with conditional RMSsis\_\{i\}, letπi=\|𝒟i\|/\|𝒟\|\\pi\_\{i\}=\|\\mathcal\{D\}\_\{i\}\|/\|\\mathcal\{D\}\|, and setgig\_\{i\}to zero on tokens not routed toii\. The empirical average squared damage over*all*calibration tokens is

1\|𝒟\|∑x∈𝒟𝟏\{i∈S\(x\)\}gi\(x\)2=πisi2\.\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\}\\mathbf\{1\}\\\{i\\in S\(x\)\\\}g\_\{i\}\(x\)^\{2\}=\\pi\_\{i\}s\_\{i\}^\{2\}\.\(20\)An all\-token RMS criterion would therefore rankπi​si\\sqrt\{\\pi\_\{i\}\}s\_\{i\}, which differs from rankingsis\_\{i\}by the exposure factorπi\\sqrt\{\\pi\_\{i\}\}\. Rankingsis\_\{i\}measures severity conditional on use, so an infrequent but severe deletion is not discounted for being rare, at the cost of minimizing neither total calibration error nor damage to capabilities the calibration set never exercises\.

To isolate aggregation, we fix the RCS token quantitygi​\(t\)=wi​\(t\)​‖fi​\(t\)−c⁡\(t\)‖2g\_\{i\}\(t\)=w\_\{i\}\(t\)\\\|f\_\{i\}\(t\)\-c\(t\)\\\|\_\{2\}and compare three reductions of the observed damage at four pruning budgets,

conditional RMS\\displaystyle\\text\{conditional RMS\}1ni​∑t∈𝒟igi​\(t\)2,\\displaystyle\\sqrt\{\\tfrac\{1\}\{n\_\{i\}\}\\textstyle\\sum\_\{t\\in\\mathcal\{D\}\_\{i\}\}g\_\{i\}\(t\)^\{2\}\},conditional mean\\displaystyle\\qquad\\text\{conditional mean\}1ni​∑t∈𝒟igi​\(t\),\\displaystyle\\tfrac\{1\}\{n\_\{i\}\}\\textstyle\\sum\_\{t\\in\\mathcal\{D\}\_\{i\}\}g\_\{i\}\(t\),\(21\)corpus sum\\displaystyle\\text\{corpus sum\}∑t∈𝒟gi​\(t\)2=\|𝒟\|​πi​si2,\\displaystyle\\textstyle\\displaystyle\\sum\_\{t\\in\\mathcal\{D\}\}g\_\{i\}\(t\)^\{2\}=\|\\mathcal\{D\}\|\\,\\pi\_\{i\}s\_\{i\}^\{2\},where𝒟i\\mathcal\{D\}\_\{i\}contains tokens routed to expertii,ni=\|𝒟i\|n\_\{i\}=\|\\mathcal\{D\}\_\{i\}\|, andgi​\(t\)=0g\_\{i\}\(t\)=0elsewhere\. The corpus sum differs from the all\-token mean squared damage in Eq\.[20](https://arxiv.org/html/2609.30465#A3.E20)only by the common factor\|𝒟\|\|\\mathcal\{D\}\|, so they induce the same ranking\. We exclude selection frequencynin\_\{i\}because it discardsgig\_\{i\}entirely\. Including it would conflate a change in token quantity with a change in aggregation\. Table[1](https://arxiv.org/html/2609.30465#S3.T1)compares frequency as a criterion in its own right\.

For the RCS quantity, reverse KL follows the strict orderingRMS<mean<sum\\text\{RMS\}<\\text\{mean\}<\\text\{sum\}in all eight model–budget settings \(Figure[6](https://arxiv.org/html/2609.30465#A3.F6)\)\. Both choices thus point the same way under this protocol, as conditioning on use beats weighting squared severity by exposure, and squaring within the conditional reduction beats not squaring\. The all\-token quantity is not thereby defective, since it targets total calibration error rather than per\-use severity\.

How much the reduction matters is easier to appreciate against the spread among criteria, which the main text reports as the headline comparison\. Between the extreme reductions, KL spans2\.4×2\.4\\timesand1\.7×1\.7\\timeson Qwen3\.6\-35B\-A3B and1\.15×1\.15\\timesand1\.05×1\.05\\timeson GLM\-4\.7\-Flash at 25% and 50% removal\. The corresponding spans among REAP, RCS, and RCS\-LOO in Table[6](https://arxiv.org/html/2609.30465#A3.T6)are1\.12×1\.12\\times,1\.13×1\.13\\times,1\.12×1\.12\\times, and1\.17×1\.17\\times\. The reduction span is thus larger on Qwen3\.6\-35B\-A3B, comparable on GLM\-4\.7\-Flash at 25%, and smaller there at 50%\. On one backbone, in other words, how the damage is aggregated matters more than which damage quantity is aggregated, which is why the factorial grid of Appendix[C\.5](https://arxiv.org/html/2609.30465#A3.SS5)crosses the two rather than varying them in sequence\.

The ordering largely carries over to the two exact deletion quantities\. These comparisons use equal layerwise budgets on the same four shards as Table[6](https://arxiv.org/html/2609.30465#A3.T6), applying the two conditional reductions to the leave\-one\-out and refill residuals\. For refill damage, RMS lowers reverse KL relative to the conditional mean at all four GLM\-4\.7\-Flash budgets, by 4\.4%–11\.4%, and at 12\.5% and 25% on Qwen3\.6\-35B\-A3B, while raising it at 37\.5% and 50% on Qwen3\.6\-35B\-A3B \(0\.20744 versus 0\.18874, and 0\.44397 versus 0\.44143\), for six of eight settings improved\. Applied to the leave\-one\-out residual, RMS lowers KL in seven of eight settings, failing only at 37\.5% on Qwen3\.6\-35B\-A3B \(0\.20504 versus 0\.19196\)\. The advantage of RMS is therefore robust to which exact deletion quantity it reduces, but it is not uniform in the removal budget\.

Figure 6:Reverse KL under different calibration reductions \(lower is better\)\.The RCS quantity uses conditional RMS, conditional mean, and corpus sum\. The RCS\-LOO and RCS\-Refill quantities each use the two conditional reductions\. Selection frequency is excluded because it discards the token quantity rather than reducing it\.
### C\.4What the refill term changes

Refill is the one component whose effect on the benchmark is positive on both backbones, so it is worth asking what it changes in the scores themselves\. Table[8](https://arxiv.org/html/2609.30465#A3.T8)answers this from the same accumulated calibration moments as Table[7](https://arxiv.org/html/2609.30465#A3.T7), and the answer is that refill changes score scale in opposite directions on the two models while leaving the shape of each expert’s damage profile nearly intact\. Router averages there are token\-weighted over routed events, whereas score ratios and token\-axis cosines use unweighted expert quantiles\. The routed\-event\-weighted meanwi=1/kw\_\{i\}=1/kis a renormalization consistency check, since each token’skkselected weights sum to one\.

##### Score scale and residual geometry\.

The median refill\-to\-fixed\-support score ratio corresponds to a2\.81%2\.81\\%increase on Qwen3\.6\-35B\-A3B and a2\.56%2\.56\\%decrease on GLM\-4\.7\-Flash\. Refill raises scores for85\.2%85\.2\\%of Qwen3\.6\-35B\-A3B experts and14\.3%14\.3\\%of GLM\-4\.7\-Flash experts\. Substituting the reported mean routing weights into\(1−wi\)/\(1−wi\+wr\)\(1\-w\_\{i\}\)/\(1\-w\_\{i\}\+w\_\{r\}\)gives0\.9220\.922and0\.8500\.850, respectively\. These values illustrate the denominator’s attenuation at average weights\. They are neither token\-averaged factors nor a decomposition of the median score ratio\.

The mean residual cosines are0\.1080\.108on Qwen3\.6\-35B\-A3B and0\.1030\.103on GLM\-4\.7\-Flash\. By Eq\.[11](https://arxiv.org/html/2609.30465#A1.E11), such a low positive cosine is compatible with either an increase or a decrease in the numerator, depending on the relative residual magnitudeρi\\rho\_\{i\}\. Reconstructingρi\\rho\_\{i\}would require the promoted residual norms, which these accumulators do not retain\. The averages also cannot be combined to recover the nonlinear factors, because the score ratios summarize experts whereas the routing statistics summarize routed events\.

##### Damage\-profile similarity and selection\.

Treatingδiloo\\delta\_\{i\}^\{\\mathrm\{loo\}\}andδi\\delta\_\{i\}as vectors over each expert’s routed tokens, their cosine has median0\.99420\.9942on Qwen3\.6\-35B\-A3B and0\.99520\.9952on GLM\-4\.7\-Flash, with fifth percentiles0\.96300\.9630and0\.97150\.9715\. Most profiles are therefore close in direction, although the minima of0\.70620\.7062and0\.54400\.5440show substantial changes for some experts\. Directional similarity alone does not fix the ranking, since nearly proportional profiles with different scale factors still reorder experts\. The checkpoint outcome also runs against the score summaries\. Despite the opposite median score shifts on the two models, refill lowers KL at both Qwen3\.6\-35B\-A3B budgets and at only one GLM\-4\.7\-Flash budget, so these local statistics do not predict its contribution to checkpoint quality\.

Table 8:Calibration statistics of routing weights and refill\-induced score changes\.

### C\.5Reference, routing factor, and aggregation

Table[9](https://arxiv.org/html/2609.30465#A3.T9)reports fourteen scoring configurations for two models at 25% and 50% removal\. Its core grid crosses two output references with two routing factors\. Rows within each block compare token quantities, while blocks compare reductions of the same quantity\. The conditional mean ofwi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}is REAP, and the conditional RMS ofwi1−wi​‖fi−c‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\-c\\\|\_\{2\}is RCS\-LOO\. Adding refill damage gives five token quantities across the corpus\-sum∑tgi​\(t\)2\\sum\_\{t\}g\_\{i\}\(t\)^\{2\}, conditional\-mean, and conditional\-RMS blocks, giving fourteen rows per model\.

One row is not exactly matched to the rest\. The conditional\-mean refill row comes from the main sweep of Figure[6](https://arxiv.org/html/2609.30465#A3.F6)rather than this grid’s evaluation collection, so comparisons involving it are cross\-collection contrasts, unlike the matched comparisons among the other thirteen rows\.

Table 9:Fourteen scoring configurations varying reference, routing factor, and aggregation\.Reverse KL is in nats at 25% and 50% removal, with lower values better\. Thirteen rows share four evaluation shards\. The conditional\-mean refill row comes from the main sweep in Figure[6](https://arxiv.org/html/2609.30465#A3.F6), which differs in one shard\. Comparisons involving this row are not fully matched\. Sum denotes∑tgi​\(t\)2\\sum\_\{t\}g\_\{i\}\(t\)^\{2\}\. No corpus\-sum result is available forwi1−wi​‖fi‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\\\|\_\{2\}\. Refill damage with RMS isRazor\.Token quantityAggregateQwen3\.6\-35B\-A3BGLM\-4\.7\-Flash25%50%25%50%wi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}Sum0\.147120\.677250\.223090\.89425wi​‖fi−c‖2w\_\{i\}\\\|f\_\{i\}\-c\\\|\_\{2\}Sum0\.160050\.671730\.235130\.96880wi1−wi​‖fi−c‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\-c\\\|\_\{2\}Sum0\.176590\.692940\.233390\.91110‖wi​ri−wr​rr‖21−wi\+wr\\frac\{\\\|w\_\{i\}r\_\{i\}\-w\_\{r\}r\_\{r\}\\\|\_\{2\}\}\{1\-w\_\{i\}\+w\_\{r\}\}Sum0\.187980\.702640\.231760\.88086wi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}Mean0\.074690\.456910\.228521\.01692wi1−wi​‖fi‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\\\|\_\{2\}Mean0\.078820\.489750\.229220\.97162wi​‖fi−c‖2w\_\{i\}\\\|f\_\{i\}\-c\\\|\_\{2\}Mean0\.073570\.440750\.228040\.94483wi1−wi​‖fi−c‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\-c\\\|\_\{2\}Mean0\.071510\.453410\.230170\.95748‖wi​ri−wr​rr‖21−wi\+wr\\frac\{\\\|w\_\{i\}r\_\{i\}\-w\_\{r\}r\_\{r\}\\\|\_\{2\}\}\{1\-w\_\{i\}\+w\_\{r\}\}Mean0\.073860\.441430\.224810\.96456wi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}RMS0\.068690\.439480\.208130\.86866wi1−wi​‖fi‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\\\|\_\{2\}RMS0\.095350\.646400\.208650\.90296wi​‖fi−c‖2w\_\{i\}\\\|f\_\{i\}\-c\\\|\_\{2\}RMS0\.066740\.403550\.204830\.92318wi1−wi​‖fi−c‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\-c\\\|\_\{2\}RMS0\.071440\.451340\.203530\.87107‖wi​ri−wr​rr‖21−wi\+wr\\frac\{\\\|w\_\{i\}r\_\{i\}\-w\_\{r\}r\_\{r\}\\\|\_\{2\}\}\{1\-w\_\{i\}\+w\_\{r\}\}RMS0\.068740\.444090\.204610\.85460##### Evaluation and aggregation\.

Within each model, thirteen of the fourteen rows share four evaluation shards with matched domain sequences and token counts\. LetLs,dL\_\{s,d\}denote mean reverse KL in shardssand domaindd, withns,dn\_\{s,d\}scored tokens\. The estimator pools tokens within domains and then weights the eight domains equally,

L=18​∑d=18∑s=03ns,d​Ls,d∑s=03ns,d\.L=\\frac\{1\}\{8\}\\sum\_\{d=1\}^\{8\}\\frac\{\\sum\_\{s=0\}^\{3\}n\_\{s,d\}L\_\{s,d\}\}\{\\sum\_\{s=0\}^\{3\}n\_\{s,d\}\}\.\(22\)The entries are shard\-based point estimates, not estimates of variability across independent runs\.

##### Relation to the matched component study\.

This collection is not identical to the main sweep\. For their shared criteria, per\-domain reverse KL agrees exactly in three of four shards\. The remaining Qwen3\.6\-35B\-A3B shard has 120 batches and 529,093 tokens here, compared with 119 batches and 532,523 tokens in the main sweep\. Both GLM\-4\.7\-Flash collections have 121 batches in that shard, but contain 528,964 and 530,090 tokens, respectively\. We therefore analyze the collections separately\. Their overlap also precludes treating these results as an independent replication\.

##### No single scoring component is uniformly beneficial\.

The grid’s overall pattern is that each component’s effect is contingent on the others, so the three paragraphs below report the contingency for the routing factor, the reduction, and the output reference in turn\. Even with mean aggregation fixed, the effect of the routing factor depends on the output reference\. On GLM\-4\.7\-Flash at 50% removal, replacingwiw\_\{i\}withwi/\(1−wi\)w\_\{i\}/\(1\-w\_\{i\}\)lowers reverse KL from 1\.01692 to 0\.97162 for output magnitudes, but raises it from 0\.94483 to 0\.95748 for consensus residuals\. For fixed\-support deletion damage, RMS lowers reverse KL relative to the conditional mean in all four displayed model–budget settings, although the Qwen3\.6\-35B\-A3B differences are small\.

##### The reduction ordering depends on the token quantity\.

The grid reproduces the RCS orderingRMS<mean<sum\\text\{RMS\}<\\text\{mean\}<\\text\{sum\}from Appendix[C\.3](https://arxiv.org/html/2609.30465#A3.SS3)at both displayed budgets on both models, but this ordering does not hold for every quantity\. Conditional RMS has the lowest KL in 15 of the 16 cells with all three reductions, including four cross\-collection comparisons involving mean refill\. The exception is refill at 50% removal on Qwen3\.6\-35B\-A3B\. On GLM\-4\.7\-Flash, corpus sum yields lower KL than conditional mean for output magnitudes at both budgets \(0\.22309 versus 0\.22852, and 0\.89425 versus 1\.01692\) and for both residual quantities at 50% removal \(0\.91110 versus 0\.95748, and 0\.88086 versus 0\.96456\), reversing the middle two terms in four of the sixteen cells\. On Qwen3\.6\-35B\-A3B, mean yields lower KL than sum in all eight available comparisons, with sum yielding 1\.48–2\.55×\\timesthe KL of mean\. RMS beats sum in all sixteen matched comparisons, but the mean–sum ordering depends on the quantity and backbone\.

##### The reference and the reduction interact\.

For the four non\-refill quantities in the core grid, RMS lowers reverse KL relative to mean in 14 of the 16 quantity–model–budget cells\. Both exceptions are the routing factor without the consensus reference,wi1−wi​‖fi‖2\\frac\{w\_\{i\}\}\{1\-w\_\{i\}\}\\\|f\_\{i\}\\\|\_\{2\}, on Qwen3\.6\-35B\-A3B, where RMS raises KL from 0\.07882 to 0\.09535 at 25% removal and from 0\.48975 to 0\.64640 at 50%\. Read in the other direction, the benefit of consensus subtraction depends on which reduction is in use\. With the routing factor held atwiw\_\{i\}, subtracting the consensus lowers KL in all four settings under the mean but in three under RMS, raising it from 0\.86866 to 0\.92318 on GLM\-4\.7\-Flash at 50% removal\. With the leave\-one\-out factor, it lowers KL in all four settings under RMS but in three under the mean, raising it from 0\.22922 to 0\.23017 on GLM\-4\.7\-Flash at 25% removal\. The two components are therefore not separable additive improvements\.

##### An aggregation\-matched baseline narrows the residual advantage\.

Comparing our criteria against REAP confounds the reference with the reduction, so the grid also aggregates REAP’s token quantity by conditional RMS\. Under that matched baseline, REAP\-RMS gives lower KL than RCS\-LOO in both Qwen3\.6\-35B\-A3B settings and slightly lower KL on GLM\-4\.7\-Flash at 50%, so the mean\-to\-RMS benefit for fixed\-support deletion damage does not carry over into a uniform advantage for the residual reference\. RCS\-Refill fares better against the same baseline\. It attains the lowest KL of all fourteen rows on GLM\-4\.7\-Flash at 50% \(0\.85460\) and beats REAP\-RMS on GLM\-4\.7\-Flash at both budgets, while trailing it on Qwen3\.6\-35B\-A3B at both budgets \(0\.06874 versus 0\.06869 and 0\.44409 versus 0\.43948\), where RCS remains lowest\. With both quantities aggregated by the conditional mean instead, refill yields lower KL than REAP in all four settings \(0\.07386 versus 0\.07469 and 0\.44143 versus 0\.45691 on Qwen3\.6\-35B\-A3B, 0\.22481 versus 0\.22852 and 0\.96456 versus 1\.01692 on GLM\-4\.7\-Flash\), although mean RCS\-LOO still edges out mean refill on Qwen3\.6\-35B\-A3B at 25% \(0\.07151\)\. These mean\-refill contrasts use the separate main\-sweep shards rather than the identical batches used elsewhere in this table\.

Taken together, the grid shows that local exactness of a token quantity does not determine its empirical ordering after joint pruning, which is the same backbone dependence visible in the main\-sweep comparison of Table[6](https://arxiv.org/html/2609.30465#A3.T6)\. This is why the choice ofRazorrests on the benchmark results rather than on these checkpoint\-fidelity contrasts, which do not identify the source of the downstream gains\.

### C\.6Matched component contrasts at two removal budgets

The factorial grid above trades exact matching for coverage\. This study makes the opposite trade\. It fixes four evaluation shards at 25% and 50% removal, pools tokens within each axis, weights the eight axes equally, and varies one scoring choice at a time with the others held fixed, so each contrast is exactly matched\. It supports the component observations in the main text\. Its results agree with the grid on the central point, namely that consensus subtraction is the one component that helps everywhere while the later refinements are backbone\-dependent\.

Consensus subtraction lowers reverse KL in every matched setting\. With routing weights and mean aggregation fixed, it does so by 1\.5%/3\.5% on Qwen3\.6\-35B\-A3B and 0\.2%/7\.1% on GLM\-4\.7\-Flash at 25%/50% removal\. With residuals and the LOO factor fixed, RMS likewise lowers KL in all four settings, by 11\.6%/9\.0% on GLM\-4\.7\-Flash but by only 0\.05%/0\.44% on Qwen3\.6\-35B\-A3B, which is too small a margin to read as separation\. Beyond these two components the picture splits by backbone\. Relative to RMS without LOO, RCS\-LOO lowers GLM\-4\.7\-Flash’s reverse KL but raises Qwen3\.6\-35B\-A3B’s at both budgets \(Table[6](https://arxiv.org/html/2609.30465#A3.T6)\)\. For refill damage, RMS changes reverse KL by−7\.0%\-7\.0\\%/\+0\.6%\+0\.6\\%on Qwen3\.6\-35B\-A3B against−9\.0%\-9\.0\\%/−11\.4%\-11\.4\\%on GLM\-4\.7\-Flash\.

The two fidelity metrics also disagree, which matters because each is a defensible summary of the same predictions\. In the RMS contrast at fixed residuals and LOO factor, RMS raises Qwen3\.6\-35B\-A3B’s excess perplexity by 16\.5%/2\.7% while lowering its reverse KL, whereas it lowers GLM\-4\.7\-Flash’s by 12\.5%/13\.0% in agreement with KL\. The refill contrast shows the same split, changing excess perplexity by−8\.5%\-8\.5\\%/\+6\.9%\+6\.9\\%on Qwen3\.6\-35B\-A3B and−8\.9%\-8\.9\\%/−13\.2%\-13\.2\\%on GLM\-4\.7\-Flash\. Appendix[E\.4](https://arxiv.org/html/2609.30465#A5.SS4)examines this disagreement directly on shared routing groups\. Here it means that no ordering of the proxies holds under both metrics at once\.

Whether the gains reach beyond calibration\-covered task types is a separate question, which we address by averaging the four covered and four held\-out axes separately with equal weight within each group, using excess perplexityexp⁡\(Δ​NLL\)−1\\exp\(\\Delta\\mathrm\{NLL\}\)\-1\. Held\-out fidelity is substantially worse for every method, as REAP’s held\-out reverse KL is2\.12\.1–2\.9×2\.9\\timesits in\-domain value across the two budgets and its held\-out excess perplexity is3\.43\.4–3\.8×3\.8\\timesits in\-domain value, so this grouping probes a different fidelity regime rather than a rescaling of the all\-axis average\. Within it, GLM\-4\.7\-Flash improves on both metrics, with RCS\-LOO reducing KL by 10\.9%/14\.3% and excess perplexity by 8\.9%/17\.1% relative to REAP, andRazorreducing them by 10\.5%/16\.0% and 8\.7%/19\.3%\. Qwen3\.6\-35B\-A3B again splits, as RCS\-LOO lowers its all\-axis KL by 4\.3%/1\.2% while changing excess perplexity by−0\.7%\-0\.7\\%/\+12\.6%\+12\.6\\%, andRazorlowers KL by 7\.9%/2\.8% while changing excess perplexity by−9\.5%\-9\.5\\%/\+6\.6%\+6\.6\\%\. Refill is the one component that helps that model on both metrics, lowering its excess perplexity by 8\.9%/5\.3% relative to RCS\-LOO\.

### C\.7Selection diagnostics

The comparisons above measure score values and checkpoint fidelity, neither of which shows what changed in the retained set\. Three probes address that question from different angles\. Conditioning on keep\-set disagreements asks whether the tokens where two criteria retain different experts are where their outputs diverge\. Joint residual statistics ask whether the residuals inside a selected deletion set cancel or reinforce\. Single\-expert ablations ask whether local score order survives deployment, once the router reselects from the remaining pool\. All three probe mechanism rather than outcome, and none measures reduced redundancy among the final retained experts, which is what the budgeted benchmark comparisons test indirectly through task utility\.

#### C\.7\.1Output differences conditioned on selection disagreements

We compare RCS\-LOO and REAP at 50% removal on 16 batches spanning eight axes for each of Qwen3\.6\-35B\-A3B and GLM\-4\.7\-Flash\. This collection is separate from the four\-shard main sweep and is subject to the checkpoint\-matching limits in Appendix[B\.5](https://arxiv.org/html/2609.30465#A2.SS5)\. It contains noRazorpredictions, so this probe is not repeated for refill\.

Each conditioning variable uses routed events from a specified layerwise keep\-set disagreement\. For routing concentration, eligible experts are retained by RCS\-LOO but not by RCS\. We select the event with the largest1/\(1−wi\)1/\(1\-w\_\{i\}\)across layers for each token and use its normalized routing entropy\. For output consensus, eligible experts are retained by residual RMS without LOO but not by REAP\. We select the event with the largest residual damage relative to the expert’s calibration mean and use itsκ=‖c‖22/∑jwj​‖fj‖22\\kappa=\\\|c\\\|\_\{2\}^\{2\}/\\sum\_\{j\}w\_\{j\}\\\|f\_\{j\}\\\|\_\{2\}^\{2\}and relative residual norm\. For relative impact, eligible experts are retained by RCS\-LOO but not by conditional\-mean fixed\-support scoring\. We take the largestδiloo/𝔼cal​\[δiloo∣i∈S\]\\delta\_\{i\}^\{\\mathrm\{loo\}\}/\\mathbb\{E\}\_\{\\mathrm\{cal\}\}\[\\delta\_\{i\}^\{\\mathrm\{loo\}\}\\mid i\\in S\]across their routed events, or zero if none is observed\. This ratio measures shift relative to an expert’s mean damage, not event rarity or preservation of rare capabilities\.

All diagnostics use model\-wide quantiles on a common finite\-token subset\. Within each bin, we average RCS\-LOO\-minus\-REAP reverse KL within batches and then equally across batches, with 5,000 paired\-batch bootstrap resamples\. The most concentrated routing, highest\-consensus, and highest\-relative\-impact bins give Qwen3\.6\-35B\-A3B and GLM\-4\.7\-Flash differences of−0\.14/−0\.33\-0\.14/\-0\.33,−0\.18/−0\.28\-0\.18/\-0\.28, and−0\.25/−0\.47\-0\.25/\-0\.47, respectively\. Tokens where disagreeing experts carry the most concentrated routing, highest consensus, or highest relative impact tend to show the largest criterion gaps\. These associations are conditional on selection disagreements and do not generalize to unconditioned tokens\.

#### C\.7\.2Interactions within the selected deletion sets

The joint identity in Eq\.[4](https://arxiv.org/html/2609.30465#S2.E4)motivates a separate test of whether cross\-expert residuals cancel or reinforce within score\-selected deletion sets\. We select sets with RCS\-LOO at 25% and 50% removal and accumulate their interaction Gram matrices on the diagnostic calibration collections of Appendix[B\.1](https://arxiv.org/html/2609.30465#A2.SS1)\. Defineui​\(x\)=wi​\(x\)​\(fi​\(x\)−c⁡\(x\)\)u\_\{i\}\(x\)=w\_\{i\}\(x\)\(f\_\{i\}\(x\)\-c\(x\)\)for routed experts and zero otherwise, and letGi​j=∑x⟨ui​\(x\),uj​\(x\)⟩G\_\{ij\}=\\sum\_\{x\}\\langle u\_\{i\}\(x\),u\_\{j\}\(x\)\\ranglewithin a layer\. For deletion set𝒜ℓ\\mathcal\{A\}\_\{\\ell\}, the interaction ratio is

ρℓ=∑i,j∈𝒜ℓGi​j∑i∈𝒜ℓGi​i\.\\rho\_\{\\ell\}=\\frac\{\\sum\_\{i,j\\in\\mathcal\{A\}\_\{\\ell\}\}G\_\{ij\}\}\{\\sum\_\{i\\in\\mathcal\{A\}\_\{\\ell\}\}G\_\{ii\}\}\.\(23\)The common layer output scale cancels in the ratio\. Values below one indicate net cancellation, while values above one indicate reinforcement\. We compute one ratio per layer and report equal\-layer quantiles over 40 Qwen3\.6\-35B\-A3B and 46 GLM\-4\.7\-Flash layers\.

At 25% and 50% removal, median ratios are 1\.080/0\.921 on Qwen3\.6\-35B\-A3B and 0\.992/0\.811 on GLM\-4\.7\-Flash\. At 50%, 29 of 40 Qwen3\.6\-35B\-A3B layers and all 46 GLM\-4\.7\-Flash layers show net cancellation\. Sets selected byRazorfrom the same accumulators give median ratios of 1\.077/0\.925 on Qwen3\.6\-35B\-A3B and 0\.997/0\.831 on GLM\-4\.7\-Flash, with cancellation in 29 of 40 and 42 of 46 layers at 50%\. The degree of cancellation is therefore similar across criteria at both budgets\. Interpreting these ratios as a quality signal is not straightforward, because the identity∑iui​\(x\)=0\\sum\_\{i\}u\_\{i\}\(x\)=0constrains the cross terms and the Gram matrices do not retain the removed routing massW𝒜​\(x\)W\_\{\\mathcal\{A\}\}\(x\)\. Without that mass, the\(1−W𝒜​\(x\)\)−2\(1\-W\_\{\\mathcal\{A\}\}\(x\)\)^\{\-2\}factor needed for the fixed\-support deletion loss is unrecoverable, and any local counterfactual where no routing mass survives is undefined\.

#### C\.7\.3Single\-expert removal with deployed routing

This probe tests whether local saliency transfers to end\-to\-end single\-expert damage under deployed routing\. We uniformly sample 32 experts and intervene on each in six layers spanning shallow, middle, and deep positions, giving 192 ablations per model\. Each intervention masks one expert’s router score to−∞\-\\inftyand reselects and renormalizes the top\-kkfrom the remaining pool, potentially admitting previously inactive experts\.

We compare the intervened modelqqwith the unpruned modelppon 16 shared batches, two per evaluation axis, using assistant\-token reverse KLDKL\(q∥p\)D\_\{\\mathrm\{KL\}\}\(q\\\|p\)\. The resulting damage estimate pools 78,091 scored tokens on Qwen3\.6\-35B\-A3B and 79,975 on GLM\-4\.7\-Flash, weighting tokens rather than axes equally\. For each scoring rule, we compute Spearman correlation over the 32 interventions within each layer and average the six correlations equally\. The saliency statistics come from the diagnostic calibration collections of Appendix[B\.1](https://arxiv.org/html/2609.30465#A2.SS1)\.

Mean correlations for RCS\-LOO, REAP, and EAN are 0\.31, 0\.34, and 0\.31 on Qwen3\.6\-35B\-A3B, and 0\.19, 0\.18, and 0\.10 on GLM\-4\.7\-Flash, respectively\. Every Qwen3\.6\-35B\-A3B correlation is positive, but several GLM\-4\.7\-Flash values are near zero and EAN is negative in two layers\. The refill scoreRazorgives 0\.31 on Qwen3\.6\-35B\-A3B and 0\.18 on GLM\-4\.7\-Flash\. The point estimates show no consistent advantage across models\. This probe measures how well local output shifts under single\-expert removal transfer to sample\-wide prediction KL after router reselection\. Better ordering of all single deletions does not guarantee better budgeted selection, and the probe accounts for neither the fixed\-support identity nor joint deletion sets\.

## Appendix DCalibration Stability and Sensitivity

Because the scores are computed from a calibration sample, two questions follow\. How much calibration data does a stable retained set require, and how much does the choice of calibration corpus matter? The two studies below answer them at different costs, which determines their differing criterion coverage\. The budget study measures overlap between expert sets selected from disjoint token subsets, needs only the per\-expert accumulators, and therefore covers REAP alongside all three family members\. The corpus study measures reference\-token PPL, needs one pruned checkpoint per configuration, and therefore covers only RCS\-LOO across 24 checkpoints\. The two answers differ in how far they can be pushed, since selection agreement rises with budget in a consistent direction whereas corpus effects do not separate cleanly from token budget and conversation structure\.

### D\.1Calibration budget and selection stability

The Qwen3\.6\-35B\-A3B collection contains 32 shards and 2,151,112 tokens\. At each budget, we draw pairs of disjoint subsets with equal shard counts, merge statistics within each subset, and select experts\. Agreement is the sum of layerwise retained\-set intersection sizes divided by the total number of retained experts across scored layers\. We report averages over four subset\-pair draws\. Because every accumulator is a plain token sum, any shard subset is a valid pack for a smaller budget, so this study requires no model forward pass and is run for all four criteria on the same subsets\. Subset agreement compares the two disjoint selections\. Full\-set agreement compares each selection with that from all 32 shards and averages the two values\. Only the former measures agreement from disjoint calibration data, because the full\-set reference includes both subsets\.

Figure 7:Calibration budget and selection agreement on Qwen3\.6\-35B\-A3B\.Subset agreement \(left pair\) compares two disjoint calibration subsets, while full\-set agreement \(right pair\) compares each against the full 32\-shard selection\. Tokens per shard count are given in the text\.Selection agreement generally increases with calibration budget for all four criteria \(Figure[7](https://arxiv.org/html/2609.30465#A4.F7)\)\. For RCS\-LOO, disjoint\-subset agreement rises from 85\.72% to 93\.39% at 50% removal as the budget grows from approximately 65,500 to 1\.08 million tokens\. At 25% removal, it rises from 92\.82% to 96\.58%\. In the final budget step, disjoint\-subset agreement increases by only 0\.01 percentage points for RCS\-LOO at 50% removal\. At 25%, the corresponding changes across criteria range from−0\.03\-0\.03to\+0\.07\+0\.07percentage points\. These small changes across four draws do not establish whether agreement has plateaued\.

Across eight calibration budgets, two removal ratios, and two agreement measures, REAP has the highest selection agreement and RCS the second highest in all 32 settings\.Razorexceeds RCS\-LOO in 31 settings\. At the smallest budget and 50% removal, disjoint\-subset agreement is 89\.50% for REAP and 85\.72% for RCS\-LOO\. At approximately 1\.08 million tokens, their disjoint\-subset gap narrows to 2\.64 percentage points, alongside a narrowing full\-set agreement gap\. The residual criteria therefore produce less stable retained sets than REAP throughout this collection, and more calibration data reduces but does not close that difference\.

This instability is worth stating plainly, since it is the one dimension on which REAP is consistently ahead\. Our results do not explain its source, and they do not connect it to the fidelity and benchmark comparisons where the residual criteria lead\. Establishing that connection would require varying calibration budget while holding the evaluation protocol fixed, which these descriptive means over four subset\-pair draws, reported without intervals or a test of the ordering, cannot support\.

### D\.2Corpus configurations and evaluation comparability

The composition study compares fullRazorCal, a tool\-filtered subset, self\-generated text, WikiText, and two complementary halves of the filtered subset at 25% and 50% removal\. Filtering retains 1,790 of 2,048 examples by excluding 249 tool\-calling and nine coding conversations with a tool\-role message before the final assistant turn\. The complementary halves combine alternating shards,\{0,2\}\\\{0,2\\\}and\{1,3\}\\\{1,3\\\}\. The filtered corpus retainsRazorCal’s mixed\-source provenance and is not a human\-authored control\. Self\-generated variants use the corresponding unpruned backbone to continue the prefix before the first assistant turn\. Since later dialogue turns are not preserved, these variants change conversation structure as well as response content\. For article\-style language\-modeling text, we use the raw WikiText\-2 training split of curated Wikipedia articles\([Merity et al\., 2017](https://arxiv.org/html/2609.30465#bib.bib29)\)\.

##### Perplexity across configurations\.

We evaluate 24 RCS\-LOO checkpoints, one per configuration, model and removal ratio\. Each is scored on 128 sequences of 2,048 tokens, yielding 262,016 next\-token targets\. PPL exponentiates the mean reference\-token NLL over all next\-token positions, not only assistant turns\. The unpruned references are5\.7815\.781on Qwen3\.6\-35B\-A3B and3\.5633\.563on GLM\-4\.7\-Flash\. Comparisons across configurations are subject to the input\-matching limits below\.

##### Which inputs are matched?

The Qwen3\.6\-35B\-A3B full\-RazorCaland self\-generated runs share the original model’s tokenization and evaluation rows\. Their earlier conversation pool differs from the eight\-axis fidelity pool\. It contains the first 128 conversations in source order that fill the token window, each truncated to its first 2,048 tokens\. On these matched inputs, self\-generated calibration yields lower PPL than fullRazorCalat both removal ratios \(6\.4686\.468versus7\.5447\.544at 25%, and6\.9916\.991versus9\.0289\.028at 50%\)\. This contrast still changes calibration token budgets and conversation structure as well as response content\.

Across the larger collection, self\-generated calibration yields lower PPL than tool\-filtered calibration on Qwen3\.6\-35B\-A3B but higher PPL on GLM\-4\.7\-Flash at both budgets\. WikiText calibration produces a 50%\-removal Qwen3\.6\-35B\-A3B checkpoint with PPL4\.8314\.831, below the unpruned reference\. The complementary half\-splits differ by0\.2950\.295on Qwen3\.6\-35B\-A3B at 25% removal and by0\.1700\.170on GLM\-4\.7\-Flash at 50%\. These orderings hold under limited input matching: matching reference and per\-sequence NLL values for shared checkpoints do not confirm that every configuration used identical input tokens, so each result reflects the combined effect of calibration source, token budget, and conversation structure rather than a controlled ranking of calibration sources or a robustness advantage over REAP\.

## Appendix EOutput Fidelity and Response Diversity

The main text reports a single aggregate fidelity gain per model and budget\. This appendix asks where that gain comes from and where it stops\. Reading it in order, the aggregate gain coexists with axis\-level regressions \(Appendix[E\.1](https://arxiv.org/html/2609.30465#A5.SS1)\), persists across domains but not on Qwen3\.6\-35B\-A3B’s held\-out group \(Appendix[E\.2](https://arxiv.org/html/2609.30465#A5.SS2)\), widens with routing concentration on some backbones only \(Appendix[E\.3](https://arxiv.org/html/2609.30465#A5.SS3)\), disagrees with reference\-token likelihood in many cells \(Appendix[E\.4](https://arxiv.org/html/2609.30465#A5.SS4)\), and does not carry over to response diversity \(Appendix[E\.5](https://arxiv.org/html/2609.30465#A5.SS5)\) or generation formatting \(Appendix[E\.6](https://arxiv.org/html/2609.30465#A5.SS6)\)\. The last two are the strongest constraint on how these gains should be read\.

Each analysis resamples a different unit, so their intervals are not interchangeable\. Per\-axis fidelity resamples paired evaluation shards, domain\-wise and routing\-stratified analyses resample paired batches, and response diversity resamples questions\. All reported 95% intervals are pointwise and unadjusted for multiple comparisons\.

### E\.1Per\-axis fidelity results

The main sweep pools four disjoint stratified shards, weighting tokens within axes and the eight axes equally\. Table[6](https://arxiv.org/html/2609.30465#A3.T6)reports its absolute and relative reverse KL at 25% and 50% removal\. The RCS\-LOO checkpoints have lower reverse KL than REAP on both models at both budgets, with larger relative reductions on GLM\-4\.7\-Flash\.

At 25% and 50% removal, the axis\-wise comparison separates four calibration\-covered axes \(*math, code, instruction following, tool use*\) from four held\-out axes \(*chat, creative, safety, SQL*\)\. RCS\-LOO has lower reverse\-KL point estimates in 11 of 16 axis–ratio cells on Qwen3\.6\-35B\-A3B and 13 of 16 on GLM\-4\.7\-Flash\. To distinguish favorable and adverse directions, we enumerate all44=2564^\{4\}=256ordered resamples of the four paired shards\. In each resample we recompute the token\-weighted KL of both methods within the axis and take100​\(Lmethod/LREAP−1\)100\(L\_\{\\mathrm\{method\}\}/L\_\{\\mathrm\{REAP\}\}\-1\)\. The 2\.5th and 97\.5th percentiles give pointwise intervals, of which 23 lie below zero, four lie above zero, and five include zero\. Per setting, the below/above/includes\-zero counts are 5/1/2 for Qwen3\.6\-35B\-A3B at both budgets, 6/1/1 for GLM\-4\.7\-Flash at 25%, and 7/1/0 for GLM\-4\.7\-Flash at 50%\. The four adverse cells are creative writing on Qwen3\.6\-35B\-A3B at both budgets, SQL on GLM\-4\.7\-Flash at 25%, and mathematics on GLM\-4\.7\-Flash at 50%\. Thus the aggregate improvement coexists with identifiable axis\-level regressions\. The same procedure applied toRazorgives lower point estimates in 11 of 16 cells on Qwen3\.6\-35B\-A3B and 14 of 16 on GLM\-4\.7\-Flash, with 24 intervals below zero, three above zero, and five including zero \(6/0/2 and 5/1/2 on Qwen3\.6\-35B\-A3B at 25%/50%, 6/1/1 and 7/1/0 on GLM\-4\.7\-Flash\)\. Refill removes the adverse Qwen3\.6\-35B\-A3B creative\-writing interval at 25%, but the other three adverse cells remain\. All comparisons share the same evaluation shards and ordered resamples within each axis\. Their intervals are paired contrasts, not estimates of variation across checkpoints or calibration draws\. This view therefore decomposes the same predictions rather than replicating the experiment\.

### E\.2Domain\-wise output fidelity

Figure[2](https://arxiv.org/html/2609.30465#S2.F2)resolves fidelity by domain for four backbones at both removal ratios, using predictions collected by applying criterion\-specific expert masks to each original model\. These mask\-based diagnostics are separate from the downstream benchmark runs\. The figure compares REAP, RCS, RCS\-LOO and RCS\-Refill using reverse KLDKL\(q∥p\)D\_\{\\mathrm\{KL\}\}\(q\\\|p\)in nats over four calibration\-covered and four held\-out domains\. Each domain estimate averages token means equally over six batches\. ID/OOD summaries then give equal weight to domain means, and intervals use 10,000 paired\-batch bootstrap draws, shared across criteria so every comparison stays paired on the same resampled batches\.

The residual criteria lead on most domains, with Qwen3\.6\-35B\-A3B’s held\-out group the consistent exception\. For RCS\-LOO versus REAP, lower\-KL domain counts are 6/4/7/6 at 25% removal and 7/5/7/7 at 50%, in GLM\-4\.7\-Flash, Qwen3\.6\-35B\-A3B, DeepSeek\-V4\-Flash\-0731, and Hy3 order, each out of eight domains\. Qwen3\.6\-35B\-A3B’s held\-out equal\-domain mean is higher than REAP’s by 7\.3% at 25% and 2\.3% at 50%, with relative\-change intervals crossing zero at both budgets, whereas the other three models have lower held\-out means at 50%\. Refill narrows this gap without closing it\. ForRazorthe corresponding counts are 5/6/6/6 at 25% and 7/5/8/7 at 50%, and its Qwen3\.6\-35B\-A3B held\-out mean changes by−3\.9%\-3\.9\\%at 25% and\+0\.2%\+0\.2\\%at 50% relative to REAP, again with intervals crossing zero and the other three models again lower at 50%\. Held\-out fidelity on that backbone is therefore where our advantage is weakest, and it is the same backbone whose excess perplexity moves against reverse KL in Appendix[C\.6](https://arxiv.org/html/2609.30465#A3.SS6)\.

Both removal ratios share reference predictions and token positions, with at most 1,024 valid positions per batch and no further subsampling\. For RCS\-LOO the 50% row uses the same predictions as Figure[8](https://arxiv.org/html/2609.30465#A5.F8), but aggregates them by domain rather than routing decile\. Neither view evaluates full generated trajectories\.

Each spoke uses an independent KL scale spanning all four point estimates, padded by the median bootstrap half\-width across criteria\. Curves are comparable along a spoke, but distances are not comparable across spokes, panels, or removal ratios\. Eight of 256 whiskers extend beyond the frame, with their full intervals recorded in the released plotted\-value table alongside relative changes against REAP\. Intermediate labels give mid\-axis KL values\.

### E\.3Routing\-stratified output fidelity

Consensus residuals measure how far an expert’s output sits from the mixture it participates in, so their advantage should plausibly be largest where few experts carry the token\. Routing concentration provides a direct test of that expectation, and the result is that concentration predicts the advantage on some backbones but not others\. Figure[8](https://arxiv.org/html/2609.30465#A5.F8)groups assistant tokens by routing concentration within each learned\-router MoE layer\. We rank tokens by negative normalized routing entropy−Hr\-H\_\{r\}and form ten quantile groups with shared ID/OOD boundaries\. A common finite\-value mask pairs methods and metrics, including the accompanying entropy and PPL measurements\. We exclude layer–batch–decile cells with fewer than 20 tokens, then average within\-cell token means equally over valid layers and batches\. This grouped estimator differs from the main sweep’s equal\-axis estimator\. The horizontal axis runs from diffuse to concentrated routing\.

Four calibration\-covered and four held\-out domains are evaluated separately\. Within each ID/OOD group, 2,000 paired\-batch bootstrap resamples preserve all ten deciles\. The 25% and 50% curves share original\-model routing profiles, token masks, quantile boundaries, and valid layer–batch–decile cells, but compute intervals separately\.

The collector does not apply learned\-router saliency scoring to DeepSeek\-V4\-Flash\-0731’s three fixed\-hash\-router layers, which are excluded from the curves\. Across criteria, these layers use identical frozen keep sets based on occurrence counts in the token\-to\-expert table, not calibration\-token frequencies\. Remapping preserveskkdistinct routes per token\.

At 50% removal, held\-out reverse KL is2\.12\.1–3\.0×3\.0\\timesits corresponding in\-distribution value\. From the least to the most concentrated routing decile, RCS\-LOO’s relative advantage widens by15\.6/17\.215\.6/17\.2percentage points on Qwen3\.6\-35B\-A3B ID/OOD and by4\.34\.3points on GLM\-4\.7\-Flash ID\. Nominal paired\-bootstrap intervals exclude zero for these three endpoint contrasts but cross zero for the other five, giving mixed evidence for a concentration\-dependent advantage\.

Aggregate point estimates favor RCS\-LOO by14\.4%14\.4\\%on GLM\-4\.7\-Flash OOD,15\.4/10\.3%15\.4/10\.3\\%on Hy3 ID/OOD, and7\.5/8\.9%7\.5/8\.9\\%on DeepSeek\-V4\-Flash\-0731 ID/OOD\. Qwen3\.6\-35B\-A3B’s OOD KL reduction is near zero \(−1\.5%\-1\.5\\%, CI\[−5\.3,\+2\.7\]\[\-5\.3,\+2\.7\]\)\. The same estimator applied toRazor, whose 50% predictions share the reference predictions and token mask, favors it by15\.6%15\.6\\%on GLM\-4\.7\-Flash OOD,17\.2/12\.9%17\.2/12\.9\\%on Hy3 ID/OOD, and10\.0/13\.1%10\.0/13\.1\\%on DeepSeek\-V4\-Flash\-0731 ID/OOD, with Qwen3\.6\-35B\-A3B OOD again nearly tied \(\+0\.5%\+0\.5\\%, CI\[−3\.7,\+4\.9\]\[\-3\.7,\+4\.9\]\)\. Relative to REAP, RCS\-LOO’s mean absolute entropy drift is lower by15%15\\%,22%22\\%,11%11\\%, and21%21\\%on GLM\-4\.7\-Flash, Qwen3\.6\-35B\-A3B, DeepSeek\-V4\-Flash\-0731, and Hy3, respectively\.

Adding the leave\-one\-out factor does not uniformly improve routing\-stratified fidelity under conditional RMS\. Averaged over deciles, RCS attains the largest KL reduction against REAP in three of the eight model–group panels\. These are both Qwen3\.6\-35B\-A3B panels \(\+15\.9%\+15\.9\\%ID and\+11\.9%\+11\.9\\%OOD, against\+9\.4%\+9\.4\\%and−0\.4%\-0\.4\\%forRazor\) and DeepSeek\-V4\-Flash\-0731 ID \(\+11\.4%\+11\.4\\%versus\+9\.7%\+9\.7\\%\)\.Razorleads the remaining five panels by1\.21\.2–4\.34\.3percentage points over the better of RCS and RCS\-LOO\. The largest RCS–RCS\-LOO gap occurs on Qwen3\.6\-35B\-A3B OOD, where RCS\-LOO andRazorboth fall below REAP on average, while RCS improves on it\. The gap narrows at the tenth decile, where the reductions are\+16\.5%\+16\.5\\%,\+11\.8%\+11\.8\\%, and\+12\.4%\+12\.4\\%for RCS, RCS\-LOO, andRazor, respectively\. These equal\-decile means at 50% removal do not establish a downstream ranking or replace the matched equal\-axis estimates in Table[6](https://arxiv.org/html/2609.30465#A3.T6)\. The RCS–RCS\-LOO contrast shows that the gate factor’s effect on fidelity depends on the backbone even with aggregation held fixed\.

Entropy drift complements reverse KL by measuring changes in predictive dispersion\. Predictive entropy itself is neither accuracy, calibration, nor response diversity, and neither higher nor lower entropy alone is preferable\.

Figure 8:Routing\-stratified fidelity at 50% removal\.All three scored criteria use conditional RMS, so the curves differ only in the token quantity\. The RCS arm was rescored for this figure on the same batches as the archive, with matching row order and domain labels and exactly matching stored original\-model NLL and predictive entropy\. All panels compare against REAP at the same decile, with a dashed zero line and higher values indicating better fidelity\. The top two rows show reverse\-KL reduction in percent\. The bottom two show REAP’s\|PPLq−PPLp\|\|\\mathrm\{PPL\}\_\{q\}\-\\mathrm\{PPL\}\_\{p\}\|minus the criterion’s, in PPL points \(Equation[24](https://arxiv.org/html/2609.30465#A5.E24)\), so positive values mean closer to the original\. Differences avoid unstable ratios where REAP’s deviation approaches zero \(0\.00160\.0016on Hy3 ID\)\. Bands are pointwise 95% paired batch\-bootstrap intervals, and panels use independent vertical scales\.Reverse KL and perplexity deviation give different rankings\.Razorlowers reverse KL relative to REAP in 73 of 80 deciles, with all seven exceptions in Qwen3\.6\-35B\-A3B’s held\-out group\. On perplexity deviation, however, it is farther from the original than REAP in7/107/10ID deciles on Qwen3\.6\-35B\-A3B and9/109/10on DeepSeek\-V4\-Flash\-0731, while remaining closer in all twenty GLM\-4\.7\-Flash ID and OOD deciles\. Across all groups it preserves PPL more closely in 47 of 80 deciles\. These are the same paired point comparisons summarized in Appendix[E\.4](https://arxiv.org/html/2609.30465#A5.SS4), not separate experiments\.

The incremental effect of refill depends on the budget and domain group\. At 50% removal,Razorexceeds RCS\-LOO by1\.01\.0–4\.34\.3percentage points in REAP\-relative KL reduction, averaging deciles equally within each of the eight model–domain groups\. Their marginal bands overlap throughout on Qwen3\.6\-35B\-A3B\. At 25% removal, the ordering varies by model\. In the calibration\-covered groups,Razorleads by8\.28\.2points on Qwen3\.6\-35B\-A3B but trails RCS\-LOO by3\.73\.7on GLM\-4\.7\-Flash and by1\.51\.5on DeepSeek\-V4\-Flash\-0731\. The refill margin also does not increase consistently with concentration, as its linear slope at 50% ranges from−0\.30\-0\.30to\+0\.06\+0\.06percentage points per decile across the eight groups\. Refill therefore shifts fidelity relative to RCS\-LOO in a direction that depends on the model and removal budget, not on routing concentration\.

### E\.4Reference\-token perplexity on the same routing groups

To distinguish distributional fidelity from reference\-token likelihood, we compute PPL at 50% removal for all four backbones on the same predictions, finite\-token mask, and layer–batch–decile cells as Figure[8](https://arxiv.org/html/2609.30465#A5.F8), whose lower two rows plot the result\. Using the reference\-token losses in Appendix[B\.4](https://arxiv.org/html/2609.30465#A2.SS4), letΔ​ℓ​\(t\)=ℓq​\(t\)−ℓp​\(t\)\\Delta\\ell\(t\)=\\ell\_\{q\}\(t\)\-\\ell\_\{p\}\(t\)\. Grouped PPL is

PPLp=exp⁡\(𝒜⁡\[ℓp\]\),PPLq=exp⁡\(𝒜⁡\[ℓp\+Δ​ℓ\]\),\\mathrm\{PPL\}\_\{p\}=\\exp\\\!\\big\(\\mathcal\{A\}\[\\ell\_\{p\}\]\\big\),\\qquad\\mathrm\{PPL\}\_\{q\}=\\exp\\\!\\big\(\\mathcal\{A\}\[\\ell\_\{p\}\+\\Delta\\ell\]\\big\),\(24\)where𝒜\\mathcal\{A\}is the routing\-stratified average over valid token cells, layers, and batches\. Exponentiation follows NLL aggregation rather than preceding it\. The result is neither mean per\-token perplexity nor exponentiated predictive entropy, and need not equal whole\-corpus token\-weighted PPL\. We measure preservation by the absolute deviation\|PPLq−PPLp\|\|\\mathrm\{PPL\}\_\{q\}\-\\mathrm\{PPL\}\_\{p\}\|, so “closer” means a smaller deviation than REAP on the same group, not simply a lower PPL\.

The bootstrap pairs the original, REAP, and RCS\-LOO predictions on the same retained positions\. Within each ID/OOD group, shared whole\-batch indices preserve all ten deciles\. Recomputing and exponentiating each resampled NLL average gives marginal intervals and paired PPL differences\.

The two fidelity views disagree widely\. Across 80 point estimates from four models, two ID/OOD groups, and ten deciles, RCS\-LOO has lower reverse KL than REAP in 72 but PPL closer to the original in only 47, andRazorgives 73 and 47 on the same cells\. The disagreement is concentrated rather than diffuse\. Both measures favor RCS\-LOO throughout GLM\-4\.7\-Flash’s groups, whereas on Qwen3\.6\-35B\-A3B PPL is closer in only two ID and one OOD decile despite lower reverse KL in every ID decile\. DeepSeek\-V4\-Flash\-0731 illustrates why the measures can part ways, since its ID NLL changes are negative, so a lower PPL there means a larger deviation from the original rather than a smaller one\. Improving the full predictive distribution and improving the likelihood of the particular reference continuation are therefore distinct objectives, and our scores optimize neither directly\.

### E\.5Response diversity

Every analysis so far scores predictions at reference tokens, which leaves sampled generation unexamined\. Pruning can preserve next\-token distributions on reference prefixes and still change what the model produces when it writes freely, and for a reasoning model that behavior is what a user sees\. This subsection measures it through lexical and embedding diversity across repeated generations, and finds that the diversity ordering does not track the fidelity ordering\. Two collections are involved and are not interchangeable\. A three\-benchmark comparison covers all four criteria on Qwen3\.6\-35B\-A3B, and a larger collection covers fewer criteria on four tasks and two backbones\.

##### Criterion\-family comparison\.

The left panel of Figure[3](https://arxiv.org/html/2609.30465#S3.F3)compares REAP, RCS, RCS\-LOO, andRazoron Qwen3\.6\-35B\-A3B using Distinct\-nnforn=2,3,4n=2,3,4\. For each benchmark, all four criteria at both budgets and the unpruned reference use 16 fixed response positions per question, drawn from the first 256 positions for AIME and 32 for each coding task, the largest pools common to every criterion\. This comparison draws from different sampling pools than the four\-task analysis below, so their percentages are not directly comparable\.

Within each benchmark, values are100​\(M¯/M¯Original−1\)100\(\\bar\{M\}/\\bar\{M\}\_\{\\mathrm\{Original\}\}\-1\)for question\-averaged metrics\. AIME 2026, HumanEval\+, and LiveCodeBench receive equal weight, so AIME contributes one third of the weight on 30 of 369 questions\. For pointwise 95% paired question\-bootstrap intervals, we average the resampled relative changes across benchmarks before taking percentiles\. These intervals condition on the 16 fixed responses rather than measure variation over new generations\. Of twelve intervals per budget \(four criteria and three orders\), all twelve exclude zero at 50% and nine do so at 25%\. They describe changes from Original, not the RCS\-LOO–REAP difference\. In per\-benchmark point estimates, RCS\-LOO exceeds REAP in all nine benchmark–order cells at 50% and six at 25%\. Across the six benchmark\-averaged budget–order cells, its median advantage over RCS is 4\.02 percentage points in relative change from Original\. Refill lowers Distinct\-nnrelative to RCS\-LOO in five of these six cells\.

##### Four\-task collection and sampling\.

The repeated\-generation collection is separate from the nine\-task benchmark\. It contains 30 AIME 2026\([Mathematical Association of America, n\.d\.](https://arxiv.org/html/2609.30465#bib.bib28)\), 219 Minerva\([Lewkowycz et al\., 2022](https://arxiv.org/html/2609.30465#bib.bib18)\), 164 HumanEval\+\([Liu et al\., 2023](https://arxiv.org/html/2609.30465#bib.bib23)\), and 175 LiveCodeBench\([Jain et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib11)\)questions\. Minerva uses 219 of OCWCourses’ 272 undergraduate STEM problems from MIT OpenCourseWare, with the source reports recording the 53 omitted problems as ungradable\. Its 175 LiveCodeBench task IDs are fixed across all variants compared here, which is what the paired comparisons require, and are drawn from a different release than the nine\-task benchmark\.

It covers Original plus EAN, Frequency, REAP, and RCS\-LOO at 25% and 50% removal on Qwen3\.6\-35B\-A3B, and Original, REAP, and RCS\-LOO at both budgets on GLM\-4\.7\-Flash\. Unlike the three\-benchmark comparison in Figure[3](https://arxiv.org/html/2609.30465#S3.F3), this analysis includes Minerva and GLM\-4\.7\-Flash\.

Mathematics has 2,048 generations per question, HumanEval\+ has 512, and LiveCodeBench has 256\. Generation uses temperature0\.70\.7, top\-p=0\.95p=0\.95, top\-k=20k=20, repetition penalty1\.01\.0, an 8,192\-token output cap, and thinking disabled, which differs from benchmark decoding in Appendix[B\.3](https://arxiv.org/html/2609.30465#A2.SS3)\. We sample 16 response positions per question without replacement, reusing the same positions across methods so that all comparisons are paired\. This yields 131,712 responses, of which 53,136 cover the three tasks shared with Figure[3](https://arxiv.org/html/2609.30465#S3.F3)\.

##### Metric definitions\.

Metrics are computed per question and then averaged over questions\. We adapt the distinct\-nnapproach of[Li et al\. \(2016\)](https://arxiv.org/html/2609.30465#bib.bib19)to model tokens, usingn∈\{2,4\}n\\in\\\{2,4\\\}in the four\-task analysis andn∈\{2,3,4\}n\\in\\\{2,3,4\\\}in Figure[3](https://arxiv.org/html/2609.30465#S3.F3)\. Distinct\-nndivides the number of unique tokennn\-grams pooled across 16 responses by their total occurrences, without crossing response boundaries\. Jaccard distance averages1−\|Gi∩Gj\|/\|Gi∪Gj\|1\-\|G\_\{i\}\\cap G\_\{j\}\|/\|G\_\{i\}\\cup G\_\{j\}\|over all 120 response pairs, whereGiG\_\{i\}andGjG\_\{j\}are response\-level token 4\-gram sets\. Both lexical measures use each model family’s own tokenizer\.

Embedding distance averages pairwise cosine distances using the fixed all\-MiniLM\-L6\-v2 encoder\.111[https://huggingface\.co/sentence\-transformers/all\-MiniLM\-L6\-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2)Our full\-response extension splits text into non\-overlapping chunks of at most 254 WordPieces, mean\-pools token representations, and normalizes each chunk vector\. We then normalize their length\-weighted mean\. This retains content beyond the first chunk but not chunk order\. The resulting distances do not measure execution behavior or mathematical\-strategy equivalence\.

##### Normalization and uncertainty\.

For each method, family, and task, this comparison reports100​\(M¯/M¯Original−1\)100\(\\bar\{M\}/\\bar\{M\}\_\{\\mathrm\{Original\}\}\-1\)using question meansM¯\\bar\{M\}, where Original denotes the unpruned MoE\. Higher values indicate more diversity\. We recompute this ratio over 2,000 paired question\-bootstrap draws and report pointwise percentile 95% intervals, conditional on the 16 fixed responses and evaluated checkpoints\. They exclude new\-response and checkpoint\-level variation\.

##### Length control and interpretation\.

The four diversity metrics are correlated, so they are not independent evidence\. We assess length sensitivity by repeating the analysis on exactly the first 128 tokens of each response\. Within a family, we retain only questions eligible for every available method at both pruning ratios\. In AIME 2026, Minerva, HumanEval\+, and LiveCodeBench order, the resulting question counts are 30/219/134/143 for Qwen3\.6\-35B\-A3B and 29/176/82/142 for GLM\-4\.7\-Flash\.

At 50% removal, the four\-task mean Distinct\-4 gap between RCS\-LOO and REAP changes from\+4\.30%\+4\.30\\%on full responses to−0\.37%\-0\.37\\%under prefix control on Qwen3\.6\-35B\-A3B\. The corresponding gaps on GLM\-4\.7\-Flash are\+10\.35%\+10\.35\\%and\+16\.50%\+16\.50\\%\. These means weight task\-relative changes against REAP equally, unlike the Original\-normalized diversity panel of Figure[3](https://arxiv.org/html/2609.30465#S3.F3)\. The control changes both length and question eligibility, so Qwen3\.6\-35B\-A3B’s reversal under prefix control reflects both factors and is not interpretable as a pure length effect\. Lexical and embedding diversity measure surface variation in token sequences and do not measure correctness, token efficiency, or the number of distinct correct solutions\.

### E\.6Response delimiters, formatting, and termination

Diversity shifts leave open whether pruned models still produce well\-formed responses at all\. This subsection measures three failure modes that are visible without judging content, namely emitting a stray closing thinking delimiter, omitting the code fence that a coding response is expected to carry, and reaching the output cap instead of terminating\. All three rise with removal budget on both backbones in rank correlation, though not stepwise and not significantly in every case, and at 75% removal they reach rates that no benchmark score in this paper would reveal\. This is the clearest evidence here that retained accuracy alone does not certify a pruned reasoning model\.

##### Protocol and operational definitions\.

This analysis follows the repeated\-generation protocol in Appendix[E\.5](https://arxiv.org/html/2609.30465#A5.SS5), using temperature0\.70\.7, top\-p=0\.95p=0\.95, top\-k=20k=20, an 8,192\-token cap, and thinking disabled\. The tables and the three response\-form panels of Figure[3](https://arxiv.org/html/2609.30465#S3.F3)use 256 shared sample positions per question, giving 44,800 completions per variant on LiveCodeBench\. These measurements characterize response form, not the benchmark accuracy of Table[1](https://arxiv.org/html/2609.30465#S3.T1)\.

A*stray*</think\>is a literal occurrence of that string in a completion\. In this tokenizer it is a single vocabulary entry that is not flagged as a special token, so ordinary detokenization leaves it in the text\. The chat template already closes the thinking block inside the prompt, so any further occurrence is emitted by the model and is not evidence about latent reasoning\. We avoid the word*leakage*because it denotes benchmark contamination in evaluation, which these counts do not measure\.*No fence*denotes the absence of a line\-delimited code fence, while*prefix*denotes a nonempty prefix before such a fence\. Both are shares of all completions, and neither is a semantic code or narration classifier\. Unfenced responses can contain code, which the evaluator may extract through its bare\-code fallback\.*Length\-limited*uses the decoder’s finish reason\.

Response lengths in this table count Unicode characters, whereas Figure[3](https://arxiv.org/html/2609.30465#S3.F3)counts tokens under the backbone’s own tokenizer, matching the units of its Distinct\-nnpanel\. Characters per token fall from3\.443\.44on the unpruned model to3\.043\.04for REAP at 50% removal, so the two units are not interchangeable\. The character counts here understate that checkpoint’s growth in generated tokens\. Intervals resample questions to retain dependence among completions sharing a prompt\.

Table 10:Response form on LiveCodeBench by criterion and budget,Qwen3\.6\-35B\-A3B\.Rates are percentages except stray</think\>per 10,000\. Lengths are in kchar\. Brackets give pointwise 95% question\-bootstrap intervals\. GLM\-4\.7\-Flash is measured but not printed here because that collection covers only REAP and RCS\-LOO, with noRazorrow for comparison\. Its RCS\-LOO ladder appears in Table[11](https://arxiv.org/html/2609.30465#A5.T11)\.Table 11:The RCS\-LOO budget ladder on LiveCodeBench\.Spearman correlations use all 10 checkpoints, not only the four shown\. Units follow Table[10](https://arxiv.org/html/2609.30465#A5.T10)\.
##### Sample\-pool sensitivity\.

Rare\-event estimates depend on the number of sampled completions\. REAP’s stray\-delimiter rate is8\.938\.93per10,00010\{,\}000with 32 completions per question versus13\.3913\.39with 256, although the criteria retain the same point\-estimate ordering at 50% removal under either pool\. Both criteria in the main\-text panel use the full 44,800\-completion pool, whose rates and intervals are reported below\.

##### Budget\-dependent changes\.

We measure ten RCS\-LOO checkpoints from 0% to 75% removal on each backbone\. Table[11](https://arxiv.org/html/2609.30465#A5.T11)shows the quarter budgets and reports each rank correlation over all ten checkpoints\. Only RCS\-LOO covers every measured budget in this collection, while Table[10](https://arxiv.org/html/2609.30465#A5.T10)compares other criteria at shared budgets\. On Qwen3\.6\-35B\-A3B, the unfenced and length\-limited shares each have Spearmanρ=\+0\.82\\rho=\+0\.82with removal budget \(p=0\.007p=0\.007\), and the rate of stray</think\>hasρ=\+0\.75\\rho=\+0\.75\(p=0\.020p=0\.020\)\. GLM\-4\.7\-Flash givesρ=\+0\.93\\rho=\+0\.93\(p<0\.001p<0\.001\),\+0\.58\+0\.58\(p=0\.099p=0\.099\), and\+0\.89\+0\.89\(p=0\.001p=0\.001\), respectively\. These descriptive associations are not stepwise monotonic\. Among the undisplayed budgets, Qwen3\.6\-35B\-A3B’s length\-limited share falls at 40% and GLM\-4\.7\-Flash’s stray\-delimiter rate peaks at 60%\. The displayed values also show Qwen3\.6\-35B\-A3B’s stray rate higher at 50% than at 75%\.

The remaining two measures show no such trend\. Median response length hasρ=\+0\.43\\rho=\+0\.43\(p=0\.24p=0\.24\) on Qwen3\.6\-35B\-A3B and\+0\.17\+0\.17\(p=0\.67p=0\.67\) on GLM\-4\.7\-Flash, and the prefix share hasρ=−0\.28\\rho=\-0\.28\(p=0\.46p=0\.46\) and−0\.88\-0\.88\(p=0\.002p=0\.002\), respectively\. Nonsignificance here is weak evidence either way, since all tests are exploratory, use ten checkpoints on shared questions, report two\-sidedpp\-values uncorrected for multiple comparisons, and do not estimate independent\-run variability\.

The end of the ladder makes the practical stake concrete\. At 75% removal, Qwen3\.6\-35B\-A3B produces25\.8%25\.8\\%unfenced and42\.0%42\.0\\%length\-limited completions, while GLM\-4\.7\-Flash produces78\.7%78\.7\\%length\-limited completions at a median length of 30,444 characters\. A checkpoint that fails to terminate on most coding prompts is unusable regardless of how its retained experts score, which is why we report these rates alongside the benchmark averages rather than after them\.

##### Refill and delimiter emissions on Qwen3\.6\-35B\-A3B\.

At 25% removal, every criterion’s tag\-rate point estimate is at or below the unpruned Qwen3\.6\-35B\-A3B rate of3\.63\.6per 10,000, with RCS\-LOO lowest at0\.20\.2\[0\.0,0\.7\]\[0\.0,0\.7\]\(Table[10](https://arxiv.org/html/2609.30465#A5.T10)\)\. At 50%, RCS\-LOO reaches68\.568\.5\[22\.8,133\.7\]\[22\.8,133\.7\], versus REAP’s13\.413\.4\[7\.6,20\.1\]\[7\.6,20\.1\]\. The refill checkpoint reduces the observed rate to8\.58\.5\[4\.2,13\.6\]\[4\.2,13\.6\], below REAP’s point estimate but with overlapping marginal intervals\.

The RCS\-LOO–REAP gap persists within baseline\-length strata\. In the fourth quintile of the unpruned length distribution, rates are100\.7100\.7and14\.214\.2per 10,000, withRazorat14\.414\.4\. Matching question and sample position against Original gives 307 newly occurring tag events for RCS\-LOO, 60 for REAP, and 38 forRazor\. These checks preserve the comparison within shared prompts and baseline\-length groups without ruling out effects of changed generated lengths or termination\. On the same collection, GLM\-4\.7\-Flash’s corresponding RCS\-LOO rate is2\.92\.9\[1\.1,4\.9\]\[1\.1,4\.9\], as shown at 50% removal in Table[11](https://arxiv.org/html/2609.30465#A5.T11)\. REAP, which is not included in that table, has a rate of1\.61\.6\[0\.4,2\.7\]\[0\.4,2\.7\]\. The magnitude of the Qwen3\.6\-35B\-A3B gap therefore does not carry over\. Across variants,9898–100%100\\%of tag\-containing completions also contain code output, but this does not establish answer correctness or preserved utility\.

##### Length shifts are small at 25% and mixed at 50%\.

At every 25% checkpoint, the median paired length change relative to Original, matched by question and sample position, is at most 17 characters in magnitude\. At 50%, REAP lengthens responses by a median 542 characters \(67\.3%67\.3\\%of completions longer\), RCS\-LOO by 126 andRazorby 114, while Frequency shortens them by 129\. The share of completions with a prefix before the code fence falls to16\.6%16\.6\\%for RCS\-LOO and19\.6%19\.6\\%forRazorbut rises to66\.3%66\.3\\%for REAP, against61\.4%61\.4\\%unpruned\.

##### Task and backbone dependence\.

HumanEval\+, the shortest task on both backbones, is the only task nearly free of visible tags throughout, with one tag\-containing completion out of 671,744\. On the other tasks, the two backbones differ\. Qwen3\.6\-35B\-A3B emits tags on LiveCodeBench but almost nowhere else, with one AIME 2026 completion out of 84,480 and none of 616,704 on Minerva\. GLM\-4\.7\-Flash instead shows this behavior on mathematics\. At 50% removal,2\.90%2\.90\\%\[1\.25,5\.18\]\[1\.25,5\.18\]of its AIME 2026 completions carry a tag under REAP and7\.25%7\.25\\%\[4\.11,11\.22\]\[4\.11,11\.22\]under RCS\-LOO\. The corresponding rates are0\.67%0\.67\\%and0\.54%0\.54\\%on Minerva, against at most0\.03%0\.03\\%on LiveCodeBench\.

Task\-level median length alone does not order these rates\. Qwen3\.6\-35B\-A3B’s AIME 2026 responses are longer than its LiveCodeBench responses \(10,862 versus 8,332 characters\) yet contain almost no tags\. This comparison does not isolate length from other task differences\.

##### Matched examples\.

For one LiveCodeBench question at a fixed sample position, the 50% checkpoints produce 1,894 and 1,897 characters under Frequency andRazor\(both opening with code\), 4,228 under REAP, 4,875 under EAN, and 6,923 under RCS\-LOO, versus 3,423 unpruned\. Only RCS\-LOO emits a closing thinking tag, preceded by 6,233 characters ending with “Let's code accordingly\.” and followed by program text\. This describes the visible response sequence, not a latent reasoning process\. The median pre\-tag length among this checkpoint’s tag\-containing completions is 12,507 characters\.

A separate RCS\-LOO 70% completion reaches the cap after 21,966 characters without a code fence, contributing to both the unfenced and length\-limited counts\. The two examples were selected by median preamble length within their respective categories\. The tag\-containing example terminates normally\. They illustrate the measured categories rather than establish comparative answer quality\.

## Appendix FExtended Related Work

MoE compression methods differ along three axes, namely the removal unit, the importance signal, and the selection or adaptation procedure applied afterwards\. Sorting the literature this way clarifies what our comparison does and does not cover, becauseRazorvaries only the importance signal\. It removes whole experts, ranks them independently under a fixed layerwise budget, and leaves retained weights and routers untouched\. The three subsections below group prior work by how far it departs from that setting, moving from alternative whole\-expert saliency signals, through methods that model expert relations or search over candidate sets, to methods that change the surviving functions or the router itself\.

### F\.1Whole\-expert saliency and conditional aggregation

Routing, activation, and parameter statistics offer inexpensive rankings\. SEER\-MoE uses activation counts or routing\-probability soft counts with layerwise or global removal\([Muzio et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib30)\)\. AIMER instead ranks experts without calibration data using a normalizedℓ1/ℓ2\\ell\_\{1\}/\\ell\_\{2\}concentration statistic over their weights\([Liu et al\., 2026a](https://arxiv.org/html/2609.30465#bib.bib24)\)\.[Jaiswal et al\. \(2025\)](https://arxiv.org/html/2609.30465#bib.bib12)compare several expert\-dropping criteria, including activation norms\. These independent scores do not reconstruct the routed mixture after deletion\.

REAP, our closest magnitude\-based comparator, averageswi​‖fi‖2w\_\{i\}\\\|f\_\{i\}\\\|\_\{2\}over tokens selecting expertii\([Lasby et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib15)\)\. Its analysis includes promoted substitution and survivor renormalization, but its saliency uses only the removed expert’s weighted output norm\. A unified formulation similarly derives an exact rerouting residual before omitting it from practical one\-pass scores\([Liu et al\., 2026b](https://arxiv.org/html/2609.30465#bib.bib25)\)\. Consensus residuals instead retain the removed expert’s direction relative to the mixture, while RCS\-Refill also includes the promoted expert’s weighted residual\.

Token\-level scoring and cross\-token aggregation remain separate choices\.Razoruses conditional RMS rather than REAP’s conditional mean, while the component comparisons hold the token quantity fixed when varying aggregation \(Appendix[C](https://arxiv.org/html/2609.30465#A3)\)\. Neither conditioning nor RMS follows from the local deletion identity, and no aggregation rule is uniformly preferred by the available results\.

### F\.2Expert relations and candidate\-set selection

Relational information is not unique to consensus\-residual scoring\. STUN clusters router\-derived behavior before representative selection or selective reconstruction and unstructured pruning\([Lee et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib16)\)\. HC\-SMoE instead clusters average expert outputs and merges each group\([Chen et al\., 2025](https://arxiv.org/html/2609.30465#bib.bib2)\)\. SHAPE uses routing co\-occurrence in a coalition\-utility proxy\([Zhang, 2026](https://arxiv.org/html/2609.30465#bib.bib42)\)\. ConMoE combines routing\-conditioned contribution with nearest\-expert parameter distance before prototype reassignment\([Yao et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib41)\)\. MAESTRO derives importance from cross\-layer routing transitions before layerwise selection\([Goel et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib8)\), with its main results using attention\-only LoRA recovery while experts and routers remain frozen\.

Our scores instead evaluate output change under specified deletion counterfactuals at a fixed layer input\. The reference is the routed mixture, not a cluster representative or a routing graph\. This distinction concerns the local score, not a claim that independent expert ranking captures all interactions\. Residuals can reinforce or cancel under joint removal \(Appendix[A\.3](https://arxiv.org/html/2609.30465#A1.SS3)\), so exact single\-deletion quantities do not yield an exact set\-selection algorithm\.

Other approaches evaluate candidate sets directly\.[Lu et al\. \(2024\)](https://arxiv.org/html/2609.30465#bib.bib26)select retained combinations by layer\-output reconstruction, EEP searches pruning and merging configurations\([Liu et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib22)\), and MoE\-I2evaluates layerwise deletion sets before jointly choosing candidates across short layer blocks\([Yang et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib40)\)\. We instead retain the highest\-scoring experts under a fixed layerwise budget\. Our local single\-deletion identities neither solve combinatorial selection nor establish superiority to subset search\.

### F\.3Merging, router adaptation, and budget allocation

Expert merging changes the surviving functions, rather than only selecting which original experts remain\. MC\-SMoE groups experts around high\-usage representatives using router\-logit similarity, aligns neurons, and averages expert parameters before low\-rank and structured compression\([Li et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib21)\)\. REAM protects high\-saliency experts as centroids and merges similar non\-centroid experts into them\([Jha et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib13)\)\. Neither operation is the router refill modeled by RCS\-Refill, which promotes an existing unselected expert without averaging its parameters with the removed expert\.

Router KD adapts only router parameters after compression, distilling the uncompressed model’s next\-token distribution while freezing the experts and other backbone parameters\([Hyeon & Do, 2026](https://arxiv.org/html/2609.30465#bib.bib10)\)\. It is a recovery procedure, not an expert\-ranking criterion\. Budget allocation is another distinct choice\. GRAPE compares residual redundancy across layers to allocate non\-uniform budgets, merging similar expert pairs within the selected layer\([Zhang et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib43)\)\. MC\-SMoE also permits adaptive layerwise merging ratios\([Li et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib21)\)\. These examples show that merging, adaptation, and budget allocation are overlapping design choices rather than mutually exclusive method families\. Our experiments use uniform layerwise removal ratios, leave retained expert weights unchanged, and perform no recovery training\. Combining residual\-based scores with merging, router adaptation, or non\-uniform allocation is outside the present comparison, so our results do not establish their compatibility or joint benefit\.

### F\.4Finer\-grained compression and comparison scope

MoE\-Pruner sparsifies weights within experts using weight magnitudes and router\-weighted input activation statistics\([Xie et al\., 2024](https://arxiv.org/html/2609.30465#bib.bib39)\)\. HEAPr treats feed\-forward intermediate neurons as atomic expert units and uses output\-space second\-order information to estimate their importance\([Li et al\., 2026](https://arxiv.org/html/2609.30465#bib.bib20)\)\. Both operate at a finer granularity than removing entire routed experts\. An atomic expert in this terminology is not one of the router’s original selectable experts\.

These approaches share an interest in preserving model behavior, but their removal units and compression procedures differ from ours, so their reported sparsity levels are not matched whole\-expert budgets\.

The scope of our comparison follows from this survey\. We hold the removal unit, the layerwise budget, the retained weights, and the router fixed, and vary only the importance signal among RCS, RCS\-LOO, and RCS\-Refill against the magnitude\-based and frequency\-based baselines, withRazordenoting RCS\-Refill under conditional RMS\. Our evidence therefore concerns which local signal best identifies replaceable experts, and it leaves open how such a signal would combine with merging, router adaptation, non\-uniform budget allocation, or sub\-expert granularity\. Because those families act on the surviving computation rather than on the ranking, a better ranking is plausibly complementary to them, but we do not test that\.

相似文章

混合专家语言模型中机器遗忘的路由感知专家校准

arXiv cs.CL

论文提出TRACE,一种用于混合专家语言模型中机器遗忘的方法,通过重新加权词元级保留损失来校准保留正则化,以解决遗忘-保留路由不匹配问题。实验表明,在多个MoE大语言模型上改善了遗忘-效用权衡。