From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models
Summary
This survey paper defines model fusion and organizes research into three levels—parameter, representation, and behavior—for integrating capabilities in large language models.
View Cached Full Text
Cached at: 09/18/26, 08:56 AM
# A Survey of Model Fusion for Large Language Models
Source: [https://arxiv.org/html/2609.19553](https://arxiv.org/html/2609.19553)
## From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models
Yanggan GuZihao WangYuanyi WangYibo YanAffiliation:The Hong Kong University of Science and Technology \(Guangzhou\)Wenjun WangAffiliation:The Hong Kong Polytechnic UniversityYuhang LiuAffiliation:InfiX\.aiGuanghao ZhuAffiliation:The Hong Kong Polytechnic UniversitySirui HuangAffiliation:The Hong Kong Polytechnic UniversityMing LiAffiliation:The Hong Kong Polytechnic UniversityHongxia YangAffiliation:The Hong Kong Polytechnic UniversityAffiliation:PolyU\-Daya Bay Technology and Innovation Research InstituteAffiliation:The Chinese University of Hong Kong
###### Abstract
Model fusion integrates the capabilities from source models into a single target model\. As of June 2026, Hugging Face hosts more than 2M models\. This growing pool provides a rich base for model reuse and capability integration\. Yet existing surveys often cover only separate parts of this space, and they do not provide a unified definition or a systematic taxonomy\. This survey defines model fusion and organizes prior work into three levels: parameter\-level, representation\-level, and behavior\-level fusion\. We also review related metrics, benchmarks, and applications, summarize current challenges, and identify future directions\. Our goal is to provide a clear map of this area and support future work on model fusion\. A comprehensive list of papers about model fusion is available at[https://github\.com/Baicaihaochi/Awesome\-Model\-Fusion\-Survey](https://github.com/Baicaihaochi/Awesome-Model-Fusion-Survey)\.
††footnotetext:\*Equal contribution\.†Corresponding authors\.## 1Introduction
As large language models and the open\-source ecosystem continue to grow, the number and variety of available models have increased quickly\. As of June 2026, Hugging Face hosts more than 2M models111[https://huggingface\.co/blog/huggingface/state\-of\-os\-hf\-spring\-2026](https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026), providing a rich and diverse base for model reuse and capability integration\. Therefore, reusing and integrating existing model capabilities within a single model is becoming an important direction\([Li et al\., 2026c](https://arxiv.org/html/2609.19553#bib.bib11);[Zheng et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib47)\)\.
Given multiple source models with diverse capabilities, model fusion constructs a target model by combining their parameters, aligning their representations, or distilling their output behaviors, as shown in Figure[1](https://arxiv.org/html/2609.19553#S1.F1)\. After fusion, the target model operates without relying on the complete source models at inference time\. Under this operational definition, traditional model merging and knowledge distillation can be viewed as parameter\-level and behavior\-level model fusion, respectively\([Yang et al\., 2026a](https://arxiv.org/html/2609.19553#bib.bib65);[Song and Zheng, 2026b](https://arxiv.org/html/2609.19553#bib.bib66);[Yadav et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib93);[Li et al\., 2026c](https://arxiv.org/html/2609.19553#bib.bib11)\)\.
Figure 1:Model fusion overview\. Multiple source models contribute parameters, representations, or behaviors to construct one target model\. After fusion, the target operates without the complete source models at inference time\.Figure 2:Progress timeline of model fusion\. Selected methods are color\-coded by their primary fusion level\. Annual counts indicate the broader growth trend; the 2026 count is a partial\-year snapshot\.Model fusion offers two practical advantages\. First, it enables efficient reuse of existing models and integrates their capabilities into one target model\. This can be done by fusing parameters, using representations to diagnose and repair drift, or distilling output behaviors\([Yadav et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib98);[Yang et al\., 2024a](https://arxiv.org/html/2609.19553#bib.bib84);[Wan et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib45);[Agarwal et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib76)\)\. As shown in Figure[2](https://arxiv.org/html/2609.19553#S1.F2), model fusion has attracted steadily increasing researcher attention since 2023\. This trend is also reflected in industrial practice, where DeepSeek\-V4\([DeepSeek\-AI et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib13)\), NVIDIA’s Nemotron\-Cascade 2\([Yang et al\., 2026f](https://arxiv.org/html/2609.19553#bib.bib73)\), and GLM\-5\([GLM\-5\-Team et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib27)\)adopt on\-policy distillation to integrate or recover model capabilities\. Second, model fusion supports continual learning by absorbing new task signals while preserving earlier capabilities\. AIMMerging\([Feng et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib6)\), RECALL\([Wang et al\., 2025a](https://arxiv.org/html/2609.19553#bib.bib81)\), and NUFILT\([Qiu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib75)\)reduce forgetting\.
Despite these advantages, recent studies also show that model fusion remains far from settled\. Weight averaging and alignment can improve accuracy and robustness, but some task\-level combinations may collapse, and current theory cannot yet predict when fusion will succeed\([Wortsman et al\., 2022](https://arxiv.org/html/2609.19553#bib.bib70);[Ainsworth et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib28);[Cao et al\., 2026b](https://arxiv.org/html/2609.19553#bib.bib21)\)\. Moreover, fusion becomes harder when source models differ in architecture, tokenizer, or modality, because parameter and representation alignment can be unstable\([Sung et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib22);[Cui et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib103);[Du et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib4)\)\. In evaluation, recent benchmarks improve standardization, but average scores can still hide local degradation and cross\-capability interference\([Tang et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib26);[He et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib57);[Tam et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib80);[Cao et al\., 2026b](https://arxiv.org/html/2609.19553#bib.bib21)\)\.
SurveyVenue & YearParam\. levelRepre\. levelBehav\. level[Gou et al\. \(2021\)](https://arxiv.org/html/2609.19553#bib.bib44)IJCV’21✓\\checkmark✓\\checkmark[Xu et al\. \(2024\)](https://arxiv.org/html/2609.19553#bib.bib92)arXiv’24✓\\checkmark✓\\checkmark[Yadav et al\. \(2025\)](https://arxiv.org/html/2609.19553#bib.bib93)TMLR’25✓\\checkmark[Yang et al\. \(2025a\)](https://arxiv.org/html/2609.19553#bib.bib39)TIST’25✓\\checkmark✓\\checkmark[Qin et al\. \(2025\)](https://arxiv.org/html/2609.19553#bib.bib43)IJIS’25✓\\checkmark✓\\checkmark[Yang et al\. \(2026a\)](https://arxiv.org/html/2609.19553#bib.bib65)CSUR’26✓\\checkmark✓\\checkmark[Song and Zheng \(2026b\)](https://arxiv.org/html/2609.19553#bib.bib66)arXiv’26✓\\checkmark✓\\checkmark[Li et al\. \(2026c\)](https://arxiv.org/html/2609.19553#bib.bib11)TNNLS’26✓\\checkmark✓\\checkmark[Song and Zheng \(2026a\)](https://arxiv.org/html/2609.19553#bib.bib94)arXiv’26✓\\checkmark[Fang et al\. \(2026\)](https://arxiv.org/html/2609.19553#bib.bib40)AIR’26✓\\checkmark✓\\checkmarkOurs✓\\checkmark✓\\checkmark✓\\checkmarkTable 1:Coverage of related surveys across the three fusion levels\. A checkmark indicates substantive treatment of methods at that level; only our survey covers parameter\-, representation\-, and behavior\-level fusion together\.As shown in Table[1](https://arxiv.org/html/2609.19553#S1.T1), existing surveys treat model merging and knowledge transfer separately\([Xu et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib92);[Song and Zheng, 2026a](https://arxiv.org/html/2609.19553#bib.bib94);[Gou et al\., 2021](https://arxiv.org/html/2609.19553#bib.bib44);[Yang et al\., 2025a](https://arxiv.org/html/2609.19553#bib.bib39);[Qin et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib43);[Fang et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib40)\), without a unified view of the scope, boundary, and evaluation of model fusion\. We address this gap with a formal definition and a three\-level taxonomy, followed by evaluation, applications, and open research problems\.
We survey more than 150 papers\. Section 2 defines model fusion, and Section 3 presents the three\-level taxonomy and evaluation settings\. Sections 4–6 present practical takeaways and challenges\.
## 2Definition and Formulation
#### Definition\.
We define model fusion as follows:
> Model fusion constructs a target model by combining parameters, aligning representations, or distilling output behaviors from multiple source models\. Its goal is to integrate the knowledge and capabilities carried by the source models while retaining efficient inference\.
Accordingly, we include a method as model fusion only if its construction of the target model uses at least one of these three operations and the resulting target model can operate at inference time without relying on the complete source models\. Whether an operation occurs during training or after training is orthogonal to this definition: model merging is often training\-free, whereas representation alignment and behavior distillation usually optimize the target model during fusion\. Task Arithmetic combines parameter updates, Representation Surgery aligns hidden representations, and FuseLLM distills output distributions; each produces a target that is independent of the complete source models at inference time\([Ilharco et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib20);[Yang et al\., 2024a](https://arxiv.org/html/2609.19553#bib.bib84);[Wan et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib45)\)\.
These criteria exclude inference\-time model combination and selection\. Ensembles aggregate predictions from multiple models, while routing systems such as RouteLLM select a source model for each query; both retain and invoke source models at inference time rather than constructing an inference\-independent target\([Chen et al\., 2026b](https://arxiv.org/html/2609.19553#bib.bib41);[Ong et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib42)\)\. LLM\-as\-a\-Judge is also excluded because it evaluates model outputs rather than constructing a target model through parameter combination, representation alignment, or behavior distillation\([Zheng et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib37)\)\. We further distinguish behavior distillation from reward\-only reinforcement learning\. On\-policy distillation qualifies when a teacher provides dense, token\-level behavioral supervision, such as output\-distribution targets or an equivalent KL\-constrained signal, on target\-generated states\([Agarwal et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib76);[Yang et al\., 2026e](https://arxiv.org/html/2609.19553#bib.bib48)\); using an LLM only for a scalar reward does not qualify because it does not distill the teacher’s output behavior\.
Let𝒳\\mathcal\{X\}and𝒴\\mathcal\{Y\}be the input space and the output space\. Givennnsource models
𝒮=\{Misrc\}i=1n,\\mathcal\{S\}=\\\{M\_\{i\}^\{\\mathrm\{src\}\}\\\}\_\{i=1\}^\{n\},\(1\)whereMisrcM\_\{i\}^\{\\mathrm\{src\}\}is theii\-th source model\. For inputx∈𝒳x\\in\\mathcal\{X\}, each source model gives a conditional output distributionpisrc\(y∣x\)p\_\{i\}^\{\\mathrm\{src\}\}\(y\\mid x\), wherey∈𝒴y\\in\\mathcal\{Y\}\. The goal is to build a target modelMθtgtM\_\{\\theta\}^\{\\mathrm\{tgt\}\}with parametersθ\\theta\. Its conditional output distribution ispθtgt\(y∣x\)p\_\{\\theta\}^\{\\mathrm\{tgt\}\}\(y\\mid x\)\. Model fusion can be written as a mapping from source models to the target model:
θ=Φ\(𝒮,𝒟\)\.\\theta=\\Phi\(\\mathcal\{S\},\\mathcal\{D\}\)\.\(2\)Here,Φ\\Phiis the fusion mapping\.𝒟\\mathcal\{D\}denotes optional fusion data and can be empty\. The general fusion goal can be expressed as
θ⋆=argminθ∑i=1n𝔼x∼𝒯i\[Dout\(pisrc\(⋅∣x\),pθtgt\(⋅∣x\)\)\]\.\\begin\{split\}\\theta^\{\\star\}&=\\arg\\min\_\{\\theta\}\\sum\_\{i=1\}^\{n\}\\mathbb\{E\}\_\{x\\sim\\mathcal\{T\}\_\{i\}\}\\Bigl\[\\\\\[\-1\.99997pt\] &\\quad D\_\{\\mathrm\{out\}\}\\\!\\left\(p\_\{i\}^\{\\mathrm\{src\}\}\(\\cdot\\mid x\),p\_\{\\theta\}^\{\\mathrm\{tgt\}\}\(\\cdot\\mid x\)\\right\)\\Bigr\]\.\\end\{split\}\(3\)Here,θ⋆\\theta^\{\\star\}denotes the fused target parameters,𝒯i\\mathcal\{T\}\_\{i\}is the task distribution of sourceii, andDoutD\_\{\\mathrm\{out\}\}measures the discrepancy between the source and target output distributions\. Equation[3](https://arxiv.org/html/2609.19553#S2.E3)states a general retention goal rather than a shared training loss; a fusion method need not optimize it explicitly\.
#### Inference Independence\.
Onceθ\\thetais fixed, the target model no longer needs the source models during inference:
pθtgt\(y∣x,𝒮\)=pθtgt\(y∣x\)\.p\_\{\\theta\}^\{\\mathrm\{tgt\}\}\(y\\mid x,\\mathcal\{S\}\)=p\_\{\\theta\}^\{\\mathrm\{tgt\}\}\(y\\mid x\)\.
## 3Taxonomy of Model Fusion
\{forest\}
Figure 3:A taxonomy of model fusion for LLMs and MLLMs\. The method branches are organized by the main object being fused or aligned: parameters, representations, or behaviors\. AdaMerging, MWA, and Fisher Merging remain parameter\-level because their learned or estimated quantities determine parameter combinations rather than provide external source behavior\. The evaluation branch summarizes representative benchmark resources\. Methods and resources are illustrative rather than exhaustive\.We organize methods by the source signal used in fusion rather than the surface form of the final result: parameters, representations, or behaviors\.
### 3\.1Parameter\-Level Fusion
Definition\.Letθ0\\theta\_\{0\}be the reference parameters and, for each sourceMisrc∈𝒮M\_\{i\}^\{\\mathrm\{src\}\}\\in\\mathcal\{S\}, letθi\\theta\_\{i\}be its parameters andΔi=θi−θ0\\Delta\_\{i\}=\\theta\_\{i\}\-\\theta\_\{0\}its update\. Parameter\-level fusion transforms and combines these updates:
ΦP\(𝒮,𝒟\)=θ0\+∑i=1nαiAi\(Δi\)\.\\Phi\_\{\\mathrm\{P\}\}\(\\mathcal\{S\},\\mathcal\{D\}\)=\\theta\_\{0\}\+\\sum\_\{i=1\}^\{n\}\\alpha\_\{i\}A\_\{i\}\(\\Delta\_\{i\}\)\.\(4\)Here,αi∈ℝ\\alpha\_\{i\}\\in\\mathbb\{R\}is a scalar source coefficient, andAiA\_\{i\}can be the identity, masking or rescaling, a permutation, or a subspace projection; both can be determined without data or estimated from𝒟\\mathcal\{D\}\.
Related work and methods\.Parameter\-level fusion methods can be organized by how they manipulate parameters\.*Arithmetic rules*combine source weights or parameter deltas with fixed or lightly tuned coefficients\. SWA\([Izmailov et al\., 2018](https://arxiv.org/html/2609.19553#bib.bib9)\)averages parameter snapshots sampled along an SGD trajectory with a cyclical or constant learning rate, approximating an ensemble with a single model and improving generalization with little additional cost\. Model soups\([Wortsman et al\., 2022](https://arxiv.org/html/2609.19553#bib.bib70)\)show that multiple fine\-tuned models can be averaged when they lie in a nearby parameter basin, while task arithmetic\([Ilharco et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib20)\)represents the difference between a fine\-tuned model and its base model as a task vector, enabling capability composition or behavior editing through vector addition and subtraction\. Fisher Merging weights individual parameters using sample\-estimated Fisher information, while MWA weights checkpoints using training metrics such as loss or training step\([Matena and Raffel, 2022](https://arxiv.org/html/2609.19553#bib.bib62);[Yu and Choi, 2025](https://arxiv.org/html/2609.19553#bib.bib78)\)\. Subsequent methods further address conflicts and redundancy among parameter deltas\. For example, TIES\-Merging, DARE, and DELLA\-Merging reduce interference through sign consistency, random dropping or rescaling\([Yadav et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib98);[Yu et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib46);[Deep et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib14)\)\.
*Subspace methods*identify, reshape, or constrain structured directions in weight or update space to improve alignment and reduce interference\. SVD\-based methods use singular directions to reshape update spaces, separate shared and task\-specific components, and reduce interference\([Stoica et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib69);[Marczak et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib74);[Gargiulo et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib96)\)\. For continual fusion, DOP\([Yang et al\., 2025b](https://arxiv.org/html/2609.19553#bib.bib117)\)approximates unavailable data subspaces with SVD subspaces of task vectors and applies dual orthogonal projections to balance stability and plasticity without accessing task data\.
*Optimization methods*formulate merge coefficients, task vectors, or transformation variables as explicit optimization problems\. AdaMerging learns task\- or layer\-wise fusion coefficients by minimizing output entropy on unlabeled data; the resulting coefficients still combine source parameters rather than distilling source behaviors\([Yang et al\., 2024c](https://arxiv.org/html/2609.19553#bib.bib3)\)\. AWD\([Xiong et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib72)\)optimizes a decomposition of task vectors into redundant and disentangled components, improving orthogonality while preserving task\-specific performance\. WUDI\([Cheng et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib105)\)uses task vectors to locate interference sources and correct the responsible components without data\. GCWM\([Wang et al\., 2026f](https://arxiv.org/html/2609.19553#bib.bib116)\)and DOGE\([Wei et al\., 2025b](https://arxiv.org/html/2609.19553#bib.bib71)\)use geometric or projected\-gradient objectives to reduce interference during multi\-task fusion\.
*Module fusion*combines LoRA, adapters, projectors, or other pluggable modules instead of full\-model parameters, suiting parameter\-efficient fine\-tuning\. AdapterSoup averages domain adapters, whereas LoRA soups average, concatenate, or weight skill\-specific modules\([Chronopoulou et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib5);[Hu et al\., 2022](https://arxiv.org/html/2609.19553#bib.bib52);[Prabhakar et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib53)\)\.
Parameter\-level fusion is efficient but requires source compatibility and interference control\.
Figure 4:Three levels of model fusion and their practical trade\-offs\. Data demand and source access indicate typical relative requirements rather than strict rules\. A pipeline may combine levels, but its primary level is determined by the signal used to construct or update the target model\.
### 3\.2Representation\-Level Fusion
Definition\.Representation\-level fusion uses intermediate representations as the main signal for capability integration\. For eachMisrc∈𝒮M\_\{i\}^\{\\mathrm\{src\}\}\\in\\mathcal\{S\}, letr~iℓ\(x\)\\widetilde\{r\}\_\{i\}^\{\\ell\}\(x\)denote its representation after any required layer matching, normalization, or dimensional projection, and letrθℓ\(x\)r\_\{\\theta\}^\{\\ell\}\(x\)be the target representation, whereℓ∈ℒ\\ell\\in\\mathcal\{L\}indexes the matched layers\. We write𝒟=\(𝒟i\)i=1n\\mathcal\{D\}=\(\\mathcal\{D\}\_\{i\}\)\_\{i=1\}^\{n\}, where𝒟i\\mathcal\{D\}\_\{i\}is the fusion\-data distribution for sourceii\. Representation\-level fusion matches these signals as
ΦR\(𝒮,𝒟\)=argminθ∑ℓ∈ℒ∑i=1n𝔼x∼𝒟i\[DR\(r~iℓ\(x\),rθℓ\(x\)\)\]\.\\begin\{split\}\\Phi\_\{\\mathrm\{R\}\}\(\\mathcal\{S\},\\mathcal\{D\}\)&=\\arg\\min\_\{\\theta\}\\sum\_\{\\ell\\in\\mathcal\{L\}\}\\sum\_\{i=1\}^\{n\}\\\\\[\-1\.99997pt\] &\\quad\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{i\}\}\\Bigl\[D\_\{\\mathrm\{R\}\}\\\!\\left\(\\widetilde\{r\}\_\{i\}^\{\\ell\}\(x\),r\_\{\\theta\}^\{\\ell\}\(x\)\\right\)\\Bigr\]\.\\end\{split\}\(5\)The discrepancyDRD\_\{\\mathrm\{R\}\}is instantiated according to the matched signal, such asL1L\_\{1\}or MSE for feature tensors, cosine or correlation discrepancy for feature directions or unit matching, and MMD or moment discrepancy for representation distributions or activation calibration\. Representation statistics on𝒟i\\mathcal\{D\}\_\{i\}can also determineαi\\alpha\_\{i\}orAiA\_\{i\}in Equation[4](https://arxiv.org/html/2609.19553#S3.E4); such methods remain representation\-level when these statistics are the main fusion signal\.
Related work and methods\.Representation\-level fusion asks how intermediate representations can guide the construction or repair of a target model\. Existing methods mainly use representations in three ways: to derive merge signals, solve local matching problems, or train repair and distillation objectives\.
*Weighting methods*compute fusion weights from representations and then combine models in parameter space\. These weights can be defined over parameters, layers, modules, or matched components\. AIM\([Heyrani Nobari et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib2)\)estimates weight saliency from activation magnitudes on a task\-agnostic calibration set\. MAGIC\([Li et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib54)\)calibrates representation and weight magnitudes, while Merging Beyond\([Yao et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib61)\)uses activation subspaces to form rotation\-aware updates\. Related alignment methods compute correspondence from representations before fusion: REPAIR\([Jordan et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib55)\)rescales preactivations, ZipIt\!\([Stoica et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib108)\)matches units by activation similarity, and Transformer Fusion\([Imfeld et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib102)\)aligns Transformer components with optimal transport\. These methods are efficient, but they depend on calibration data, layer correspondence, and reliable representation similarity\.
*Closed\-form solvers*formulate representation matching as local regression problems and solve them analytically, which is most practical for linear modules\. RegMean\([Jin et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib10)\)uses input covariance to merge each linear module so that its output matches source\-module outputs\. RegMean\+\+\([Nguyen et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib83)\)improves this local view by adding intra\-layer and cross\-layer dependencies\. LOT\-Merging\([Sun et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib100)\)and FeatCal\([Gu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib101)\)further treat representation drift as the main target: the former derives layer\-wise analytic updates, while the latter calibrates merged weights in forward order by separating upstream propagation from local mismatch\. For LoRA fusion, LoRM\([Salami et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib36)\)applies output matching to low\-rank modules, and IterIS\([Chen et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib35)\)refines the matching objective through iterative inference\-solving\. Compared with weighting methods, these solvers use representations more directly by fitting local matching objectives, not only by estimating fusion weights\.
*Backpropagation methods*train the target model or added repair modules with representation losses, allowing nonlinear repair and the use of multiple internal signals such as hidden states and attention maps\. Patient Knowledge Distillation\([Sun et al\., 2019](https://arxiv.org/html/2609.19553#bib.bib79)\)augments output distillation with hidden\-state matching at selected source layers, while TinyBERT\([Jiao et al\., 2020](https://arxiv.org/html/2609.19553#bib.bib99)\)combines prediction\-layer supervision with matching of embeddings, attention maps, and hidden states\. TED\([Liang et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib49)\)likewise augments output distillation with task\-aware filters that retain task\-relevant hidden representations before alignment\. Fusion repair methods use the same idea after an initial parameter\-level fusion step\. Representation Surgery\([Yang et al\., 2024a](https://arxiv.org/html/2609.19553#bib.bib84)\)learns a lightweight module to correct final\-layer representation bias\. Surgeryv2\([Yang et al\., 2024b](https://arxiv.org/html/2609.19553#bib.bib91)\)extends this repair across multiple layers\. ProbSurgery\([Wei et al\., 2025a](https://arxiv.org/html/2609.19553#bib.bib85)\)models the correction as a distribution to capture uncertainty from parameter interference\. Compared with closed\-form solvers, these methods can handle more complex mismatch, but they require slower training, larger repair data, and careful regularization to avoid overfitting\.
Representation\-level fusion is most useful when hidden states expose drift or layer mismatch\. Weighting methods are inexpensive but sensitive to calibration; closed\-form solvers need aligned linear modules and more representation samples; backpropagation handles complex mismatch but requires more training and data\. Open problems include efficient drift repair and alignment across dissimilar models\.
### 3\.3Behavior\-Level Fusion
Definition\.Behavior\-level fusion uses observable source behaviors to train a target model\. Letssdenote an input state, such as\(x,y<t\)\(x,y\_\{<t\}\)for an autoregressive model, and letμ𝒟\\mu\_\{\\mathcal\{D\}\}be the empirical state distribution from fixed or target\-generated sequences\. Letq𝒮\(⋅∣s\)q\_\{\\mathcal\{S\}\}\(\\cdot\\mid s\)be a behavior target constructed by selecting, aggregating, or aligning source outputs\. Output\-distribution distillation uses
ΦB\(𝒮,𝒟\)=argminθ𝔼s∼μ𝒟\[DB\(q𝒮\(⋅∣s\),pθtgt\(⋅∣s\)\)\]\.\\begin\{split\}\\Phi\_\{\\mathrm\{B\}\}\(\\mathcal\{S\},\\mathcal\{D\}\)&=\\arg\\min\_\{\\theta\}\\mathbb\{E\}\_\{s\\sim\\mu\_\{\\mathcal\{D\}\}\}\\Bigl\[\\\\\[\-1\.99997pt\] &\\quad D\_\{\\mathrm\{B\}\}\\\!\\left\(q\_\{\\mathcal\{S\}\}\(\\cdot\\mid s\),p\_\{\\theta\}^\{\\mathrm\{tgt\}\}\(\\cdot\\mid s\)\\right\)\\Bigr\]\.\\end\{split\}\(6\)Here,pθtgt\(⋅∣s\)p\_\{\\theta\}^\{\\mathrm\{tgt\}\}\(\\cdot\\mid s\)is the target output distribution\. Common choices forDBD\_\{\\mathrm\{B\}\}include forward KL, reverse KL, and generalized JSD\. Source models therefore act as*behavior providers*, rather than parameter or representation providers\. Methods using only target\-model entropy, confidence, or uncertainty for parameter fusion lack external behavioral supervision and fall outside this category\.
Related work and methods\.Behavior\-level fusion can be grouped by the transferred behavior type into distribution fusion, demonstration fusion, and feedback fusion, with an orthogonal distinction between off\-policy supervision on fixed data and on\-policy supervision on target\-induced states\.
*Distribution fusion*transfers source\-provided soft labels, output distributions, token probabilities, or logits\. Classical knowledge distillation matches a teacher’s softened output distribution\([Hinton et al\., 2015](https://arxiv.org/html/2609.19553#bib.bib18)\), while DistilBERT shows its effectiveness for language model compression\([Sanh et al\., 2019](https://arxiv.org/html/2609.19553#bib.bib17)\)\. For model fusion, such signals can integrate complementary capabilities across models: InfiGFusion further models logits as relational graphs and aligns their geometry via an efficient Gromov–Wasserstein approximation, moving beyond independent token\-level matching\([Wang et al\., 2025b](https://arxiv.org/html/2609.19553#bib.bib110)\)\.*Demonstration fusion*learns from source\-generated responses, rationales, reasoning traces, tool\-use traces, or trajectories\. Instruction distillation and GPT4All\-style training use stronger\-model outputs to train independently deployable targets\([Sun et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib34);[Anand et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib29)\), while rationale or step\-level distillation transfers intermediate reasoning processes\([Hsieh et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib19);[Magister et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib97)\)\. These methods require only sampled outputs, but may inherit source errors, spurious reasoning, or stylistic bias\.
*Feedback fusion*transfers preferences, scores, critiques, corrections, reward signals, or verifier labels, making it useful for alignment and safety transfer when parameters, hidden states, or full distributions are unavailable\. Chat\-oriented fusion can construct data from multi\-source responses, rankings, and preferences, as in FuseChat and Zephyr\([Wan et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib25);[Tunstall et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib107)\)\. InfiFPO further formulates fusion as implicit preference optimization, absorbing source\-model advantages without direct pivot model access\([Gu et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib111)\)\. Source\-derived preferences, critiques, or verifier feedback qualify when they transfer identifiable source behavior, whereas generic reward\-only RL without such behavioral transfer is excluded\.
From the state\-distribution perspective,*off\-policy fusion*uses fixed behavior data and is simple to scale, but suffers from mismatch when the target visits poorly covered states\.*On\-policy fusion*instead lets the target generate prefixes, responses, or trajectories, and then obtains supervision on these target\-induced states\. This connects to dataset aggregation in imitation learning\([Ross et al\., 2011](https://arxiv.org/html/2609.19553#bib.bib82)\); in LLMs, GKD instantiates it by distilling from teacher feedback on student\-generated sequences\([Agarwal et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib76)\)\. Recent OPD variants study self\-distillation, black\-box or semi\-on\-policy supervision, offline logit reuse, token\-efficient supervision, and stabilization\([Zhao et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib90);[Chen et al\., 2026a](https://arxiv.org/html/2609.19553#bib.bib112);[Wu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib113);[Xu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib114);[Luo et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib115)\), and extend OPD to multimodal trajectories such as video grounding and speech LLM alignment\([Li et al\., 2026a](https://arxiv.org/html/2609.19553#bib.bib104);[Cao et al\., 2026a](https://arxiv.org/html/2609.19553#bib.bib106)\)\. This formulation is especially suitable for heterogeneous fusion, where source and target models may differ in architecture, tokenizer, modality interface, decoding policy, or capability profile\.
Behavior\-level fusion suits closed\-source and heterogeneous models because it avoids parameter and hidden\-state access\. Demonstrations may carry imitation bias; distributions often require logit access; feedback depends on verifier or reward quality\. Open problems include robust multi\-source aggregation, budget\-aware on\-policy queries, and joint process\- and outcome\-level feedback\.
### 3\.4Evaluation
#### Metrics\.
Evaluation for model fusion can start from two simple metrics\. \(1\)Avg performancereports the average performance of the target model over the task pool\. It gives a direct view of overall quality and is easy to compare across methods\. \(2\)Normalized performancecompares the target model with the corresponding source model on each task\. MergeBench uses this metric to measure how much source task performance is retained by the target model\([He et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib57)\)\. This is important because a target model can improve the average score while losing one source capability\. Other metrics cover interference, generalization, alignment, cost, and safety when applicable\. Appendix Table[11](https://arxiv.org/html/2609.19553#A7.T11)summarizes these metrics\.
#### Benchmarks\.
Model fusion benchmarks involve more than a task leaderboard\. They usually define a model pool and a task pool, so methods can be compared under shared source and evaluation settings\. FusionBench\([Tang et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib26)\)gives unified settings for comparing many parameter\-level fusion methods across model and task pools\. MergeBench\([He et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib57)\)focuses on domain source models and reports retention, generalization, and cost\. Appendix Table[10](https://arxiv.org/html/2609.19553#A6.T10)compares representative resources by modality coverage, model pool, task pool, heterogeneity, fusion type, and evaluation focus\. The comparison shows that current resources still mainly support parameter\-level fusion\. Representation\-level fusion often relies on drift analysis in method papers\. Behavior\-level fusion often borrows task, response, or safety benchmarks from distillation studies\([Xu et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib92);[Song and Zheng, 2026a](https://arxiv.org/html/2609.19553#bib.bib94)\)\. Future benchmarks should share source settings and report drift, behavior transfer, judge protocols, and fusion cost\.
MethodLevelAvg\.RLVR Expert Fusion: Qwen3\-4B, 5 domainsTIESP61\.00TA\+DAREP60\.99MT\-OPDB60\.46Domain Expert Fusion: Llama\-3\.1\-8B, 5\-domain Acc\. \(%\)Task ArithmeticP48\.7TIESP46\.8RegMeanR46\.3DAREP45\.2Instruct\+Code: AlpacaEval 2\.0, HumanEval, MBPPTask Arithmetic \(0 shots\)P0\.2448ProDistill \(16 shots\)R0\.2367ProDistill \(32 shots\)R0\.2380ProDistill \(64 shots\)R0\.2496
CLIP values are average accuracy \(%\) under matched 16\-shot settings\. B32 and L14 denote ViT\-B/32 and ViT\-L/14; 8, 14, and 20 denote TA\-8, TALL\-14, and TALL\-20\.
Table 2:Within\-paper evidence for Takeaway 2 from M2RL\([Wang et al\., 2026a](https://arxiv.org/html/2609.19553#bib.bib120)\), MergeBench\([He et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib57)\), ProDistill\([Xu et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib88)\), and SAMerging\([Dalili and Mahdavi, 2025](https://arxiv.org/html/2609.19553#bib.bib68)\)\. Values are reported point estimates, and ProDistill shot counts are shown with the method names\.P,RandBdenote parameter\-, representation\-, and behavior\-level fusion\.Table 3:Four within\-paper comparisons for Takeaway 3 from FeatCal\([Gu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib101)\), Patient KD\([Sun et al\., 2019](https://arxiv.org/html/2609.19553#bib.bib79)\), TED\([Liang et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib49)\), and TinyBERT\([Jiao et al\., 2020](https://arxiv.org/html/2609.19553#bib.bib99)\)\. Each row compares a single\-level baseline with its hybrid counterpart under the same paper setting; scores are not compared across rows\.
## 4Practical Takeaways
❶Fusion methods should be selected under practical constraints\.Figure[4](https://arxiv.org/html/2609.19553#S3.F4)compares their signals, data, and access needs\. The key choice is which signal is available and which failure mode is most likely\. When source models share architecture, initialization, and tokenizer, and their weights are available, parameter\-level fusion is often a simple first option\. When drift or internal loss appears, representation\-level fusion can use hidden states to find and calibrate layer mismatch\. If only responses are available or models differ, behavior\-level fusion is practical\.
❷Cross\-level rankings vary by setting and budget\.Because models, settings, and budgets differ across papers, Table[2](https://arxiv.org/html/2609.19553#S3.T2)uses only within\-paper comparisons\. In M2RL, parameter\-level TIES and TA\+DARE are close to behavior\-level MT\-OPD \(61\.00 and 60\.99 versus 60\.46\), while MT\-OPD uses 967\.9 GPU\-hours of training\. In MergeBench, representation\-level RegMean \(46\.3\) lies between parameter\-level TIES \(46\.8\) and DARE \(45\.2\)\. On Instruct\+Code, ProDistill trails Task Arithmetic at 16 and 32 shots but slightly exceeds it at 64 shots \(0\.2496 versus 0\.2448\)\. Under matched 16\-shot CLIP settings, SAMerging has higher reported point estimates than ProDistill in five of six configurations, whereas ProDistill is higher on L14/20\. Thus, cross\-level rankings depend on the evaluation setting, data access, and compute budget\.
❸Combining fusion levels can yield a stronger practical pipeline\.Different levels address complementary failure modes: parameter fusion provides a low\-cost target, representation fusion corrects drift or mismatch, and behavior supervision recovers missing outputs when cheaper signals are insufficient\. Across four within\-paper comparisons, adding a second fusion level improves the reported score: FeatCal raises Task Arithmetic from 63\.5 to 65\.8, Patient KD raises RACE test accuracy from 58\.74 to 60\.34, TED raises the GLUE development average from 87\.0 to 87\.5, and TinyBERT raises the three\-task GLUE development average from 73\.5 to 75\.6\. Together, the four comparisons spanP\+R,B\+R, andR\+Bpipelines; they do not imply universal gains\.
## 5Challenges and Future Directions
❶Unclear Theoretical Foundations and Applicability Conditions\.Model fusion is already used in practice: DeepSeek\-V4 independently trains domain experts and consolidates them through on\-policy distillation\([DeepSeek\-AI et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib13)\)\. Yet fusion quality is usually known only after the target model has been built and evaluated, and no general theory predicts applicability across all three fusion levels\. Source and method selection therefore still rely heavily on trial and error, increasing cost and the risk of capability loss\. Recent theory explains specific cases:[Li et al\. \(2026b\)](https://arxiv.org/html/2609.19553#bib.bib121)analyzes parameter averaging under heterogeneous fine\-tuning throughL2L\_\{2\}\-stability, while other work studies shared initialization, nearby loss basins, or hidden\-unit alignment\([Wortsman et al\., 2022](https://arxiv.org/html/2609.19553#bib.bib70);[Ainsworth et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib28);[Zhou et al\., 2026a](https://arxiv.org/html/2609.19553#bib.bib15)\)\. These results do not yet provide comparable conditions for representation alignment or behavior distillation\. Theory should link outcomes to source compatibility, fusion data, and target capacity, and identify settings where source capabilities may not be retained\.
❷Difficulty in Aligning Heterogeneous Source Models\.Source models can differ in architecture, parameterization, tokenizer, or modality interface, leaving no direct correspondence between their components\. This breaks parameter correspondence at the parameter level, makes hidden\-state matching harder at the representation level, and complicates behavior distillation when output spaces or reasoning styles differ\. Transport and Merge\([Cui et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib103)\)uses optimal transport for cross\-architecture LLM fusion, while AdaMMS\([Du et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib4)\)learns coefficients for heterogeneous MLLMs\. Without reliable alignment, source capabilities can interfere rather than combine\. Future work should align architectures and translate representations across model families and domains\.
❸Evaluating Multi\-Source Capability Retention\.Model fusion succeeds only if one target model retains the capabilities of multiple sources\. Aggregate scores can hide source\-specific capability loss and interference, making an unsuccessful fusion appear effective\. Appendix Table[10](https://arxiv.org/html/2609.19553#A6.T10)surveys 14 benchmarks and related evaluation resources, but reusable multi\-method protocols remain concentrated in a small subset\. FusionBench\([Tang et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib26)\)and MergeBench\([He et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib57)\)provide broad evaluation settings, whereas most other resources focus on narrower tasks, individual methods, empirical studies, or search and deployment tooling\. Future benchmarks should use shared source settings and report per\-capability retention, worst\-task degradation, fusion cost, and compliance with the single\-model inference condition\.
❹Preventing Risk Transfer During Fusion\.Unsafe behavior can survive parameter merging\([Hammoud et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib64)\); LoRATK\([Liu et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib50)\)and Merge Hijacking\([Yuan et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib56)\)further show how malicious updates and backdoors can enter a fused model\. Backdoors can also transfer through behavior distillation\([Wang and Zhao, 2025](https://arxiv.org/html/2609.19553#bib.bib122);[De Muri et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib123)\)and remain in the target model after the source models have been removed at inference time\. Fusion also creates privacy, ownership, and collaboration risks: Merger\-as\-a\-Stealer\([Lu et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib60)\)studies private information leakage, while Among Us\([Yang et al\., 2026g](https://arxiv.org/html/2609.19553#bib.bib7)\)studies malicious contributions in model collaboration\. Because consolidating these attacks may lower the barrier to reproduction or misuse, future work should combine source screening, provenance tracking, contribution attribution, and post\-fusion safety testing with defensive mechanisms such as MergeGuard\([Cong et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib31)\)\. Appendix[H](https://arxiv.org/html/2609.19553#A8)discusses forgetting and deployment cost\.
## 6Conclusion
We define model fusion through three operations: parameter combination, representation alignment, and behavior distillation\. The fused target must run without the complete source models at inference\. This boundary supports a three\-level taxonomy and exposes level\-specific problems in compatibility, capability retention, and risk transfer\.
No fusion level dominates across the reviewed settings\. Comparisons must account for source compatibility, available signals, fusion data, and compute budgets\. Multi\-level pipelines can combine an inexpensive initial merge with representation repair or behavior supervision, but the reported gains remain setting dependent\. Progress therefore requires applicability conditions for each level, reliable alignment across heterogeneous sources, evaluation of per\-source capability retention, and safeguards against risk transfer\.
## Limitations
This survey may miss recent fusion work, especially fast\-moving preprints and industrial systems with limited public details\. Relevant papers may also be overlooked because the topic appears as model merging or knowledge transfer\. We collected papers from surveys, benchmarks, and method papers and repeatedly checked their taxonomy and references, but errors may remain\. Some methods combine multiple levels\. Our labels identify the main signal used to construct or update the target rather than mutually exclusive pipeline classes, and boundary cases may admit alternative readings\.
Our benchmark summary relies on reported results and may not account for model scale, data access, or tuning budget\. The cross\-paper numbers in Section 4 are illustrative rather than controlled head\-to\-head comparisons, and some sources report neither variance nor significance\. The conclusions are therefore setting\-dependent, not universal rankings among fusion levels\. Reported fusion costs are likewise not independently normalized across hardware and implementation settings\.
## Acknowledgments
This paper is fully supported by five grants from the Research Grants Council of the Hong Kong Special Administrative Region, China \(No\. 15215325, 15208824, 15228325, 25208626, T41\-517/25\-N\) and an Innovation and Technology Fund from the Innovation and Technology Commission of the Hong Kong Special Administrative Region, China \(No\. ITP/003/26LP\)\.
## References
- Agarwalet al\.\(2024\)R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. Ramos Garea, M\. Geist, and O\. BachemOn\-Policy Distillation of Language Models: Learning from Self\-Generated Mistakes\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 21246–21263\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/5be69a584901a26c521c2b51e40a4c20-Abstract-Conference.html)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.4.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1),[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p3.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Ainsworthet al\.\(2023\)S\. K\. Ainsworth, J\. Hayase, and S\. S\. SrinivasaGit Re\-Basin: Merging Models modulo Permutation Symmetries\.InProceedings of the International Conference on Learning Representations, ICLR 2023,External Links:[Link](https://arxiv.org/abs/2209.04836)Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p4.1),[§5](https://arxiv.org/html/2609.19553#S5.p1.1)\.
- Akizukiet al\.\(2025\)R\. Akizuki, Y\. Kudo, N\. Yoshinari, Y\. Hirose, T\. Nishimoto, K\. Uchida, and S\. ShirakawaSurrogate Benchmarks for Model Merging Optimization\.InAutoML 2025 Non\-Archival Content Track,External Links:[Link](https://arxiv.org/abs/2509.02555)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.8.1.1.1)\.
- Anandet al\.\(2023\)Y\. Anand, Z\. Nussbaum, B\. Duderstadt, B\. Schmidt, and A\. MulyarGPT4All: Training an Assistant\-Style Chatbot with Large Scale Data Distillation from GPT\-3\.5\-Turbo\.Technical reportNomic AI\.External Links:[Link](https://s3.amazonaws.com/static.nomic.ai/gpt4all/2023_GPT4All_Technical_Report.pdf)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.11.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Biggset al\.\(2024\)B\. Biggs, A\. Seshadri, Y\. Zou, A\. Jain, A\. Golatkar, Y\. Xie, A\. Achille, A\. Swaminathan, and S\. SoattoDiffusion Soup: Model Merging for Text\-to\-Image Diffusion Models\.InComputer Vision – ECCV 2024,Lecture Notes in Computer Science, Vol\.15121,pp\. 257–274\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-73036-8%5F15),[Link](https://doi.org/10.1007/978-3-031-73036-8_15)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.11.1)\.
- Caoet al\.\(2026a\)D\. Cao, D\. Fu, H\. Yu, S\. Zheng, X\. Tan, and T\. JinX\-OPD: Cross\-Modal On\-Policy Distillation for Capability Alignment in Speech LLMs\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.24596),[Link](https://arxiv.org/abs/2603.24596),2603\.24596Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.11.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Caoet al\.\(2026b\)Y\. Cao, D\. Ran, Y\. Guo, M\. Wu, S\. Chen, L\. Li, W\. Yang, and T\. XieAn Empirical Study and Theoretical Explanation on Task\-Level Model\-Merging Collapse\.CoRRabs/2603\.09463\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.09463),[Link](https://arxiv.org/abs/2603.09463),2603\.09463Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p4.1)\.
- Ceritliet al\.\(2025\)T\. Ceritli, O\. Bohdal, M\. Ozay, J\. Moon, K\. Lee, H\. Ko, and U\. MichieliHydraOpt: navigating the efficiency\-performance trade\-off of adapter merging\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 26887–26909\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1365),[Link](https://aclanthology.org/2025.emnlp-main.1365/)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.14.1)\.
- Chenet al\.\(2025\)H\. Chen, Z\. Wang, R\. Li, B\. Zhu, and L\. ChenIterIS: Iterative Inference\-Solving Alignment for LoRA Merging\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 4829–4838\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00455),[Link](https://openaccess.thecvf.com/content/CVPR2025/html/Chen_IterIS_Iterative_Inference-Solving_Alignment_for_LoRA_Merging_CVPR_2025_paper.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.14.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p4.1)\.
- Chenet al\.\(2026a\)X\. Chen, J\. Wang, W\. Zhu, P\. Qiu, X\. Dong, H\. Sang, Z\. Wang, A\. Geramifard, and F\. LuoSODA: Semi On\-Policy Black\-Box Distillation for Large Language Models\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.03873),[Link](https://arxiv.org/abs/2604.03873v3),2604\.03873v3Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.6.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Chenet al\.\(2026b\)Z\. Chen, X\. Lu, J\. Li, P\. Chen, Z\. Li, K\. Sun, Y\. Luo, Q\. Mao, M\. Li, L\. Xiao, D\. Yang, X\. Huang, Y\. Ban, H\. Sun, and P\. S\. YuHarnessing multiple large language models: a survey on llm ensemble\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.18036),[Link](https://arxiv.org/abs/2502.18036),2502\.18036Cited by:[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p3.1)\.
- Chenget al\.\(2025\)R\. Cheng, F\. Xiong, Y\. Wei, W\. Zhu, and C\. YuanWhoever Started the Interference Should End It: Guiding Data\-Free Model Merging via Task Vectors\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 10121–10143\.External Links:[Link](https://proceedings.mlr.press/v267/cheng25h.html)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.5.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p4.1)\.
- Chronopoulouet al\.\(2023\)A\. Chronopoulou, M\. Peters, A\. Fraser, and J\. DodgeAdapterSoup: weight averaging to improve generalization of pretrained language models\.InFindings of the Association for Computational Linguistics: EACL 2023,pp\. 2054–2063\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-eacl.153),[Link](https://aclanthology.org/2023.findings-eacl.153/)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.13.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p5.1)\.
- Conget al\.\(2024\)T\. Cong, D\. Ran, Z\. Liu, X\. He, J\. Liu, Y\. Gong, Q\. Li, A\. Wang, and X\. WangHave You Merged My Model? On The Robustness of Large Language Model IP Protection Methods Against Model Merging\.InProceedings of the 1st ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis,pp\. 69–76\.External Links:[Document](https://dx.doi.org/10.1145/3689217.3690614),[Link](https://arxiv.org/abs/2404.05188)Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Cuiet al\.\(2026\)C\. Cui, B\. Yang, F\. Shen, Y\. Chen, J\. Zheng, X\. Wang, A\. Zhang, and T\. ChuaTransport and Merge: Cross\-Architecture Merging for Large Language Models\.CoRRabs/2602\.05495\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.05495),[Link](https://arxiv.org/abs/2602.05495),2602\.05495Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p4.1),[§5](https://arxiv.org/html/2609.19553#S5.p2.1)\.
- Dalili and Mahdavi \(2025\)S\. A\. Dalili and M\. MahdaviModel Merging via Multi\-Teacher Knowledge Distillation\.CoRRabs/2512\.21288\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.21288),[Link](https://arxiv.org/abs/2512.21288),2512\.21288Cited by:[Table 2](https://arxiv.org/html/2609.19553#S3.T2)\.
- De Muriet al\.\(2026\)G\. De Muri, M\. Vero, R\. Staab, and M\. VechevPay Attention to the Triggers: Constructing Backdoors That Survive Distillation\.InICLR 2026 Workshop on Agents in the Wild,External Links:[Link](https://arxiv.org/abs/2510.18541),2510\.18541Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Deepet al\.\(2024\)P\. T\. Deep, R\. Bhardwaj, and S\. PoriaDELLA\-Merging: Reducing Interference in Model Merging through Magnitude\-Based Sampling\.CoRRabs/2406\.11617\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.11617),[Link](https://arxiv.org/abs/2406.11617),2406\.11617Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.10.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- DeepSeek\-AIet al\.\(2026\)DeepSeek\-AI, A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling, C\. Lu, C\. Zhao, C\. Deng, C\. Hou, C\. Xu, C\. Shao, C\. Ruan, C\. Sun, D\. Dai, D\. Guo, D\. Yang, D\. Chen, D\. Li, D\. Ji, E\. Li, F\. Wei, F\. Lin, F\. Yuan, F\. Xia, F\. Dai, G\. Hao, G\. Chen, G\. Cao, G\. Meng, G\. Li, H\. Yu, H\. Zhang, H\. Xu, H\. Li, H\. Liang, H\. Zhang, H\. Luo, H\. Wei, H\. Yuan, H\. Zhang, H\. Luo, H\. Chen, H\. Ji, H\. Zhang, H\. Ding, H\. Tang, H\. Cao, H\. Gao, H\. Qu, H\. Zeng, J\. Yang, J\. Zhu, J\. Luo, J\. Song, J\. Yu, J\. Huang, J\. Cai, J\. Liang, J\. Zhou, J\. Ye, J\. Li, J\. Xu, J\. Hu, J\. Yang, J\. Chen, J\. Yan, J\. Chen, J\. Zhou, J\. Xiang, J\. Yuan, J\. Cheng, J\. Zhou, J\. Zhu, J\. Yu, J\. Sun, J\. Ran, J\. Jiang, J\. Qiu, J\. Li, J\. Zheng, J\. Song, K\. Dong, K\. Gao, K\. Guan, K\. Zhou, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Wang, L\. Xia, L\. Zhang, L\. Zhao, L\. Guo, L\. Luo, L\. Ma, L\. Zhu, L\. Wang, L\. Cai, L\. Zhang, L\. Chen, M\. Di, M\. Xu, M\. Mei, M\. Wang, M\. Zhang, M\. Zhang, M\. Tang, M\. Li, M\. Zhou, M\. Han, N\. Wang, P\. Huang, P\. Wang, P\. Cong, P\. Wang, P\. Zhang, Q\. Wang, Q\. Zhu, Q\. Li, Q\. Chen, Q\. Du, Q\. Jiang, R\. Tian, R\. Xu, R\. Lu, R\. Xu, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. Chen, R\. Yin, R\. Xu, R\. Shen, R\. Zhang, R\. Chen, S\. Liu, S\. Lu, S\. Sun, S\. Zhou, S\. Chen, S\. Cai, S\. Nie, S\. Wu, S\. Chen, S\. Hu, S\. Liu, S\. Hu, S\. Ma, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. Yu, S\. Zhou, T\. Ni, T\. Yun, T\. Jin, T\. Pei, T\. Ye, T\. Lin, T\. Ji, T\. Cui, T\. Yue, T\. Yu, T\. Wang, W\. Zhang, W\. Xiao, W\. Zeng, W\. An, W\. Zhao, W\. Liu, W\. Liang, W\. Pang, W\. Luo, W\. Yao, W\. Gao, W\. Yang, W\. Huang, W\. Hou, W\. Zhang, W\. Ma, X\. Gao, X\. He, X\. Wang, X\. Wang, X\. Bi, X\. Liu, X\. Wang, X\. Chen, X\. Zhang, X\. Nie, X\. Sun, X\. Wang, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Liu, X\. Yu, X\. Li, X\. Yang, X\. Zhang, X\. Chen, X\. Wang, X\. Su, X\. Chen, X\. Lin, X\. Fu, Y\. Yan, Y\. Wang, Y\. Ma, Y\. Luo, Y\. Zhang, Y\. Xu, Y\. Ma, Y\. Huang, Y\. Li, Y\. Li, Y\. Xu, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Qian, Y\. Shao, Y\. Yu, Y\. Zhang, Y\. Ding, Y\. Shi, Y\. Wu, Y\. Xiong, Y\. Ma, Y\. He, Y\. Tang, Y\. Zhou, Y\. Luo, Y\. Zhong, Y\. Piao, Y\. Wang, Y\. Zhang, Y\. Chen, Y\. Tan, Y\. Wei, Y\. Ma, Y\. Liu, Y\. Yang, Y\. Guo, Y\. Wu, Y\. Wu, Y\. Li, Y\. Cheng, Y\. Ou, Y\. Xu, Y\. Li, Y\. Wang, Y\. Yang, Y\. Xu, Y\. Wu, Y\. Meng, Y\. Zou, Y\. Zha, Y\. Xiong, Y\. Chen, Y\. Lin, Y\. Cao, Y\. Wang, Y\. Zhang, Y\. Yan, Y\. Lin, Y\. Gu, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. Zhou, Y\. Huang, Z\. Wu, Z\. Wang, Z\. Zhao, Z\. Ren, Z\. Zhang, Z\. Sha, Z\. Fu, Z\. Ju, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Gao, Z\. Hao, Z\. Gou, Z\. Ma, Z\. Yan, Z\. Shao, Z\. Huang, Z\. Chen, Z\. Wu, Z\. Ren, Z\. Wu, Z\. Li, Z\. Zhang, Z\. Xu, Z\. Wang, Z\. Qu, Z\. Gu, Z\. Zhu, Z\. Li, Z\. Zhang, Z\. Xie, Z\. Gao, Z\. Wan, Z\. Pan, and Z\. YaoDeepSeek\-V4: Towards Highly Efficient Million\-Token Context Intelligence\.Technical reportDeepSeek\-AI\.External Links:2606\.19348,[Document](https://dx.doi.org/10.48550/arXiv.2606.19348),[Link](https://arxiv.org/abs/2606.19348)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1),[§5](https://arxiv.org/html/2609.19553#S5.p1.1)\.
- Djuheraet al\.\(2025\)A\. Djuhera, S\. R\. Kadhe, F\. Ahmed, S\. Zawad, and H\. BocheSafeMERGE: preserving safety alignment in fine\-tuned large language models via selective layer\-wise model merging\.External Links:2503\.17239,[Document](https://dx.doi.org/10.48550/arXiv.2503.17239),[Link](https://arxiv.org/abs/2503.17239)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px3.p1.1)\.
- Dmonteet al\.\(2026\)A\. Dmonte, V\. Gupta, D\. J\. Perry, and M\. ArehartImproving training efficiency and reducing maintenance costs via language specific model merging\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 5: Industry Track\),pp\. 562–570\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-industry.43),[Link](https://aclanthology.org/2026.eacl-industry.43/)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1)\.
- Duet al\.\(2025\)Y\. Du, X\. Wang, C\. Chen, J\. Ye, Y\. Wang, P\. Li, M\. Yan, J\. Zhang, F\. Huang, Z\. Sui, M\. Sun, and Y\. LiuAdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025,pp\. 9413–9422\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.00879),[Link](https://doi.org/10.1109/CVPR52734.2025.00879)Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p4.1),[§5](https://arxiv.org/html/2609.19553#S5.p2.1)\.
- Fanget al\.\(2026\)L\. Fang, X\. Yu, J\. Cai, Y\. Chen, S\. Wu, Z\. Liu, Z\. Yang, H\. Lu, X\. Gong, Y\. Liu, T\. Ma, W\. Ruan, A\. Abbasi, J\. Zhang, T\. Wang, E\. Latif, W\. Liu, W\. Zhang, S\. Kolouri, X\. Zhai, D\. Zhu, W\. Zhong, T\. Liu, and P\. MaKnowledge distillation and dataset distillation of large language models: emerging trends, challenges, and future directions\.Artificial Intelligence Review59\(1\),pp\. 17\.External Links:[Document](https://dx.doi.org/10.1007/s10462-025-11423-3),[Link](https://doi.org/10.1007/s10462-025-11423-3)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.11.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p5.1)\.
- Fenget al\.\(2025\)Y\. Feng, J\. Li, X\. Dong, P\. Xu, X\. Zhou, Y\. Zhang, Z\. Lu, Y\. Wang, A\. Zhao, X\. Chu, and X\. WuAIMMerging: adaptive iterative model merging using training trajectories for language model continual learning\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 13420–13437\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.678),[Link](https://aclanthology.org/2025.emnlp-main.678/)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px1.p1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1)\.
- Gargiuloet al\.\(2025\)A\. A\. Gargiulo, D\. Crisostomi, M\. S\. Bucarelli, S\. Scardapane, F\. Silvestri, and E\. RodolàTask Singular Vectors: Reducing Task Interference in Model Merging\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 18695–18705\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52734.2025.01742),[Link](https://doi.org/10.1109/CVPR52734.2025.01742)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.16.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p3.1)\.
- GLM\-5\-Teamet al\.\(2026\)GLM\-5\-Team, A\. Zeng, X\. Lv, Z\. Hou, Z\. Du, Q\. Zheng, B\. Chen, D\. Yin, C\. Ge, C\. Huang, C\. Xie, C\. Zhu, C\. Yin, C\. Wang, G\. Pan, H\. Zeng, H\. Zhang, H\. Wang, H\. Chen, J\. Zhang, J\. Jiao, J\. Guo, J\. Wang, J\. Du, J\. Wu, K\. Wang, L\. Li, L\. Fan, L\. Zhong, M\. Liu, M\. Zhao, P\. Du, Q\. Dong, R\. Lu, Shuang\-Li, S\. Cao, S\. Liu, T\. Jiang, X\. Chen, X\. Zhang, X\. Huang, X\. Dong, Y\. Xu, Y\. Wei, Y\. An, Y\. Niu, Y\. Zhu, Y\. Wen, Y\. Cen, Y\. Bai, Z\. Qiao, Z\. Wang, Z\. Wang, Z\. Zhu, Z\. Liu, Z\. Li, B\. Wang, B\. Wen, C\. Huang, C\. Cai, C\. Yu, C\. Li, C\. Hu, C\. Zhang, D\. Zhang, D\. Lin, D\. Yang, D\. Wang, D\. Ai, E\. Zhu, F\. Yi, F\. Chen, G\. Wen, H\. Sun, H\. Zhao, H\. Hu, H\. Zhang, H\. Liu, H\. Zhang, H\. Peng, H\. Tai, H\. Zhang, H\. Liu, H\. Wang, H\. Yan, H\. Ge, H\. Liu, H\. Chu, J\. Zhao, J\. Wang, J\. Zhao, J\. Ren, J\. Wang, J\. Zhang, J\. Gui, J\. Zhao, J\. Li, J\. An, J\. Li, J\. Yuan, J\. Du, J\. Liu, J\. Zhi, J\. Duan, K\. Zhou, K\. Wei, K\. Wang, K\. Luo, L\. Zhang, L\. Sha, L\. Xu, L\. Wu, L\. Ding, L\. Chen, M\. Li, N\. Lin, P\. Ta, Q\. Zou, R\. Song, R\. Yang, S\. Tu, S\. Yang, S\. Wu, S\. Zhang, S\. Li, S\. Li, S\. Fan, W\. Qin, W\. Tian, W\. Zhang, W\. Yu, W\. Liang, X\. Kuang, X\. Cheng, X\. Li, X\. Yan, X\. Hu, X\. Ling, X\. Fan, X\. Xia, X\. Zhang, X\. Zhang, X\. Pan, X\. Zou, X\. Zhang, Y\. Liu, Y\. Wu, Y\. Li, Y\. Wang, Y\. Zhu, Y\. Tan, Y\. Zhou, Y\. Pan, Y\. Zhang, Y\. Su, Y\. Geng, Y\. Yan, Y\. Tan, Y\. Bi, Y\. Shen, Y\. Yang, Y\. Li, Y\. Liu, Y\. Wang, Y\. Li, Y\. Wu, Y\. Zhang, Y\. Duan, Y\. Zhang, Z\. Liu, Z\. Jiang, Z\. Yan, Z\. Zhang, Z\. Wei, Z\. Chen, Z\. Feng, Z\. Yao, Z\. Chai, Z\. Wang, Z\. Zhang, B\. Xu, M\. Huang, H\. Wang, J\. Li, Y\. Dong, and J\. TangGLM\-5: from Vibe Coding to Agentic Engineering\.External Links:2602\.15763,[Document](https://dx.doi.org/10.48550/arXiv.2602.15763),[Link](https://arxiv.org/abs/2602.15763)Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p3.1)\.
- Goddardet al\.\(2024\)C\. Goddard, S\. Siriwardhana, M\. Ehghaghi, L\. Meyers, V\. Karpukhin, B\. Benedict, M\. McQuade, and J\. SolawetzArcee’s MergeKit: a toolkit for merging large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 477–485\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-industry.36),[Link](https://aclanthology.org/2024.emnlp-industry.36/)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.4.1.1.1),[Table 10](https://arxiv.org/html/2609.19553#A6.T10.3.5.2.1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px2.p1.1)\.
- Gouet al\.\(2021\)J\. Gou, B\. Yu, S\. J\. Maybank, and D\. TaoKnowledge distillation: a survey\.Int\. J\. Comput\. Vision129\(6\),pp\. 1789–1819\.External Links:ISSN 0920\-5691,[Link](https://doi.org/10.1007/s11263-021-01453-z),[Document](https://dx.doi.org/10.1007/s11263-021-01453-z)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.2.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p5.1)\.
- Guet al\.\(2026\)Y\. Gu, S\. Cai, Z\. Wang, W\. Wang, Y\. Wang, P\. Wang, S\. Huang, S\. Lu, J\. Wu, and H\. YangFeatCal: Feature Calibration for Post\-Merging Models\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.13030),[Link](https://arxiv.org/abs/2605.13030),2605\.13030Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1),[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.16.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p4.1),[Table 3](https://arxiv.org/html/2609.19553#S3.T3)\.
- Guet al\.\(2025\)Y\. Gu, Y\. Wang, Z\. Yan, Y\. Zhang, Q\. Zhou, F\. Wu, and H\. YangInfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 15645–15672\.External Links:[Document](https://dx.doi.org/10.52202/085713-0529),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/1730c56d7d415ec0fb1f9fa70b271c8f-Abstract-Conference.html)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.17.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p4.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Ding, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Chen, J\. Yuan, J\. Tu, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. You, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. ZhangDeepSeek\-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning\.Nature645\(8081\),pp\. 633–638\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09422-z),[Link](https://doi.org/10.1038/s41586-025-09422-z)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px4.p1.1)\.
- Hammoudet al\.\(2024\)H\. A\. A\. K\. Hammoud, U\. Michieli, F\. Pizzati, P\. Torr, A\. Bibi, B\. Ghanem, and M\. OzayModel merging and safety alignment: one bad model spoils the bunch\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 13033–13046\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.762),[Link](https://aclanthology.org/2024.findings-emnlp.762/)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Heet al\.\(2025\)Y\. He, S\. Zeng, Y\. Hu, R\. Yang, T\. Zhang, and H\. ZhaoMergeBench: A Benchmark for Merging Domain\-Specialized LLMs\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Document](https://dx.doi.org/10.52202/085713-3374),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/91f7f71cb04699f387dc863da42a1fe3-Abstract-Datasets_and_Benchmarks_Track.html)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.12.1.1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p4.1),[§3\.4](https://arxiv.org/html/2609.19553#S3.SS4.SSS0.Px1.p1.1),[§3\.4](https://arxiv.org/html/2609.19553#S3.SS4.SSS0.Px2.p1.1),[Table 2](https://arxiv.org/html/2609.19553#S3.T2),[§5](https://arxiv.org/html/2609.19553#S5.p3.1)\.
- Heet al\.\(2026\)Z\. He, R\. Ding, Z\. Huang, R\. Yang, T\. Li, and X\. HuangCompress then merge: from multiple LoRAs into one low\-rank adapter\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://icml.cc/virtual/2026/poster/61546)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.20.1)\.
- Heyrani Nobariet al\.\(2025\)A\. Heyrani Nobari, K\. Alimohammadi, A\. ArjomandBigdeli, A\. Srivastava, F\. Ahmed, and N\. AzizanActivation\-Informed Merging of Large Language Models\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 72975–72998\.External Links:[Document](https://dx.doi.org/10.52202/085713-2446),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/69b576be5219a18195f2dba1ae1ffceb-Abstract-Conference.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.6.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p3.1)\.
- Hintonet al\.\(2015\)G\. Hinton, O\. Vinyals, and J\. DeanDistilling the Knowledge in a Neural Network\.External Links:1503\.02531,[Document](https://dx.doi.org/10.48550/arXiv.1503.02531),[Link](https://arxiv.org/abs/1503.02531)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.3.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Hititet al\.\(2025\)O\. K\. Hitit, L\. Girrbach, and Z\. AkataA systematic study of in\-the\-wild model merging for large language models\.External Links:2511\.21437,[Document](https://dx.doi.org/10.48550/arXiv.2511.21437),[Link](https://arxiv.org/abs/2511.21437)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.9.1.1.1)\.
- Hsiehet al\.\(2023\)C\. Hsieh, C\. Li, C\. Yeh, H\. Nakhost, Y\. Fujii, A\. Ratner, R\. Krishna, C\. Lee, and T\. PfisterDistilling step\-by\-step\! outperforming larger language models with less training data and smaller model sizes\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 8003–8017\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.507),[Link](https://aclanthology.org/2023.findings-acl.507/)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.12.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: Low\-Rank Adaptation of Large Language Models\.InProceedings of the International Conference on Learning Representations, ICLR 2022,External Links:[Link](https://arxiv.org/abs/2106.09685)Cited by:[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p5.1)\.
- Huet al\.\(2026\)J\. Hu, D\. Zhu, X\. Luo, D\. Zhang, S\. He, Y\. Lei, H\. Zheng, S\. Feng, J\. He, Y\. Sun, H\. Wu, and H\. WangCORD: bridging the audio–text reasoning gap via weighted on\-policy cross\-modal distillation\.External Links:2601\.16547,[Document](https://dx.doi.org/10.48550/arXiv.2601.16547),[Link](https://arxiv.org/abs/2601.16547)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.12.1.1.1)\.
- Huanget al\.\(2024\)C\. Huang, P\. Ye, T\. Chen, T\. He, X\. Yue, and W\. OuyangEMR\-Merging: Tuning\-Free High\-Performance Model Merging\.InAdvances in Neural Information Processing Systems 37, NeurIPS 2024,Vol\.37,pp\. 122741–122769\.External Links:[Document](https://dx.doi.org/10.52202/079017-3900),[Link](https://doi.org/10.52202/079017-3900)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.5.1.1.1)\.
- Ilharcoet al\.\(2023\)G\. Ilharco, M\. T\. Ribeiro, M\. Wortsman, S\. Gururangan, L\. Schmidt, H\. Hajishirzi, and A\. FarhadiEditing models with task arithmetic\.InProceedings of the International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2212.04089)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.7.1),[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p2.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- Imfeldet al\.\(2024\)M\. Imfeld, J\. Graldi, M\. Giordano, T\. Hofmann, S\. Anagnostidis, and S\. P\. SinghTransformer Fusion with Optimal Transport\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 29313–29336\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/7e0af0d1bc0ec2a90fc294be2e00447e-Abstract-Conference.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.5.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p3.1)\.
- Izmailovet al\.\(2018\)P\. Izmailov, D\. Podoprikhin, T\. Garipov, D\. P\. Vetrov, and A\. G\. WilsonAveraging Weights Leads to Wider Optima and Better Generalization\.InProceedings of the Conference on Uncertainty in Artificial Intelligence, UAI 2018,pp\. 876–885\.External Links:[Link](https://auai.org/uai2018/proceedings/papers/313.pdf)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.3.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- Jianget al\.\(2026\)T\. Jiang, X\. Yu, C\. Yi, Y\. Wu, Y\. Li, R\. Cheng, D\. Jiang, and J\. ZhangEvoGM: learning to merge LLMs via evolutionary generative optimization\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://icml.cc/virtual/2026/poster/64090)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.9.1)\.
- Jiaoet al\.\(2020\)X\. Jiao, Y\. Yin, L\. Shang, X\. Jiang, X\. Chen, L\. Li, F\. Wang, and Q\. LiuTinyBERT: distilling BERT for natural language understanding\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 4163–4174\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.372),[Link](https://aclanthology.org/2020.findings-emnlp.372/)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.19.1.1.1),[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.6.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p5.1),[Table 3](https://arxiv.org/html/2609.19553#S3.T3)\.
- Jinet al\.\(2026\)W\. Jin, T\. Min, Y\. Yang, D\. Wei, Y\. Zhou, S\. R\. Kadhe, N\. Baracaldo, and K\. LeeEntropy\-aware on\-policy distillation of language models\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://icml.cc/virtual/2026/poster/64855)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.13.1.1.1)\.
- Jinet al\.\(2023\)X\. Jin, X\. Ren, D\. Preotiuc\-Pietro, and P\. ChengDataless Knowledge Fusion by Merging Weights of Language Models\.InProceedings of the International Conference on Learning Representations, ICLR 2023,External Links:[Link](https://arxiv.org/abs/2212.09849)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1),[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.11.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p4.1)\.
- Jordanet al\.\(2023\)K\. Jordan, H\. Sedghi, O\. Saukh, R\. Entezari, and B\. NeyshaburREPAIR: Renormalizing Permuted Activations for Interpolation Repair\.InProceedings of the International Conference on Learning Representations, ICLR 2023,External Links:[Link](https://arxiv.org/abs/2211.08403)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.3.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p3.1)\.
- Khademet al\.\(2026\)F\. Khadem, S\. Mousavi, Y\. Fang, and Y\. LiuDP\-OPD: differentially private on\-policy distillation for language models\.External Links:2604\.04461,[Document](https://dx.doi.org/10.48550/arXiv.2604.04461),[Link](https://arxiv.org/abs/2604.04461)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.20.1.1.1)\.
- Khalifaet al\.\(2024\)M\. Khalifa, Y\. Tan, A\. Ahmadian, T\. Hosking, H\. Lee, L\. Wang, A\. Üstün, T\. Sherborne, and M\. GalléIf You Can’t Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs\.CoRRabs/2412\.04144\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2412.04144),[Link](https://arxiv.org/abs/2412.04144),2412\.04144Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.6.1.1.1)\.
- Leeet al\.\(2026\)W\. Lee, W\. Jeong, and K\. YoonLabel\-free cross\-task LoRA merging with null\-space compression\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 847–859\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2026/html/Lee_Label-Free_Cross-Task_LoRA_Merging_with_Null-Space_Compression_CVPR_2026_paper.html)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.19.1)\.
- Leeet al\.\(2025\)Y\. Lee, C\. Ko, T\. Pedapati, I\. Chung, M\. Yeh, and P\. ChenSTAR: spectral truncation and rescale for model merging\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 2: Short Papers\),pp\. 496–505\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-short.42),[Link](https://aclanthology.org/2025.naacl-short.42/)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.17.1)\.
- Leiet al\.\(2026\)H\. Lei, Y\. Li, H\. Zhang, S\. Zhang, Q\. Cheng, X\. Qu, G\. Cui, B\. Zhou, N\. Ding, Y\. Luo, and Y\. ChengDraft\-OPD: on\-policy distillation for speculative draft models\.External Links:2605\.29343,[Document](https://dx.doi.org/10.48550/arXiv.2605.29343),[Link](https://arxiv.org/abs/2605.29343)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.22.1.1.1)\.
- Liet al\.\(2026a\)J\. Li, H\. Yin, H\. Xu, B\. Xu, W\. Tan, Z\. He, J\. Ju, Z\. Luo, and J\. LuanVideo\-OPD: Efficient Post\-Training of Multimodal Large Language Models for Temporal Video Grounding via On\-Policy Distillation\.CoRRabs/2602\.02994\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.02994),[Link](https://arxiv.org/abs/2602.02994),2602\.02994Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.10.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Liet al\.\(2026b\)Q\. Li, A\. Tang, M\. Zhang, M\. Wang, Q\. Yin, and L\. ShenA Unified Generalization Framework for Model Merging: Trade\-offs, Non\-Linearity, and Scaling Laws\.Note:arXiv preprint arXiv:2601\.21690External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.21690),[Link](https://arxiv.org/abs/2601.21690),2601\.21690Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p1.1)\.
- Liet al\.\(2026c\)W\. Li, Y\. Peng, M\. Zhang, L\. Ding, H\. Hu, and L\. ShenDeep model fusion: a survey\.IEEE Transactions on Neural Networks and Learning Systems37\(5\),pp\. 2008–2024\.External Links:[Document](https://dx.doi.org/10.1109/TNNLS.2025.3628666),[Link](https://doi.org/10.1109/TNNLS.2025.3628666)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.9.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p2.1)\.
- Liet al\.\(2025\)Y\. Li, J\. Zhang, J\. Guo, Z\. Cheng, L\. Qi, Y\. Shi, and Y\. GaoMAGIC: achieving superior model merging via magnitude calibration\.CoRRabs/2512\.19320\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.19320),[Link](https://arxiv.org/abs/2512.19320),2512\.19320Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.8.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p3.1)\.
- Lianget al\.\(2023\)C\. Liang, S\. Zuo, Q\. Zhang, P\. He, W\. Chen, and T\. ZhaoLess is more: Task\-aware layer\-wise distillation for language model compression\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 20852–20867\.External Links:[Link](https://proceedings.mlr.press/v202/liang23j.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.20.1.1.1),[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.7.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p5.1),[Table 3](https://arxiv.org/html/2609.19553#S3.T3)\.
- Lin and Kang \(2026\)L\. Lin and H\. KangUnlocking the potential of continual model merging: an ODE perspective\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://icml.cc/virtual/2026/poster/63291)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.10.1)\.
- Liuet al\.\(2025\)H\. Liu, S\. Zhong, X\. Sun, M\. Tian, M\. Hariri, Z\. Liu, R\. Tang, Z\. Jiang, J\. Yuan, Y\. Chuang, L\. Li, S\. Choi, R\. Chen, V\. Chaudhary, and X\. HuLoRATK: LoRA Once, Backdoor Everywhere in the Share\-and\-Play Ecosystem\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 23009–23047\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1253),[Link](https://aclanthology.org/2025.findings-emnlp.1253/)Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Liuet al\.\(2026\)Y\. Liu, J\. Lou, X\. Guan, Y\. Ji, H\. Lin, B\. He, X\. Han, L\. Sun, X\. Yu, and Y\. LuYour teacher can’t help you here: combating supervision fidelity decay in on\-policy distillation\.External Links:2605\.30833,[Document](https://dx.doi.org/10.48550/arXiv.2605.30833),[Link](https://arxiv.org/abs/2605.30833)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.17.1.1.1)\.
- Luet al\.\(2025\)L\. Lu, Z\. Zuo, Z\. Sheng, and P\. ZhouMerger\-as\-a\-stealer: stealing targeted PII from aligned LLMs with model merging\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 5806–5825\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.295),[Link](https://aclanthology.org/2025.emnlp-main.295/)Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Luoet al\.\(2026\)F\. Luo, Y\. Chuang, G\. Wang, Z\. Xu, X\. Han, T\. Zhang, and V\. BravermanDemystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models\.arXiv preprint arXiv:2604\.08527\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.08527),[Link](https://arxiv.org/abs/2604.08527),2604\.08527Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.9.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Magisteret al\.\(2023\)L\. C\. Magister, J\. Mallinson, J\. Adamek, E\. Malmi, and A\. SeverynTeaching small language models to reason\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 1773–1781\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-short.151),[Link](https://aclanthology.org/2023.acl-short.151/)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.13.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Marczaket al\.\(2025\)D\. Marczak, S\. Magistri, S\. Cygert, B\. Twardowski, A\. D\. Bagdanov, and J\. V\. D\. WeijerNo Task Left Behind: Isotropic Model Merging with Common and Task\-Specific Subspaces\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 43177–43199\.External Links:[Link](https://proceedings.mlr.press/v267/marczak25a.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.15.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p3.1)\.
- Matena and Raffel \(2022\)M\. S\. Matena and C\. RaffelMerging Models with Fisher\-Weighted Averaging\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 17703–17716\.External Links:[Document](https://dx.doi.org/10.52202/068431-1287),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/70c26937fbf3d4600b69a129031b66ec-Abstract-Conference.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.6.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- Minutet al\.\(2025\)A\. R\. Minut, T\. Mencattini, A\. Santilli, D\. Crisostomi, and E\. RodolàMergenetic: a simple evolutionary model merging library\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 572–582\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-demo.55),[Link](https://aclanthology.org/2025.acl-demo.55/)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.10.1.1.1),[Table 10](https://arxiv.org/html/2609.19553#A6.T10.3.5.2.1.1)\.
- Miyano and Arase \(2025\)R\. Miyano and Y\. AraseAdaptive LoRA merge with parameter pruning for low\-resource generation\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 19353–19366\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.990),[Link](https://aclanthology.org/2025.findings-acl.990/)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.18.1)\.
- Nguyenet al\.\(2025\)T\. Nguyen, D\. Huu\-Tien, T\. Suzuki, and L\. NguyenRegMean\+\+: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging\.External Links:2508\.03121,[Document](https://dx.doi.org/10.48550/arXiv.2508.03121),[Link](https://arxiv.org/abs/2508.03121)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.12.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p4.1)\.
- Niuet al\.\(2026\)Y\. Niu, H\. Xiao, D\. Liu, Z\. Wang, D\. Gong, Y\. Wang, and J\. LiBreaking the tokenizer barrier: on\-policy distillation across model families\.External Links:2606\.09456,[Document](https://dx.doi.org/10.48550/arXiv.2606.09456),[Link](https://arxiv.org/abs/2606.09456)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.24.1.1.1)\.
- Onget al\.\(2025\)I\. Ong, A\. Almahairi, V\. Wu, W\. Chiang, T\. Wu, J\. E\. Gonzalez, M\. W\. Kadous, and I\. StoicaRouteLLM: Learning to Route LLMs from Preference Data\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 34433–34448\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/5503a7c69d48a2f86fc00b3dc09de686-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p3.1)\.
- Prabhakaret al\.\(2025\)A\. Prabhakar, Y\. Li, K\. Narasimhan, S\. Kakade, E\. Malach, and S\. JelassiLoRA soups: merging LoRAs for practical skill composition tasks\.InProceedings of the 31st International Conference on Computational Linguistics: Industry Track,pp\. 644–655\.External Links:[Link](https://aclanthology.org/2025.coling-industry.55/)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px4.p1.1),[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.15.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p5.1)\.
- Qinet al\.\(2025\)L\. Qin, T\. Zhu, W\. Zhou, and P\. S\. YuKnowledge distillation in federated learning: a survey on long lasting challenges and new solutions\.International Journal of Intelligent Systems2025\(1\),pp\. 7406934\.External Links:ISSN 1098\-111X,[Link](https://arxiv.org/abs/2406.10861),[Document](https://dx.doi.org/10.1155/int/7406934)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.6.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p5.1)\.
- Qinet al\.\(2026\)R\. Qin, Q\. Wang, D\. Liu, Q\. Li, Z\. Wei, and W\. ShenMultilingual Safety Alignment via Self\-Distillation\.arXiv preprint arXiv:2605\.02971\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.02971),[Link](https://arxiv.org/abs/2605.02971),2605\.02971Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px3.p1.1)\.
- Qiuet al\.\(2026\)Z\. Qiu, L\. Wang, Y\. Cao, R\. Zhang, B\. Su, Y\. Xu, F\. Meng, L\. Xu, Q\. Wu, and H\. LiNull\-Space Filtering for Data\-Free Continual Model Merging: Preserving Stability, Promoting Plasticity\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 128492–128518\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/d0da30e312b75a3fffd9e9191f8bc1b0-Abstract-Conference.html)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px1.p1.1),[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.25.1.1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1)\.
- Rosset al\.\(2011\)S\. Ross, G\. J\. Gordon, and J\. A\. BagnellA Reduction of Imitation Learning and Structured Prediction to No\-Regret Online Learning\.InProceedings of the 14th International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.15,pp\. 627–635\.External Links:[Link](https://proceedings.mlr.press/v15/ross11a.html)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.3.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Salamiet al\.\(2025\)R\. Salami, P\. Buzzega, M\. Mosconi, J\. Bonato, L\. Sabetta, and S\. CalderaraClosed\-Form Merging of Parameter\-Efficient Modules for Federated Continual Learning\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 86612–86637\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/d79b8c42f61b51722992705022134ef3-Abstract-Conference.html)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px4.p1.1),[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.13.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p4.1)\.
- Sanhet al\.\(2019\)V\. Sanh, L\. Debut, J\. Chaumond, and T\. WolfDistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter\.Note:Accepted at the 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing \(NeurIPS 2019\); no archival proceedings record locatedExternal Links:1910\.01108,[Document](https://dx.doi.org/10.48550/arXiv.1910.01108),[Link](https://arxiv.org/abs/1910.01108)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.4.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Shenajet al\.\(2025\)D\. Shenaj, O\. Bohdal, T\. Ceritli, M\. Ozay, P\. Zanuttigh, and U\. MichieliK\-merge: online continual merging of adapters for on\-device large language models\.External Links:2510\.13537,[Document](https://dx.doi.org/10.48550/arXiv.2510.13537),[Link](https://arxiv.org/abs/2510.13537)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px1.p1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px1.p1.1)\.
- Shenfeldet al\.\(2026\)I\. Shenfeld, M\. Damani, J\. Hübotter, and P\. AgrawalSelf\-Distillation Enables Continual Learning\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.19897),[Link](https://arxiv.org/abs/2601.19897v1),2601\.19897v1Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px1.p1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px1.p1.1)\.
- Song and Zheng \(2026a\)M\. Song and M\. ZhengA survey of on\-policy distillation for large language models\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.00626),[Link](https://arxiv.org/abs/2604.00626),2604\.00626Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.10.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p5.1),[§3\.4](https://arxiv.org/html/2609.19553#S3.SS4.SSS0.Px2.p1.1)\.
- Song and Zheng \(2026b\)M\. Song and M\. ZhengModel merging in the era of large language models: methods, applications, and future directions\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.09938),[Link](https://arxiv.org/abs/2603.09938),2603\.09938Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.8.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p2.1)\.
- Stoicaet al\.\(2024\)G\. Stoica, D\. Bolya, J\. Bjorner, P\. Ramesh, T\. Hearn, and J\. HoffmanZipIt\! Merging Models from Different Tasks without Training\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 29215–29237\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/7dbb5bfab324e3b86af9bd0df15498dd-Abstract-Conference.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.4.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p3.1)\.
- Stoicaet al\.\(2025\)G\. Stoica, P\. Ramesh, B\. Ecsedi, L\. Choshen, and J\. HoffmanModel merging with SVD to tie the Knots\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 4501–4519\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/0d4f8a5109c5083b5307fcd0bddae7a7-Abstract-Conference.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.13.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p3.1)\.
- Sunet al\.\(2026a\)Q\. Sun, S\. Zhang, Y\. Chen, Y\. Xue, R\. Peng, and C\. ZhaoFrom “weak” signals to strong models: preference delta aggregation with LoRA merging\.External Links:2606\.00357,[Document](https://dx.doi.org/10.48550/arXiv.2606.00357),[Link](https://arxiv.org/abs/2606.00357)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.21.1)\.
- Sunet al\.\(2019\)S\. Sun, Y\. Cheng, Z\. Gan, and J\. LiuPatient knowledge distillation for BERT model compression\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 4323–4332\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1441),[Link](https://aclanthology.org/D19-1441/)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.18.1.1.1),[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.5.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p5.1),[Table 3](https://arxiv.org/html/2609.19553#S3.T3)\.
- Sunet al\.\(2023\)W\. Sun, Z\. Chen, X\. Ma, L\. Yan, S\. Wang, P\. Ren, Z\. Chen, D\. Yin, and Z\. RenInstruction Distillation Makes Large Language Models Efficient Zero\-shot Rankers\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2311.01555),[Link](https://arxiv.org/abs/2311.01555),2311\.01555Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.10.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Sunet al\.\(2025\)W\. Sun, Q\. Li, W\. Wang, Y\. Liu, Y\. Geng, and B\. LiTowards Minimizing Feature Drift in Model Merging: Layer\-wise Task Vector Fusion for Adaptive Knowledge Integration\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 172635–172663\.External Links:[Document](https://dx.doi.org/10.52202/085713-5745),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/fbb4f2269267de23e0240e1c01242981-Abstract-Conference.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.15.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p4.1)\.
- Sunet al\.\(2026b\)Y\. Sun, Z\. Hou, H\. Ma, Y\. Jia, J\. Fang, H\. Guo, H\. An, W\. Wang, and J\. WangResMerge: residual\-based spectral merging of large language models\.External Links:2606\.02252,[Document](https://dx.doi.org/10.48550/arXiv.2606.02252),[Link](https://arxiv.org/abs/2606.02252)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.21.1)\.
- Sunget al\.\(2023\)Y\. Sung, L\. Li, K\. Lin, Z\. Gan, M\. Bansal, and L\. WangAn empirical study of multimodal model merging\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 1563–1575\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.105),[Link](https://aclanthology.org/2023.findings-emnlp.105/)Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p4.1)\.
- Tamet al\.\(2024\)D\. Tam, Y\. Kant, B\. Lester, I\. Gilitschenski, and C\. RaffelRealistic Evaluation of Model Merging for Compositional Generalization\.CoRRabs/2409\.18314\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2409.18314),[Link](https://arxiv.org/abs/2409.18314),2409\.18314Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.2.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p4.1)\.
- Tanget al\.\(2025\)A\. Tang, L\. Shen, Y\. Luo, E\. Yang, H\. Hu, L\. Zhang, B\. Du, and D\. TaoFusionBench: a unified library and comprehensive benchmark for deep model fusion\.Journal of Machine Learning Research26\(307\),pp\. 1–38\.External Links:[Link](https://jmlr.org/papers/v26/25-1243.html)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.11.1.1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p4.1),[§3\.4](https://arxiv.org/html/2609.19553#S3.SS4.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.19553#S5.p3.1)\.
- Tekinet al\.\(2026\)S\. F\. Tekin, F\. Ilhan, S\. Hu, T\. Huang, Y\. Xu, Z\. Yahn, and L\. LiuH3Fusion: helpful, harmless, honest fusion of aligned LLMs\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6993–7013\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.329),[Link](https://aclanthology.org/2026.eacl-long.329/)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.7.1.1.1)\.
- Tunstallet al\.\(2024\)L\. Tunstall, E\. Beeching, N\. Lambert, N\. Rajani, K\. Rasul, Y\. Belkada, S\. Huang, L\. von Werra, C\. Fourrier, N\. Habib, N\. Sarrazin, O\. Sanseviero, A\. M\. Rush, and T\. WolfZephyr: Direct Distillation of LM Alignment\.InFirst Conference on Language Modeling,External Links:[Link](https://arxiv.org/abs/2310.16944)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.16.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p4.1)\.
- Wanet al\.\(2024\)F\. Wan, X\. Huang, D\. Cai, X\. Quan, W\. Bi, and S\. ShiKnowledge Fusion of Large Language Models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 18303–18322\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/4fd5cfd2e31bebbccfa5ffa354c04bdc-Abstract-Conference.html)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1),[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p2.1)\.
- Wanet al\.\(2025\)F\. Wan, L\. Zhong, Z\. Yang, R\. Chen, and X\. QuanFuseChat: knowledge fusion of chat models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 21618–21642\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1096),[Link](https://aclanthology.org/2025.emnlp-main.1096/)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.15.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p4.1)\.
- Wanget al\.\(2025a\)B\. Wang, H\. Wan, L\. Shi, C\. Yang, P\. He, Y\. Ma, H\. Han, W\. Li, T\. Tan, Y\. Li, F\. Liu, G\. Yifan, and S\. ZhangRECALL: REpresentation\-aligned catastrophic\-forgetting ALLeviation via hierarchical model merging\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 16381–16395\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.829),[Link](https://aclanthology.org/2025.emnlp-main.829/)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px1.p1.1),[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.24.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1)\.
- Wanget al\.\(2024\)H\. Wang, B\. Ping, S\. Wang, X\. Han, Y\. Chen, Z\. Liu, and M\. SunLoRA\-flow: dynamic LoRA fusion for large language models in generative tasks\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12871–12882\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.695),[Link](https://aclanthology.org/2024.acl-long.695/)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.16.1)\.
- Wanget al\.\(2026a\)H\. Wang, X\. Long, Z\. Li, Y\. Xu, T\. Li, and Y\. TangTo Mix or To Merge: Toward Multi\-Domain Reinforcement Learning for Large Language Models\.CoRRabs/2602\.12566\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.12566),[Link](https://arxiv.org/abs/2602.12566),2602\.12566Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.15.1.1.1),[Table 2](https://arxiv.org/html/2609.19553#S3.T2)\.
- Wanget al\.\(2026b\)J\. Wang, Y\. Liu, J\. Chen, X\. Hu, Q\. Zhang, Y\. Cao, J\. Wang, H\. Yang, Y\. Xie, and Q\. ChenMAD\-OPD: breaking the ceiling in on\-policy distillation via multi\-agent debate\.External Links:2605\.01347,[Document](https://dx.doi.org/10.48550/arXiv.2605.01347),[Link](https://arxiv.org/abs/2605.01347)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.21.1.1.1)\.
- Wanget al\.\(2026c\)J\. Wang, W\. Zhang, W\. Shi, Y\. Li, and J\. ChengTCOD: exploring temporal curriculum in on\-policy distillation for multi\-turn autonomous agents\.External Links:2604\.24005,[Document](https://dx.doi.org/10.48550/arXiv.2604.24005),[Link](https://arxiv.org/abs/2604.24005)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.15.1.1.1)\.
- Wang and Zhao \(2025\)S\. Wang and D\. ZhaoBackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine\-tuning\.Note:arXiv preprint arXiv:2511\.12046External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.12046),[Link](https://arxiv.org/abs/2511.12046),2511\.12046Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Wanget al\.\(2026d\)Y\. Wang, Y\. Gu, Z\. Wang, K\. Li, Y\. Yang, Z\. Yan, C\. Xie, J\. Wu, and H\. YangMergePipe: A Budget\-Aware Parameter Management System for Scalable LLM Merging\.CoRRabs/2602\.13273\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.13273),[Link](https://arxiv.org/abs/2602.13273),2602\.13273Cited by:[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2026e\)Y\. Wang, Y\. Gu, Y\. Zhang, Q\. Zhou, Z\. Yan, C\. Xie, X\. Wang, J\. Yuan, and H\. YangModel Merging Scaling Laws in Large Language Models\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2509.24244),[Link](https://arxiv.org/abs/2509.24244),2509\.24244Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.14.1.1.1)\.
- Wanget al\.\(2025b\)Y\. Wang, Z\. Yan, Y\. Zhang, Q\. Zhou, Y\. Gu, F\. Wu, and H\. YangInfiGFusion: Graph\-on\-Logits Distillation via Efficient Gromov\-Wasserstein for Model Fusion\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 119677–119713\.External Links:[Document](https://dx.doi.org/10.52202/085713-3996),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/ad7c21456da003836ef6be2261e940df-Abstract-Conference.html)Cited by:[Table 8](https://arxiv.org/html/2609.19553#A5.T8.2.8.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p3.1)\.
- Wanget al\.\(2026f\)Y\. Wang, Y\. Yang, S\. Lu, Y\. Gu, P\. Wang, W\. Wang, Z\. Yan, C\. Xie, J\. Wu, J\. Cao, S\. Cheung, and H\. YangGeometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post\-Training\.arXiv preprint arXiv:2605\.09608\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.09608),[Link](https://arxiv.org/abs/2605.09608),2605\.09608Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.7.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p4.1)\.
- Weiet al\.\(2025a\)Q\. Wei, S\. He, E\. Yang, T\. Liu, H\. Wang, L\. Feng, and B\. AnRepresentation Surgery in Model Merging with Probabilistic Modeling\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 66015–66032\.External Links:[Link](https://proceedings.mlr.press/v267/wei25c.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.23.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p5.1)\.
- Weiet al\.\(2026a\)Y\. Wei, R\. Cheng, W\. Jin, E\. Yang, L\. Shen, L\. Hou, S\. Du, C\. Yuan, X\. Cao, and D\. TaoOptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 61571–61595\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/64741c1e80cc3bbed205b1c2b040dfa8-Abstract-Conference.html)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.13.1.1.1)\.
- Weiet al\.\(2026b\)Y\. Wei, R\. Cheng, X\. Zhang, L\. Shen, C\. Yuan, P\. Cui, and D\. TaoClosed\-form spectral regularization for multi\-task model merging\.External Links:2606\.07289,[Document](https://dx.doi.org/10.48550/arXiv.2606.07289),[Link](https://arxiv.org/abs/2606.07289)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.11.1)\.
- Weiet al\.\(2025b\)Y\. Wei, A\. Tang, L\. Shen, Z\. Hu, C\. Yuan, and X\. CaoModeling Multi\-Task Model Merging as Adaptive Projective Gradient Descent\.InProceedings of the International Conference on Machine Learning, ICML 2025,Proceedings of Machine Learning Research, Vol\.267,pp\. 66178–66193\.External Links:[Link](https://proceedings.mlr.press/v267/wei25k.html)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.6.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p4.1)\.
- Wortsmanet al\.\(2022\)M\. Wortsman, G\. Ilharco, S\. Y\. Gadre, R\. Roelofs, R\. Gontijo\-Lopes, A\. S\. Morcos, H\. Namkoong, A\. Farhadi, Y\. Carmon, S\. Kornblith, and L\. SchmidtModel soups: averaging weights of multiple fine\-tuned models improves accuracy without increasing inference time\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 23965–23998\.External Links:[Link](https://proceedings.mlr.press/v162/wortsman22a.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.5.1),[§1](https://arxiv.org/html/2609.19553#S1.p4.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1),[§5](https://arxiv.org/html/2609.19553#S5.p1.1)\.
- Wuet al\.\(2026\)Y\. Wu, S\. Han, and H\. CaiLightning OPD: Efficient Post\-Training for Large Reasoning Models with Offline On\-Policy Distillation\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.13010),[Link](https://arxiv.org/abs/2604.13010),2604\.13010Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.7.1.1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Xionget al\.\(2024\)F\. Xiong, R\. Cheng, W\. Chen, Z\. Zhang, Y\. Guo, C\. Yuan, and R\. XuMulti\-Task Model Merging via Adaptive Weight Disentanglement\.CoRRabs/2411\.18729\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2411.18729),[Link](https://arxiv.org/abs/2411.18729),2411\.18729Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.4.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p4.1)\.
- Xuet al\.\(2025\)J\. Xu, J\. Li, and J\. ZhangScalable Model Merging with Progressive Layer\-wise Distillation\.InProceedings of the International Conference on Machine Learning, ICML 2025,Proceedings of Machine Learning Research, Vol\.267,pp\. 69451–69477\.External Links:[Link](https://proceedings.mlr.press/v267/xu25r.html)Cited by:[Table 2](https://arxiv.org/html/2609.19553#S3.T2)\.
- Xuet al\.\(2024\)X\. Xu, M\. Li, C\. Tao, T\. Shen, R\. Cheng, J\. Li, C\. Xu, D\. Tao, and T\. ZhouA Survey on Knowledge Distillation of Large Language Models\.CoRRabs/2402\.13116\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2402.13116),[Link](https://arxiv.org/abs/2402.13116),2402\.13116Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.3.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p5.1),[§3\.4](https://arxiv.org/html/2609.19553#S3.SS4.SSS0.Px2.p1.1)\.
- Xuet al\.\(2026\)Y\. Xu, H\. Sang, Z\. Zhou, R\. He, Z\. Wang, and A\. GeramifardTIP: Token Importance in On\-Policy Distillation\.arXiv preprint arXiv:2604\.14084\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.14084),[Link](https://arxiv.org/abs/2604.14084),2604\.14084Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.8.1.1.1),[Appendix H](https://arxiv.org/html/2609.19553#A8.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Yadavet al\.\(2025\)P\. Yadav, C\. Raffel, M\. Muqeeth, L\. Caccia, H\. Liu, T\. Chen, M\. Bansal, L\. Choshen, and A\. SordoniA Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning\.Transactions on Machine Learning Research\.External Links:[Link](https://arxiv.org/abs/2408.07057)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.4.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p2.1)\.
- Yadavet al\.\(2023\)P\. Yadav, D\. Tam, L\. Choshen, C\. A\. Raffel, and M\. BansalTIES\-Merging: Resolving Interference When Merging Models\.InAdvances in Neural Information Processing Systems 36,Vol\.36,pp\. 7093–7115\.External Links:[Document](https://dx.doi.org/10.52202/075280-0310),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/1644c9af28ab7916874f6fd6228a9bcf-Abstract-Conference.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.8.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- Yanget al\.\(2025a\)C\. Yang, Y\. Zhu, W\. Lu, Y\. Wang, Q\. Chen, C\. Gao, B\. Yan, and Y\. ChenSurvey on knowledge distillation for large language models: methods, evaluation, and application\.ACM Trans\. Intell\. Syst\. Technol\.16\(6\),pp\. 1–27\.External Links:ISSN 2157\-6904,[Document](https://dx.doi.org/10.1145/3699518),[Link](https://arxiv.org/abs/2407.01885)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.5.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p5.1)\.
- Yanget al\.\(2026a\)E\. Yang, L\. Shen, G\. Guo, X\. Wang, X\. Cao, J\. Zhang, and D\. TaoModel merging in llms, mllms, and beyond: methods, theories, applications, and opportunities\.ACM Comput\. Surv\.58\(8\),pp\. 1–41\.External Links:ISSN 0360\-0300,[Document](https://dx.doi.org/10.1145/3787849),[Link](https://arxiv.org/abs/2408.07666)Cited by:[Table 1](https://arxiv.org/html/2609.19553#S1.T1.2.7.1.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p2.1)\.
- Yanget al\.\(2024a\)E\. Yang, L\. Shen, Z\. Wang, G\. Guo, X\. Chen, X\. Wang, and D\. TaoRepresentation Surgery for Multi\-Task Model Merging\.InProceedings of the International Conference on Machine Learning, ICML 2024,Proceedings of Machine Learning Research, Vol\.235,pp\. 56332–56356\.External Links:[Link](https://proceedings.mlr.press/v235/yang24t.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.21.1.1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1),[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p2.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p5.1)\.
- Yanget al\.\(2024b\)E\. Yang, L\. Shen, Z\. Wang, G\. Guo, X\. Wang, X\. Cao, J\. Zhang, and D\. TaoSurgeryv2: Bridging the gap between model merging and multi\-task learning with deep representation surgery\.arXiv preprint arXiv:2410\.14389\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.14389),[Link](https://arxiv.org/abs/2410.14389),2410\.14389Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1),[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.22.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p5.1)\.
- Yanget al\.\(2025b\)E\. Yang, A\. Tang, L\. Shen, G\. Guo, X\. Wang, X\. Cao, and J\. ZhangContinual Model Merging without Data: Dual Projections for Balancing Stability and Plasticity\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 39275–39305\.External Links:[Document](https://dx.doi.org/10.52202/085713-1311),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/37d9f19150fce07bced2a81fc87d47a6-Abstract-Conference.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.14.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p3.1)\.
- Yanget al\.\(2024c\)E\. Yang, Z\. Wang, L\. Shen, S\. Liu, G\. Guo, X\. Wang, and D\. TaoAdaMerging: Adaptive Model Merging for Multi\-Task Learning\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 22743–22763\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/62868cc2fc1eb5cdf321d05b4b88510c-Abstract-Conference.html)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px2.p1.1),[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.3.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p4.1)\.
- Yanget al\.\(2026b\)E\. Yang, Q\. Yang, P\. Wang, A\. Tang, G\. Guo, X\. Cao, and L\. ShenMergOPT: a merge\-aware optimizer for robust model merging\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 48400–48425\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/500809109da5faa7641add2224e6bb92-Abstract-Conference.html)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.8.1)\.
- Yanget al\.\(2026c\)S\. Yang, G\. Zhu, B\. Song, H\. Wang, M\. Xia, X\. Zheng, Y\. Ma, Z\. Chen, W\. Wang, J\. Zhao, and G\. ChenOPRD: on\-policy representation distillation\.External Links:2606\.06021,[Document](https://dx.doi.org/10.48550/arXiv.2606.06021),[Link](https://arxiv.org/abs/2606.06021)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.26.1.1.1)\.
- Yanget al\.\(2026d\)S\. Yang, K\. Shi, and W\. LiuOrthogonal model merging\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://icml.cc/virtual/2026/poster/61531)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.19.1)\.
- Yanget al\.\(2026e\)W\. Yang, W\. Liu, R\. Xie, K\. Yang, S\. Yang, and Y\. LinLearning beyond Teacher: Generalized On\-Policy Distillation with Reward Extrapolation\.CoRRabs/2602\.12125\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.12125),[Link](https://arxiv.org/abs/2602.12125),2602\.12125Cited by:[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p3.1)\.
- Yanget al\.\(2026f\)Z\. Yang, Z\. Liu, Y\. Chen, W\. Dai, B\. Wang, S\. Lin, C\. Lee, Y\. Chen, D\. Jiang, J\. He, R\. Pi, G\. Lam, N\. Lee, A\. Bukharin, M\. Shoeybi, B\. Catanzaro, and W\. PingNemotron\-cascade 2: post\-training llms with cascade RL and multi\-domain on\-policy distillation\.CoRRabs/2603\.19220\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.19220),[Link](https://arxiv.org/abs/2603.19220),2603\.19220Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.19553#S1.p3.1)\.
- Yanget al\.\(2026g\)Z\. Yang, W\. Ding, S\. Feng, and Y\. TsvetkovAmong us: measuring and mitigating malicious contributions in model collaboration systems\.External Links:2602\.05176,[Document](https://dx.doi.org/10.48550/arXiv.2602.05176),[Link](https://arxiv.org/abs/2602.05176)Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Yaoet al\.\(2025\)Y\. Yao, S\. Liu, Z\. Liu, Q\. Li, M\. Liu, X\. Han, Z\. Guo, H\. Wu, and L\. SongActivation\-Guided Consensus Merging for Large Language Models\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 124041–124069\.External Links:[Document](https://dx.doi.org/10.52202/085713-4133),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/b395e7bd5eaf3e980386debf2e9d7f11-Abstract-Conference.html)Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.7.1.1.1)\.
- Yaoet al\.\(2026\)Y\. Yao, H\. Sheng, Q\. Lv, H\. Wu, S\. Liu, Z\. Liu, Z\. Liu, J\. Gao, H\. Tan, X\. Fu, H\. Bai, H\. C\. So, Z\. Guo, and L\. SongMerging Beyond: Streaming LLM Updates via Activation\-Guided Rotations\.CoRRabs/2602\.03237\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.03237),[Link](https://arxiv.org/abs/2602.03237),2602\.03237Cited by:[Table 7](https://arxiv.org/html/2609.19553#A4.T7.2.9.1.1.1),[§3\.2](https://arxiv.org/html/2609.19553#S3.SS2.p3.1)\.
- Yiet al\.\(2024\)X\. Yi, S\. Zheng, L\. Wang, X\. Wang, and L\. HeA Safety Realignment Framework via Subspace\-Oriented Model Fusion for Large Language Models\.Knowledge\-Based Systems306,pp\. 112701\.External Links:[Document](https://dx.doi.org/10.1016/j.knosys.2024.112701),[Link](https://doi.org/10.1016/j.knosys.2024.112701)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px3.p1.1)\.
- Yuet al\.\(2024\)L\. Yu, B\. Yu, H\. Yu, F\. Huang, and Y\. LiLanguage Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 57755–57775\.External Links:[Link](https://proceedings.mlr.press/v235/yu24p.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.9.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- Yu and Choi \(2025\)S\. J\. Yu and S\. ChoiParameter\-efficient checkpoint merging via metrics\-weighted averaging\.CoRRabs/2504\.18580\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2504.18580),[Link](https://arxiv.org/abs/2504.18580),2504\.18580Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.4.1),[§3\.1](https://arxiv.org/html/2609.19553#S3.SS1.p2.1)\.
- Yuanet al\.\(2025\)Z\. Yuan, Y\. Xu, J\. Shi, P\. Zhou, and L\. SunMerge hijacking: backdoor attacks to model merging of large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 32688–32703\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1571),[Link](https://aclanthology.org/2025.acl-long.1571/)Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p4.1)\.
- Zamanet al\.\(2024\)K\. Zaman, L\. Choshen, and S\. SrivastavaFuse to forget: bias reduction and selective memorization through model fusion\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 18763–18783\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1045),[Link](https://aclanthology.org/2024.emnlp-main.1045/)Cited by:[Appendix B](https://arxiv.org/html/2609.19553#A2.SS0.SSS0.Px3.p1.1)\.
- Zenget al\.\(2025\)F\. Zeng, H\. Guo, F\. Zhu, L\. Shen, and H\. TangRobustMerge: parameter\-efficient model merging for MLLMs with direction robustness\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 71071–71095\.External Links:[Document](https://dx.doi.org/10.52202/085713-2390),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/67101f97-Abstract-Conference.html)Cited by:[Table 6](https://arxiv.org/html/2609.19553#A3.T6.2.17.1)\.
- Zhanget al\.\(2026a\)D\. Zhang, Z\. Yang, S\. Janghorbani, J\. Han, A\. Ressler II, Q\. Qian, G\. D\. Lyng, S\. S\. Batra, and R\. E\. TillmanFast and effective on\-policy distillation from reasoning prefixes\.External Links:2602\.15260,[Document](https://dx.doi.org/10.48550/arXiv.2602.15260),[Link](https://arxiv.org/abs/2602.15260)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.19.1.1.1)\.
- Zhanget al\.\(2026b\)H\. Zhang, Z\. Zhou, M\. Luo, S\. Di, M\. Zhang, and T\. WeiDC\-Merge: improving model merging with directional consistency\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 22248–22258\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2026/html/Zhang_DC-Merge_Improving_Model_Merging_with_Directional_Consistency_CVPR_2026_paper.html)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.18.1)\.
- Zhanget al\.\(2026c\)Y\. Zhang, J\. Chai, Y\. Fu, S\. Tu, X\. Wang, W\. Lin, G\. Yin, Q\. Zhang, Y\. Zhu, and D\. ZhaoAre full rollouts necessary for on\-policy distillation?\.External Links:2605\.31490,[Document](https://dx.doi.org/10.48550/arXiv.2605.31490),[Link](https://arxiv.org/abs/2605.31490)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.18.1.1.1)\.
- Zhaoet al\.\(2026\)S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. GroverSelf\-Distilled Reasoner: On\-Policy Self\-Distillation for Large Language Models\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.18734),[Link](https://arxiv.org/abs/2601.18734),2601\.18734Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.5.1.1.1),[§3\.3](https://arxiv.org/html/2609.19553#S3.SS3.p5.1)\.
- Zhaoet al\.\(2024\)X\. Zhao, G\. Sun, R\. Cai, Y\. Zhou, P\. Li, P\. Wang, B\. Tan, Y\. He, L\. Chen, Y\. Liang, B\. Chen, B\. Yuan, H\. Wang, A\. Li, Z\. Wang, and T\. ChenModel\-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild\.InAdvances in Neural Information Processing Systems 37, NeurIPS 2024,pp\. 13349–13371\.External Links:[Document](https://dx.doi.org/10.52202/079017-0426),[Link](https://doi.org/10.52202/079017-0426)Cited by:[Table 10](https://arxiv.org/html/2609.19553#A6.T10.2.3.1.1.1)\.
- Zhenget al\.\(2026\)B\. Zheng, X\. Ma, Y\. Liang, J\. Ruan, X\. Fu, K\. Lin, B\. Zhu, K\. Zeng, and X\. CaiSCOPE: signal\-calibrated on\-policy distillation enhancement with dual\-path adaptive weighting\.External Links:2604\.10688,[Document](https://dx.doi.org/10.48550/arXiv.2604.10688),[Link](https://arxiv.org/abs/2604.10688)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.14.1.1.1)\.
- Zhenget al\.\(2025\)H\. Zheng, L\. Shen, A\. Tang, Y\. Luo, H\. Hu, B\. Du, Y\. Wen, and D\. TaoLearning from Models Beyond Fine\-Tuning\.Nature Machine Intelligence7\(1\),pp\. 6–17\.External Links:ISSN 2522\-5839,[Document](https://dx.doi.org/10.1038/s42256-024-00961-0),[Link](https://doi.org/10.1038/s42256-024-00961-0)Cited by:[§1](https://arxiv.org/html/2609.19553#S1.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track,Vol\.36,pp\. 46595–46623\.External Links:[Document](https://dx.doi.org/10.52202/075280-2020),[Link](https://doi.org/10.52202/075280-2020)Cited by:[§2](https://arxiv.org/html/2609.19553#S2.SS0.SSS0.Px1.p3.1)\.
- Zhonget al\.\(2026\)Q\. Zhong, M\. Zheng, M\. Song, X\. Lin, J\. Sun, H\. Jiang, X\. Wang, and J\. FangSOD: step\-wise on\-policy distillation for small language model agents\.External Links:2605\.07725v1,[Document](https://dx.doi.org/10.48550/arXiv.2605.07725),[Link](https://arxiv.org/abs/2605.07725v1)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.16.1.1.1)\.
- Zhouet al\.\(2026a\)L\. Zhou, B\. Zhao, R\. Yu, and E\. RodolàDemystifying Mergeability: Interpretable Properties to Predict Model Merging Success\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.22285),[Link](https://arxiv.org/abs/2601.22285),2601\.22285Cited by:[§5](https://arxiv.org/html/2609.19553#S5.p1.1)\.
- Zhouet al\.\(2026b\)W\. Zhou, B\. Wang, H\. Zhang, C\. Jia, W\. Chen, and X\. ChengExtra\-Merge: tracing the rank\-1 subspace of model merging in language model pre\-training\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://icml.cc/virtual/2026/poster/64576)Cited by:[Table 5](https://arxiv.org/html/2609.19553#A3.T5.2.20.1)\.
- Zhouet al\.\(2026c\)Y\. Zhou, L\. Zhang, Y\. Wu, M\. Wang, B\. Peng, J\. Liu, X\. Fan, and Z\. ZhaoOmniOPD: logit\-free on\-policy distillation via speculative verification\.External Links:2606\.01476,[Document](https://dx.doi.org/10.48550/arXiv.2606.01476),[Link](https://arxiv.org/abs/2606.01476)Cited by:[Table 9](https://arxiv.org/html/2609.19553#A5.T9.2.23.1.1.1)\.
## Appendix ASymbol Definitions
Table 4:Notation used in the model fusion formulation\.
## Appendix BApplications
We cover continual learning, capability integration, safety, and compression\.
#### Model Fusion in Continual Learning\.
In continual learning, model fusion can add new task or domain knowledge while limiting forgetting\. AIMMerging\([Feng et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib6)\)and NUFILT\([Qiu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib75)\)apply parameter\-level fusion to add new task updates while reducing forgetting and interference\. K\-Merge\([Shenaj et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib38)\)extends this setting to online LoRA fusion for on\-device LLMs\. RECALL\([Wang et al\., 2025a](https://arxiv.org/html/2609.19553#bib.bib81)\)uses hidden representations for hierarchical fusion without historical data\. SDFT\([Shenfeld et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib89)\)uses behavior\-level fusion to learn new skills while reducing forgetting\.
#### Multi\-Task Learning and Domain Capability Integration\.
Model fusion integrates source models trained for different tasks, domains, or languages\. Compared with training one model on mixed task data, it can reuse existing source models and reduce reliance on original data or full retraining\([Jin et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib10);[Yang et al\., 2024c](https://arxiv.org/html/2609.19553#bib.bib3)\)\. Language Specific Model Merging\([Dmonte et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib33)\)fuses language\-specific models to lower multilingual training and update costs\. SurgeryV2\([Yang et al\., 2024b](https://arxiv.org/html/2609.19553#bib.bib91)\)and FeatCal\([Gu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib101)\)repair representation drift after fusion\. FuseLLM\([Wan et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib45)\)and DeepSeek\-V4\([DeepSeek\-AI et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib13)\)use behavior\-level fusion to integrate source capabilities\.
#### Safety and Control\.
For safety control, model fusion can transfer, keep, or weaken behavior attributes after training\. SafeMERGE\([Djuhera et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib86)\)and Fuse to Forget\([Zaman et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib24)\)use parameter\-level fusion to preserve safety or reduce unwanted behavior\. Safety Realignment\([Yi et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib87)\)uses subspace\-oriented model fusion to realign unsafe models\. Multilingual Safety Alignment via Self\-Distillation\([Qin et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib109)\)transfers safety behavior across languages through behavior\-level fusion\. However, unsafe source models can also propagate misalignment during fusion\([Hammoud et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib64)\)\.
#### Model Compression\.
Model compression uses source models to build smaller target models with similar capabilities\. LoRA soups\([Prabhakar et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib53)\)and LoRM\([Salami et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib36)\)can fold several lightweight modules into one target module\. DeepSeek\-R1\([Guo et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib12)\)transfers reasoning patterns into six dense models with 1\.5B to 70B parameters\. Nemotron\-Cascade 2\([Yang et al\., 2026f](https://arxiv.org/html/2609.19553#bib.bib73)\)builds a compact 30B MoE model, with 3B active parameters, for math, code, and agentic tasks\. Smaller target models can lower serving cost and speed up inference in resource\-limited settings\.
## Appendix CParameter\-Level Fusion Analysis
This appendix compares representative parameter\-level fusion methods by source\-model relation, fusion object, and evaluated backbones or settings\.
Table 5:Parameter\-level fusion methods compared by source\-model relation, fusion object, and evaluated backbones or settings\. Method groups follow the taxonomy in Section 3 and Figure[3](https://arxiv.org/html/2609.19553#S3.F3)\.
Table 6:Parameter\-level fusion methods compared by source\-model relation, fusion object, and evaluated backbones or settings \(continued\)\.
## Appendix DRepresentation\-Level Fusion Analysis
This appendix compares representation methods by source relation, fused scope, and evaluation setting\.
MethodVenueSource relationFused params/modulesEvaluated backbones/settingsWeighting and representation matching†REPAIR\([Jordan et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib55)\)ICLR’23Asame architecture after permutation alignmentNECNN classifiersZipIt\!\([Stoica et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib108)\)ICLR’24Aarchitecture\-compatible sourcesLAEvision Transformers/classifiersTransformer Fusion\([Imfeld et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib102)\)ICLR’24Aaligned Transformer variantsLAEViT and BERT encodersAIM\([Heyrani Nobari et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib2)\)NeurIPS’25Ssame\-base LLM checkpointsFDdecoder\-only LLMsACM\([Yao et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib1)\)NeurIPS’25Ssame\-base LLM checkpointsFDdecoder\-only LLMsMAGIC\([Li et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib54)\)arXiv’25Ssame\-base or aligned sourcesFNEDCV and Llama mergingMerging Beyond\([Yao et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib61)\)arXiv’26Ssequential same\-backbone updatesFLDstreaming LLM updatesClosed\-form representation solversRegMean\([Jin et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib10)\)ICLR’23Ssame architecture and tokenizerLETRoBERTa/DeBERTa and T5RegMean\+\+\([Nguyen et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib83)\)arXiv’25Asame\-family compatible modelsLETDencoder, enc–dec, decoder\-onlyLoRM\([Salami et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib36)\)ICLR’25SPEFT modules over compatible basesLPETViT\-B/16 and T5\-smallIterIS\([Chen et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib35)\)CVPR’25Scompatible LoRA adaptersPMtext\-to\-image, VLM, and LLM adaptersLOT\-Merging\([Sun et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib100)\)NeurIPS’25Ssame\-base task\-vector checkpointsLNEViT and RoBERTa checkpointsFeatCal\([Gu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib101)\)P\+RarXiv’26Ssame\-base task experts and merged modelsLNETDCLIP, FLAN\-T5, and LlamaBackpropagation\-based representation transferPatient KD\([Sun et al\., 2019](https://arxiv.org/html/2609.19553#bib.bib79)\)B\+REMNLP’19Tteacher–student fusionFEBERT\-style encodersTinyBERT\([Jiao et al\., 2020](https://arxiv.org/html/2609.19553#bib.bib99)\)R\+BEMNLP Findings’20Tteacher–student fusionFEBERT\-style encodersTED\([Liang et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib49)\)B\+RICML’23Tteacher–student fusionFPElanguage\-model compressionRepresentation Surgery\([Yang et al\., 2024a](https://arxiv.org/html/2609.19553#bib.bib84)\)P\+RICML’24Ssame\-base multi\-task merged modelsPEencoder\-based multi\-task modelsSurgeryv2\([Yang et al\., 2024b](https://arxiv.org/html/2609.19553#bib.bib91)\)P\+RarXiv’24Ssame\-base multi\-task merged modelsPEViT and BERT multi\-task modelsProbSurgery\([Wei et al\., 2025a](https://arxiv.org/html/2609.19553#bib.bib85)\)P\+RICML’25Ssame\-base multi\-task merged modelsPEmulti\-task model mergingRECALL\([Wang et al\., 2025a](https://arxiv.org/html/2609.19553#bib.bib81)\)EMNLP’25Scontinual same\-family checkpointsLNCin\-domain checkpoint sequencesNUFILT\([Qiu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib75)\)ICLR’26Scontinual same\-backbone checkpointsFPCdata\-free continual mergingOPRD\([Yang et al\., 2026c](https://arxiv.org/html/2609.19553#bib.bib138)\)arXiv’26THcross\-family teacherFPDheterogeneous reasoning LLMs
Source relationSsame base, tokenizer, or checkpoint trajectory;Aarchitecture\-compatible sources requiring alignment/matching;Hheterogeneous or projection\-needed sources;Tteacher–student representation fusion\.Fused scopeFfull parameters or task vectors;Llinear/projection weights;Aattention components;Nnormalization, bias, or activation statistics;Pprojection, LoRA, adapter, or repair module\.Backbones/settingsEencoder or vision backbone;Ddecoder\-only LLM;Tencoder–decoder;Mmultimodal or vision–language setting;Ccheckpoint sequence or continual\-merging setting\.Cross\-level labelsP\+Rparameter fusion with representation calibration or repair;B\+RandR\+Boutput distillation combined with representation matching\.Boundary cases†Non\-LLM representation repair included because it motivates representation\-level post\-merge correction; teacher–student methods are treated as representation\-level fusion when intermediate hidden states or attention maps provide the main fusion signal\.
Table 7:Representation\-level fusion methods compared by source\-model relation, fused parameter/module scope, and evaluated model backbones or settings\. Method groups follow the taxonomy in Section 3\.
## Appendix EBehavior\-Level Fusion Analysis
This appendix compares behavior methods by source relation, signal, and evaluation setting\.
MethodVenueSource relationBehavior signalEvaluated backbones/settingsDistribution fusionKnowledge Distillation\([Hinton et al\., 2015](https://arxiv.org/html/2609.19553#bib.bib18)\)NeurIPS’15Tteacher–studentDOsoft outputsgeneral neural networksDistilBERT\([Sanh et al\., 2019](https://arxiv.org/html/2609.19553#bib.bib17)\)NeurIPS’19TBERT teacher–studentDOtoken distributionsEBERT encodersPatient KD\([Sun et al\., 2019](https://arxiv.org/html/2609.19553#bib.bib79)\)B\+REMNLP’19TBERT teacher–studentDOoutputs \+ hidden statesEBERT encodersTinyBERT\([Jiao et al\., 2020](https://arxiv.org/html/2609.19553#bib.bib99)\)R\+BEMNLP Findings’20TBERT teacher–studentDOlogits \+ representationsEBERT encodersTED\([Liang et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib49)\)B\+RICML’23TLM teacher–studentDOoutputs \+ filtered reps\.Eencoder LMsInfiGFusion\([Wang et al\., 2025b](https://arxiv.org/html/2609.19553#bib.bib110)\)NeurIPS’25Mmulti\-source LLMsDOlogit geometryDdecoder\-only LLMsDemonstration fusionInstruction Distillation\([Sun et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib34)\)arXiv’23Tstronger teacherXOinstruction responsesDLLM rankersGPT4All\([Anand et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib29)\)GitHub’23BAPI teacherXOassistant demosCchatbot tuningDistilling Step\-by\-Step\([Hsieh et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib19)\)ACL’23Tlarger LLM teacherXOrationales and labelsDsmall reasoning LMsTeaching Small LMs to Reason\([Magister et al\., 2023](https://arxiv.org/html/2609.19553#bib.bib97)\)ACL’23Treasoning teacherXOCoT rationalesDsmall LMsFeedback fusionFuseChat\([Wan et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib25)\)EMNLP’25Mchat\-model sourcesXFOresponses and preferencesCchat fusionZephyr\([Tunstall et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib107)\)COLM’24Taligned teacherFOpreference dataCchat alignmentInfiFPO\([Gu et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib111)\)NeurIPS’25Msource preferencesFOpreference optimizationDLLM fusionTable 8:Behavior\-level fusion methods compared by source\-model relation, behavior signal, state distribution, and evaluated backbones or settings\. Method groups follow the taxonomy in Section 3\.3\.Source relationTteacher–student or expert–learner relation;Mmultiple source models or model zoo;Bblack\-box or API\-access source;Sself\-distillation source;Hheterogeneous or cross\-modal source–target setting\.Behavior signalDoutput distributions, token probabilities, or logits;Xdemonstrations, responses, rationales, or traces;Fpreferences, rankings, scores, critiques, rewards, or verifier labels;Rtarget\-induced states, rollouts, or trajectories\.State distributionOoff\-policy fixed behavior data;Pon\-policy supervision on states induced by the target model;Ssemi\-on\-policy, cached, or offline\-reused on\-policy\-style supervision\.Backbones/settingsEencoder\-only LM;Ddecoder\-only LLM;Cchat or instruction\-following LLM;Mmultimodal, speech, or video\-language setting;Aimitation\-learning or agent\-policy setting\.Cross\-level labelsB\+RandR\+Bdenote output distillation combined with representation matching; these methods are cross\-referenced in Table[7](https://arxiv.org/html/2609.19553#A4.T7)\.Boundary cases†DAgger is included as the classical on\-policy imitation\-learning analogue of behavior\-level fusion; methods are grouped according to the behavior\-level branch in Section 3\.3\. Labels are descriptive and non\-exclusive, since a method may combine distributions, demonstrations, and feedback\.
Table 9:Behavior\-level fusion methods compared by source\-model relation, behavior signal, state distribution, and evaluated backbones or settings \(continued\)\.
## Appendix FModel Fusion Benchmarks
We compare fusion benchmarks by modality, task, heterogeneity, type, and evaluation focus\.
Pool definitionModel Poolindicates whether the resource explicitly defines source models, expert models, or fine\-tuned checkpoints to be fused;Task Poolindicates whether it defines downstream tasks or domain pools for post\-merge evaluation\.HeterogeneityHetero\.indicates whether the resource explicitly evaluates heterogeneous fusion, including cross\-family, cross\-architecture, cross\-modal, or heterogeneous\-output\-space settings\.OpenOpenindicates whether the resource provides public code, scripts, model pools, evaluation resources, or reproducible configurations\.SymbolsYexplicit support;Ppartial, implicit, or recipe\-dependent support;Nnot covered or not the focus\.Boundary casesToolkits and search ecosystems, such as MergeKit\([Goddard et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib8)\)and Mergenetic\([Minut et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib58)\), are included when they provide reusable fusion pipelines or practical evaluation settings\. They are not treated as fixed benchmark suites\.
Table 10:Representative model fusion benchmarks and related evaluation resources\. Unlike traditional LLM benchmarks that mainly define task instances and metrics, model fusion benchmarks often define source model pools, target tasks, fusion settings, and cost or retention axes\.
## Appendix GModel Fusion Metrics
Table[11](https://arxiv.org/html/2609.19553#A7.T11)compares model fusion metrics and their use conditions\.
Table 11:Common metrics for model fusion\. Avg score and normalized performance are the most direct metrics\. Other metrics are useful but depend on the setting, access level, or safety goal\.
## Appendix HAdditional Deployment Challenges
We discuss continual forgetting and scaling beyond Section[5](https://arxiv.org/html/2609.19553#S5)\.
#### Continual Fusion Can Easily Cause Forgetting\.
In real deployment, the target model may continually absorb new source models, domain updates, or safety patches\. Each fusion step can overwrite earlier knowledge or weaken previously aligned behavior, especially when old training data, source models, or evaluation signals are unavailable\. AIMMerging\([Feng et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib6)\), NUFILT\([Qiu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib75)\), and K\-Merge\([Shenaj et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib38)\)study continual fusion for language models, but stable long\-term fusion remains open\. Behavior\-level methods also face forgetting when new skills are learned from source feedback\([Shenfeld et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib89)\)\. Future methods should add skills without retraining while preserving capabilities and safety\.
#### Large\-Scale Fusion Remains Underexplored\.
Model fusion can be cheaper than retraining or using all source models at inference time, but large\-scale fusion brings new costs\. For parameter\-level fusion, MergeKit\([Goddard et al\., 2024](https://arxiv.org/html/2609.19553#bib.bib8)\)makes LLM fusion easier to run\. MergePipe\([Wang et al\., 2026d](https://arxiv.org/html/2609.19553#bib.bib59)\)further shows that expert\-parameter I/O and repeated scans become key bottlenecks as the source pool grows\. For method search, FusionBench\([Tang et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib26)\)and MergeBench\([He et al\., 2025](https://arxiv.org/html/2609.19553#bib.bib57)\)improve standard comparison, but large models still make candidate evaluation costly\. For behavior\-level fusion, source feedback can also be expensive\. Lightning OPD\([Wu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib113)\)and TIP\([Xu et al\., 2026](https://arxiv.org/html/2609.19553#bib.bib114)\)reduce live teacher serving or token\-level supervision cost\.Similar Articles
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
This survey paper systematically reviews the paradigm evolution of unified vision-language perception in multimodal large language models (MLLMs), proposing a five-stage taxonomy and identifying open challenges toward general multimodal intelligence.
Memory for Large Language Models
This survey presents a systematic taxonomy of memory mechanisms in large language models, classifying along axes of representation, update dynamics, and persistence, and formalizing the underlying mechanistic components.
Model Merging Scaling Laws in Large Language Models
This paper establishes empirical scaling laws for language model merging, identifying power-law relationships between model size, expert count, and performance to enable predictive planning for optimal model composition.
Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation
This paper introduces a dimension-level evaluation method for measuring intent fidelity in large language models using structured prompt ablation.
A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models
A comprehensive survey on fake review detection, analyzing methods from traditional ML to LLM-based approaches, with a focus on information fusion and benchmark trends.