CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging
Summary
CT-Merging proposes a method to merge LoRA adapters by estimating consensus directions from task subspace projectors and assigning task-level RMS coefficient scales, achieving superior performance on the DC-Merge CLIP adapter benchmark.
View Cached Full Text
Cached at: 07/24/26, 05:12 AM
# CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging
Source: [https://arxiv.org/html/2607.20561](https://arxiv.org/html/2607.20561)
Keumseo Ryum Joonhyuk Kang KAIST Daejeon, Republic of Korea keumseo@kaist\.ac\.kr jkang@kaist\.ac\.kr
###### Abstract
LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter storage and task selection at inference time\. Model merging addresses this issue by combining independently trained adapters into one multi\-task adapter\. Recent SVD\-based LoRA merging methods mainly focus on constructing shared or task specific directions, while the coefficients assigned to the final directions are often directly from the original task SVD\. On a fixed merged basis, inherited coefficients preserve component order with high rank correlation, yet their magnitudes differ substantially from the coefficients induced by the task updates\. To address this mismatch, we propose CT\-Merging, a LoRA\-aware merging algorithm that estimates consensus directions from average task subspace projectors and assigns task\-level RMS coefficient scales in the final update\. CT\-Merging uses repeated support across task SVD subspaces to construct the common basis, while reducing reliance on rank wise SVD magnitudes after direction construction\. On the DC\-Merge CLIP adapter benchmark, CT\-Merging achieves superior average normalized accuracy compared to state\-of\-the\-art merging methods and further improves over DC\-Merge by 2\.56 points on ViT\-B/32 and 1\.51 points on ViT\-L/14 KnoTS\-trained checkpoints\.
## 1Introduction
Large pretrained vision and vision language models are increasingly specialized to downstream tasks through parameter efficient fine tuning, since full fine tuning is costly in memory and storage\[[9](https://arxiv.org/html/2607.20561#bib.bib20)\]\. Low rank adaptation \(LoRA\)\[[10](https://arxiv.org/html/2607.20561#bib.bib8)\]has become a standard method to store such specialization\. LoRA represents the weight update of a pretrained model as a product of two low rank matrices, allowing a single pretrained model to be paired with many lightweight task adapters\. As these adapters accumulate across datasets, domains, and user needs, a practical problem arises\. The adapters should be combined into one deployable adapter that serves multiple tasks without joint training, without revisiting the original training data, and without keeping a separate adapter for every task\.
Model merging addresses this problem by composing independently trained models or task updates into a single set of weights\[[30](https://arxiv.org/html/2607.20561#bib.bib11),[12](https://arxiv.org/html/2607.20561#bib.bib3),[34](https://arxiv.org/html/2607.20561#bib.bib24)\]\. Merging is beneficial for LoRA adapters because it can combine task specializations after training, using only the released adapter weights\[[11](https://arxiv.org/html/2607.20561#bib.bib21),[25](https://arxiv.org/html/2607.20561#bib.bib28)\]\. Early methods operate directly in parameter space by directly averaging fine tuned checkpoints\[[30](https://arxiv.org/html/2607.20561#bib.bib11)\], or combining the task vectors obtained by subtracting the pretrained weights\[[12](https://arxiv.org/html/2607.20561#bib.bib3),[23](https://arxiv.org/html/2607.20561#bib.bib26)\]\. Since direct summation can cause interference, later methods reduce coordinate level conflicts through sign resolution, trimming, random dropping, masking, localization, or learned merging weights\[[33](https://arxiv.org/html/2607.20561#bib.bib9),[36](https://arxiv.org/html/2607.20561#bib.bib10),[29](https://arxiv.org/html/2607.20561#bib.bib27),[35](https://arxiv.org/html/2607.20561#bib.bib23)\]\. These approaches are broadly applicable, but they treat each parameter coordinate as the basic merging unit and therefore do not directly use the low rank structure exposed by LoRA updates\.
A more recent line of work specializes merging to LoRA modules by decomposing each task update into singular vectors and then recomposing a merged update from these components\. Recent methods construct common or task specific subspaces, align singular directions across tasks, filter conflicting components, or reshape the merged spectrum\[[28](https://arxiv.org/html/2607.20561#bib.bib5),[5](https://arxiv.org/html/2607.20561#bib.bib16),[19](https://arxiv.org/html/2607.20561#bib.bib1),[37](https://arxiv.org/html/2607.20561#bib.bib12),[17](https://arxiv.org/html/2607.20561#bib.bib13)\]\. However, recomposition still depends on how the final directions are paired with coefficients\. After common and residual directions are projected and aligned, a common choice is to copy each singular value from the original task SVD component by component\. The copied magnitudes were measured before projection and alignment, so they need not remain calibrated for the recomposed directions\. A merged adapter must also avoid severe task collapse\. In multi task deployment, a high average score can hide a failed task, especially when one adapter is expected to serve all tasks without per input routing\.
Building on this analysis, we propose CT\-Merging, a data free method for merging LoRA adapters through consensus directions and task\-level coefficient assignment\. CT\-Merging first estimates a common basis from average projectors over task SVD subspaces\. The projector form avoids the sign ambiguity of individual singular vectors and selects directions that receive repeated support across task adapters\. CT\-Merging then constructs task residual directions by removing the common component from each task SVD direction\.
The coefficient rule uses the task SVD coefficients through their residual energy\. All residual directions from the same task receive one RMS scale, so the total coefficient energy assigned to that task is preserved while the rank\-wise singular value magnitudes are discarded\. This design targets the coefficient mismatch observed after projection and alignment without erasing task scale differences\. On the DC\-Merge benchmark, CT\-Merging gives the best result on most of the models, and on KnOTS checkpoints, CT\-Merging gives larger gains over DC\-Merge on both ViT\-B/32 and ViT\-L/14\. Our contributions are as follows\.
- •Coefficient transfer analysis in LoRA mergingWe analyze what happens when singular values measured in isolated task SVD bases are reused after projection and basis construction\. The analysis shows that rank order can remain stable while coefficient magnitudes change substantially in the final basis\.
- •Consensus directions with task\-level scalingWe propose CT\-Merging, which builds consensus directions from average projectors, removes the common component from task residual directions, and assigns coefficients from task\-level RMS energy\. The rule keeps a separate residual energy budget for each task without directly copying rank wise singular value magnitudes\.
- •Empirical validation on released CLIP LoRA benchmarksWe evaluate CT\-Merging on DC\-Merge and KnOTS style CLIP adapter checkpoints\. CT\-Merging achieves the best result on eight of nine DC\-Merge benchmark settings and gives larger gains on KnOTS style checkpoints\. We provide ablations on the contribution of the hyperparameters and the consensus source\.
Figure 1:Overview of CT\-Merging\. CT\-Merging estimates consensus directions from task SVD subspaces, projects task residual directions away from the consensus subspace, and recomposes the merged adapter with task\-level RMS coefficients\. The coefficient rule preserves task energy while removing unreliable component\-wise magnitudes after alignment\.
## 2Related Work
##### Model merging
Model merging combines independently fine tuned models or task updates into a single model without joint training\. Early methods operate directly in weight space\. Model soups average fine tuned weights\[[30](https://arxiv.org/html/2607.20561#bib.bib11)\], and task arithmetic composes task vectors obtained by subtracting the pretrained weights\[[12](https://arxiv.org/html/2607.20561#bib.bib3)\]\. Since direct summation can introduce interference, later methods reduce coordinate level conflicts through sign resolution, pruning, random dropping, masking, or learned merging weights\[[33](https://arxiv.org/html/2607.20561#bib.bib9),[36](https://arxiv.org/html/2607.20561#bib.bib10),[29](https://arxiv.org/html/2607.20561#bib.bib27),[35](https://arxiv.org/html/2607.20561#bib.bib23)\]\. Data dependent methods such as Fisher merging and RegMean use curvature or feature statistics to guide the merge\[[20](https://arxiv.org/html/2607.20561#bib.bib6),[13](https://arxiv.org/html/2607.20561#bib.bib18)\]\. CT\-Merging follows the data free setting and focuses on the low rank structure exposed by LoRA task updates\.
##### LoRA\-aware merging
SVD based merging methods decompose task updates by SVD, construct shared or task specific directions, and recompose a merged update from singular directions and coefficients\. Task Singular Vectors orthogonalizes task specific singular directions to reduce interference\[[5](https://arxiv.org/html/2607.20561#bib.bib16)\]\. KnOTS performs a joint SVD of stacked task updates and merges in the rotated basis\[[28](https://arxiv.org/html/2607.20561#bib.bib5)\]\. Iso\-CTS combines common and task specific subspaces and applies isotropic scaling to the resulting spectrum\[[19](https://arxiv.org/html/2607.20561#bib.bib1)\]\. DC\-Merge smooths leading singular values and aligns task updates in a shared orthogonal subspace\[[37](https://arxiv.org/html/2607.20561#bib.bib12)\]\. AdaRank adapts the effective rank of each task update\[[17](https://arxiv.org/html/2607.20561#bib.bib13)\]\. SVC studies spectral over counting in an already merged update and calibrates inflated singular values while keeping the merged singular directions fixed\[[18](https://arxiv.org/html/2607.20561#bib.bib15)\]\. These works show that singular value assignment matters in merging\. CT\-Merging focuses on coefficient assignment after SVD based direction construction\.
## 3Observation on Recomposition Coefficients
SVD based LoRA merging methods decompose each task update into singular components and then recompose a merged update in a new basis\. In this process, the singular values from the isolated task SVD are often reused as coefficients for the recomposed directions\. These values are measured before projection and recomposition, so their magnitudes need not remain calibrated in the final basis\. We examine this transfer on a fixed recomposition basis, separating the effect of coefficient assignment from the choice of directions\.
For each taskttand LoRA moduleℓ\\ell, letCt,ℓC\_\{t,\\ell\}be the coefficient matrix obtained by expressing the task updateΔt,ℓ\\Delta\_\{t,\\ell\}in the recomposition basis,
Ct,ℓ=U~ℓ⊤Δt,ℓV~ℓ,dt,ℓ=diag\(Ct,ℓ\)\.C\_\{t,\\ell\}=\\widetilde\{U\}\_\{\\ell\}^\{\\top\}\\Delta\_\{t,\\ell\}\\widetilde\{V\}\_\{\\ell\},\\qquad d\_\{t,\\ell\}=\\operatorname\{diag\}\(C\_\{t,\\ell\}\)\.Here,dt,ℓ,jd\_\{t,\\ell,j\}is the coefficient induced by the task update on thejjth final direction, whilest,ℓ,js\_\{t,\\ell,j\}is the coefficient inherited from the isolated task SVD\. Across layer task pairs,st,ℓs\_\{t,\\ell\}anddt,ℓd\_\{t,\\ell\}remain strongly rank correlated\. The median Spearman correlation is 0\.9956 on CLIP ViT B/32 and 0\.9912 on CLIP ViT L, with no sign flips in the diagonal entries\. The component ordering therefore remains largely stable when the task update is expressed in the recomposition basis\.
The coefficient magnitudes, however, change substantially\. At the layer task level, the median vector relative distance∥dt,ℓ−st,ℓ∥2/∥st,ℓ∥2\\lVert d\_\{t,\\ell\}\-s\_\{t,\\ell\}\\rVert\_\{2\}/\\lVert s\_\{t,\\ell\}\\rVert\_\{2\}is 0\.35 on ViT B/32 and 0\.39 on ViT L\. At the component level, the median relative coefficient error is about 0\.20 on both backbones, where
et,ℓ,j=\|dt,ℓ,j−st,ℓ,j\|\|st,ℓ,j\|\+ϵ\.e\_\{t,\\ell,j\}=\\frac\{\|d\_\{t,\\ell,j\}\-s\_\{t,\\ell,j\}\|\}\{\|s\_\{t,\\ell,j\}\|\+\\epsilon\}\.
Figure 2:Rank wise coefficient error on CLIP LoRA adapters\. For each residual rankjj, curves show the median ofet,ℓ,je\_\{t,\\ell,j\}over the eight tasks and all LoRA modules, and shaded regions show the interquartile range\. The leading residual ranks show the largest coefficient error on both backbones\.Figure 3:Task SVD energy differs strongly across CLIP ViT B/32 LoRA adapters\. Values are normalized by the mean task energy\.Figure[2](https://arxiv.org/html/2607.20561#S3.F2)plotsmediant,ℓet,ℓ,j\\operatorname\{median\}\_\{t,\\ell\}e\_\{t,\\ell,j\}for each residual rankjj, with interquartile ranges over tasks and modules\. The leading residual ranks have the largest error, and the mismatch remains visible across the retained ranks\. These measurements show that inherited coefficients preserve a useful ordering signal, but their magnitudes are not directly calibrated for the recomposition basis\.
The same diagnostic also shows why all tasks should not be collapsed to one shared residual scale\. For tasktt, we compute the raw SVD energy averaged across modules:
Etraw=1L∑ℓ=1L∑jst,ℓ,j2,E¯raw=1T∑tEtraw\.E^\{\\mathrm\{raw\}\}\_\{t\}=\\frac\{1\}\{L\}\\sum\_\{\\ell=1\}^\{L\}\\sum\_\{j\}s\_\{t,\\ell,j\}^\{2\},\\qquad\\bar\{E\}^\{\\mathrm\{raw\}\}=\\frac\{1\}\{T\}\\sum\_\{t\}E^\{\\mathrm\{raw\}\}\_\{t\}\.Figure[3](https://arxiv.org/html/2607.20561#S3.F3)showsEtraw/E¯rawE^\{\\mathrm\{raw\}\}\_\{t\}/\\bar\{E\}^\{\\mathrm\{raw\}\}on CLIP ViT B/32\. The relative task energy ranges from 0\.10 to 2\.47 across tasks\. Assigning the same RMS magnitude to all residual directions would erase these differences across tasks\. CT\-Merging therefore uses a per task RMS rule as a simple default\. The rule avoids directly copying rank wise magnitudes from the isolated SVD basis while preserving a separate residual energy budget for each task\.
## 4Method
CT\-Merging separates direction construction from coefficient assignment in SVD\-based LoRA merging\. The direction step builds a common subspace and task residual directions from independently trained adapters\. The coefficient step then assigns task\-level RMS scales to the constructed directions\. This separation follows the observation in Section[3](https://arxiv.org/html/2607.20561#S3)that singular values measured in an isolated task SVD basis do not necessarily provide calibrated magnitudes after recomposition\.
Algorithm 1CT\-Merging1:LoRA updates
\{Δt=BtAt\}t=1T\\\{\\Delta\_\{t\}=B\_\{t\}A\_\{t\}\\\}\_\{t=1\}^\{T\}, common rank
kk, residual rank
rresr\_\{\\mathrm\{res\}\}, merge scale
γ\\gamma
2:
⊳\\trianglerightTask SVD
3:for
t=1,…,Tt=1,\\ldots,Tdo
4:Compute
Δt≈UtStVt⊤\\Delta\_\{t\}\\approx U\_\{t\}S\_\{t\}V\_\{t\}^\{\\top\}
5:Keep the top
rresr\_\{\\mathrm\{res\}\}directions and coefficients
\(Ut,Vt,st\)\(U\_\{t\},V\_\{t\},s\_\{t\}\)
6:endfor
7:
⊳\\trianglerightConsensus directions
8:
PU←T−1∑t=1TUtUt⊤P\_\{U\}\\leftarrow T^\{\-1\}\\sum\_\{t=1\}^\{T\}U\_\{t\}U\_\{t\}^\{\\top\}
9:
Uc←TopEigk\(PU\)U\_\{c\}\\leftarrow\\operatorname\{TopEig\}\_\{k\}\(P\_\{U\}\)
10:
RU←T−1∑t=1TUc⊤ΔtR\_\{U\}\\leftarrow T^\{\-1\}\\sum\_\{t=1\}^\{T\}U\_\{c\}^\{\\top\}\\Delta\_\{t\}
11:Compute
RU=PΣcVm⊤R\_\{U\}=P\\Sigma\_\{c\}V\_\{m\}^\{\\top\}
12:
Ucom←UcPU\_\{\\mathrm\{com\}\}\\leftarrow U\_\{c\}P,
Vcom←VmV\_\{\\mathrm\{com\}\}\\leftarrow V\_\{m\},
sc←diag\(Σc\)s\_\{c\}\\leftarrow\\operatorname\{diag\}\(\\Sigma\_\{c\}\)
13:
⊳\\trianglerightResidual directions
14:for
t=1,…,Tt=1,\\ldots,Tdo
15:
Ut⟂←\(I−UcUc⊤\)UtU\_\{t\}^\{\\perp\}\\leftarrow\(I\-U\_\{c\}U\_\{c\}^\{\\top\}\)U\_\{t\}
16:
Vtres←VtV\_\{t\}^\{\\mathrm\{res\}\}\\leftarrow V\_\{t\}
17:endfor
18:
Urec←\[Ucom∣U1⟂∣⋯∣UT⟂\]U\_\{\\mathrm\{rec\}\}\\leftarrow\[U\_\{\\mathrm\{com\}\}\\mid U\_\{1\}^\{\\perp\}\\mid\\cdots\\mid U\_\{T\}^\{\\perp\}\]
19:
Vrec←\[Vcom∣V1res∣⋯∣VTres\]V\_\{\\mathrm\{rec\}\}\\leftarrow\[V\_\{\\mathrm\{com\}\}\\mid V\_\{1\}^\{\\mathrm\{res\}\}\\mid\\cdots\\mid V\_\{T\}^\{\\mathrm\{res\}\}\]
20:
U~←Polar\(Urec\)\\widetilde\{U\}\\leftarrow\\operatorname\{Polar\}\(U\_\{\\mathrm\{rec\}\}\),
V~←Polar\(Vrec\)\\widetilde\{V\}\\leftarrow\\operatorname\{Polar\}\(V\_\{\\mathrm\{rec\}\}\)
21:
⊳\\trianglerightTask\-level coefficients
22:
ρc←k−1∑isc,i2\\rho\_\{c\}\\leftarrow\\sqrt\{k^\{\-1\}\\sum\_\{i\}s\_\{c,i\}^\{2\}\}
23:for
t=1,…,Tt=1,\\ldots,Tdo
24:
ρt←rres−1∑ist,i2\\rho\_\{t\}\\leftarrow\\sqrt\{r\_\{\\mathrm\{res\}\}^\{\-1\}\\sum\_\{i\}s\_\{t,i\}^\{2\}\}
25:endfor
26:
snew←\[ρc𝟏k,ρ1𝟏rres,…,ρT𝟏rres\]s\_\{\\mathrm\{new\}\}\\leftarrow\[\\rho\_\{c\}\\mathbf\{1\}\_\{k\},\\rho\_\{1\}\\mathbf\{1\}\_\{r\_\{\\mathrm\{res\}\}\},\\ldots,\\rho\_\{T\}\\mathbf\{1\}\_\{r\_\{\\mathrm\{res\}\}\}\]
27:return
Δmerge=γU~diag\(snew\)V~⊤\\Delta\_\{\\mathrm\{merge\}\}=\\gamma\\widetilde\{U\}\\operatorname\{diag\}\(s\_\{\\mathrm\{new\}\}\)\\widetilde\{V\}^\{\\top\}
### 4\.1Setup
We first define the model merging setting\. LetW0W\_\{0\}denote the pretrained base model and letWtW\_\{t\}denote the model fine\-tuned on tasktt\. A task update is defined using the concept of task vectors as in\[[12](https://arxiv.org/html/2607.20561#bib.bib3)\]
Δt=Wt−W0,\\Delta\_\{t\}=W\_\{t\}\-W\_\{0\},\(1\)and a merged model is obtained by adding a composed update to the base model,
Wm=W0\+Δmerge\.W\_\{m\}=W\_\{0\}\+\\Delta\_\{\\mathrm\{merge\}\}\.\(2\)Many merging methods constructΔmerge\\Delta\_\{\\mathrm\{merge\}\}by arithmetic operations on the task updates, for example
Δmerge=∑t=1TαtΔt\.\\Delta\_\{\\mathrm\{merge\}\}=\\sum\_\{t=1\}^\{T\}\\alpha\_\{t\}\\Delta\_\{t\}\.\(3\)CT\-Merging follows the data free merging setting, where only the independently trained task adapters are available and no task data are used during merging\. We considerTTLoRA adapters trained independently from the same pretrained model\. For each LoRA module, tasktthas a low\-rank update
Δt=BtAt,Bt∈ℝdout×rL,At∈ℝrL×din\.\\Delta\_\{t\}=B\_\{t\}A\_\{t\},\\qquad B\_\{t\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times r\_\{\\mathrm\{L\}\}\},\\quad A\_\{t\}\\in\\mathbb\{R\}^\{r\_\{\\mathrm\{L\}\}\\times d\_\{\\mathrm\{in\}\}\}\.\(4\)
### 4\.2Consensus Basis from Average Projectors
CT\-Merging defines common directions as directions that are repeatedly supported by task SVD subspaces\. Direct averaging of singular vectors is sensitive to sign ambiguity and rotations inside each task SVD subspace\. CT\-Merging therefore estimates the common basis from an average projector over the left singular subspaces,
PU=1T∑t=1TUtUt⊤\.P\_\{U\}=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}U\_\{t\}U\_\{t\}^\{\\top\}\.\(5\)The consensus basis is given by the topkkeigenvectors,
Uc=TopEigk\(PU\)\.U\_\{c\}=\\operatorname\{TopEig\}\_\{k\}\(P\_\{U\}\)\.\(6\)For any unit vectorqq,q⊤PUqq^\{\\top\}P\_\{U\}qis the average squared projection ofqqonto the task left singular subspaces\. The projectorUtUt⊤U\_\{t\}U\_\{t\}^\{\\top\}represents the retained subspace rather than individual signed singular vectors, so it is invariant to sign changes and rotations of the retained basis\.
The consensus basisUcU\_\{c\}specifies only the left side of the update\. To form matrix directions for recomposition, CT\-Merging must determine the right directions paired with this basis\. We compute the average projected task response,
RU=1T∑t=1TUc⊤Δt,RU=PΣcVm⊤\.R\_\{U\}=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}U\_\{c\}^\{\\top\}\\Delta\_\{t\},\\qquad R\_\{U\}=P\\Sigma\_\{c\}V\_\{m\}^\{\\top\}\.The SVD ofRUR\_\{U\}rotates the consensus basis byPPand selects right directionsVmV\_\{m\}that explain the average task update within the consensus subspace\. The resulting common directions and their raw coefficients are
Ucom=UcP,Vcom=Vm,sc=diag\(Σc\)\.U\_\{\\mathrm\{com\}\}=U\_\{c\}P,\\qquad V\_\{\\mathrm\{com\}\}=V\_\{m\},\\qquad s\_\{c\}=\\operatorname\{diag\}\(\\Sigma\_\{c\}\)\.\(7\)The SVD ofRUR\_\{U\}pairs the left consensus basis with right directions that explain the average projected task update\.
Table 1:Average normalized accuracy on the DC\-Merge adapter benchmark\. Best value in each column is bold\.
### 4\.3Residual Direction Construction
After selecting the common basis, CT\-Merging constructs residual directions for each task by removing the common component from its retained left singular directions,
ut,j⟂=\(I−UcUc⊤\)ut,j,vt,jres=vt,j\.u\_\{t,j\}^\{\\perp\}=\(I\-U\_\{c\}U\_\{c\}^\{\\top\}\)u\_\{t,j\},\\qquad v\_\{t,j\}^\{\\mathrm\{res\}\}=v\_\{t,j\}\.\(8\)The projection removes the part ofut,ju\_\{t,j\}already explained by the common basis\. We keepvt,jv\_\{t,j\}unchanged, so that each projected left direction remains paired with the right singular direction from the same task SVD component\.
The common and residual directions are concatenated into recomposition direction matrices,
Urec=\[Ucom∣\{ut,j⟂\}t,j\],Vrec=\[Vcom∣\{vt,jres\}t,j\]\.U\_\{\\mathrm\{rec\}\}=\[U\_\{\\mathrm\{com\}\}\\mid\\\{u\_\{t,j\}^\{\\perp\}\\\}\_\{t,j\}\],\\qquad V\_\{\\mathrm\{rec\}\}=\[V\_\{\\mathrm\{com\}\}\\mid\\\{v\_\{t,j\}^\{\\mathrm\{res\}\}\\\}\_\{t,j\}\]\.\(9\)CT\-Merging applies polar projection to obtain orthonormal recomposition directions,
U~=Polar\(Urec\),V~=Polar\(Vrec\),\\widetilde\{U\}=\\operatorname\{Polar\}\(U\_\{\\mathrm\{rec\}\}\),\\qquad\\widetilde\{V\}=\\operatorname\{Polar\}\(V\_\{\\mathrm\{rec\}\}\),\(10\)wherePolar\(X\)=LR⊤\\operatorname\{Polar\}\(X\)=LR^\{\\top\}for the compact SVDX=LΣR⊤X=L\\Sigma R^\{\\top\}\. After polar projection, each entry of the final coefficient vector multiplies one left direction and one right direction in the recomposition basis\.
Table 2:Average normalized accuracy on the CLIP ViT\-B/32 adapters provided by KnOTS\. Best results are in bold\.Table 3:Average normalized accuracy on the CLIP ViT\-L/14 adapters provided by KnOTS\. Best results are in bold\.
### 4\.4Task\-Level Coefficients
The coefficientsst,js\_\{t,j\}are measured in the isolated task SVD basis, while recomposition uses\(U~,V~\)\(\\widetilde\{U\},\\widetilde\{V\}\)\. CT\-Merging uses the task SVD coefficients only through their squared sum,
Et=∑j=1rresst,j2\.E\_\{t\}=\\sum\_\{j=1\}^\{r\_\{\\mathrm\{res\}\}\}s\_\{t,j\}^\{2\}\.\(11\)All residual directions from taskttreceive the same RMS coefficient,
st,jnew=Etrres,j=1,…,rres\.s\_\{t,j\}^\{\\mathrm\{new\}\}=\\sqrt\{\\frac\{E\_\{t\}\}\{r\_\{\\mathrm\{res\}\}\}\},\\qquad j=1,\\ldots,r\_\{\\mathrm\{res\}\}\.\(12\)The total squared coefficient energy assigned to taskttremainsEtE\_\{t\}, while the rank\-wise magnitude profile of the isolated task SVD is removed\.
The common directions use the same RMS form\. With
Ec=∑j=1ksc,j2,E\_\{c\}=\\sum\_\{j=1\}^\{k\}s\_\{c,j\}^\{2\},\(13\)CT\-Merging sets
sc,jnew=Eck,j=1,…,k\.s\_\{c,j\}^\{\\mathrm\{new\}\}=\\sqrt\{\\frac\{E\_\{c\}\}\{k\}\},\\qquad j=1,\\ldots,k\.\(14\)The final coefficient vector is
snew=\[sc,1new,…,sc,knew,s1,1new,…,sT,rresnew\]\.s\_\{\\mathrm\{new\}\}=\\big\[s\_\{c,1\}^\{\\mathrm\{new\}\},\\ldots,s\_\{c,k\}^\{\\mathrm\{new\}\},s\_\{1,1\}^\{\\mathrm\{new\}\},\\ldots,s\_\{T,r\_\{\\mathrm\{res\}\}\}^\{\\mathrm\{new\}\}\\big\]\.\(15\)After constructing the aligned recomposition directions and assigning coefficients, CT\-Merging forms the merged LoRA update as
Δmerge=γU~diag\(snew\)V~⊤,\\Delta\_\{\\mathrm\{merge\}\}=\\gamma\\widetilde\{U\}\\operatorname\{diag\}\(s\_\{\\mathrm\{new\}\}\)\\widetilde\{V\}^\{\\top\},\(16\)whereU~\\widetilde\{U\}andV~\\widetilde\{V\}are the aligned recomposition directions,snews\_\{\\mathrm\{new\}\}is the coefficient vector, andγ\\gammais a global merge scale\. The resulting update is added to the pretrained model as in the standard model merging formulation\. The merged update has rank at mostk\+Trresk\+Tr\_\{\\mathrm\{res\}\}, consisting ofkkcommon directions andrresr\_\{\\mathrm\{res\}\}residual directions per task\. In experiments, we fix this total rank budget when varyingkkandrresr\_\{\\mathrm\{res\}\}\. Algorithm[1](https://arxiv.org/html/2607.20561#alg1)summarizes the procedure\.
## 5Experiments
### 5\.1Settings
We evaluate CT\-Merging on the individual LoRA checkpoints of ViT\-B/16, ViT\-B/32, and ViT\-L/14\[[4](https://arxiv.org/html/2607.20561#bib.bib2),[26](https://arxiv.org/html/2607.20561#bib.bib7)\]provided by\[[37](https://arxiv.org/html/2607.20561#bib.bib12)\], and the 8\-task, 12\-task, 16\-task setting accordingly\. The eight task vision benchmark from\[[28](https://arxiv.org/html/2607.20561#bib.bib5)\]consists of Cars\[[14](https://arxiv.org/html/2607.20561#bib.bib29)\], DTD\[[2](https://arxiv.org/html/2607.20561#bib.bib30)\], EuroSAT\[[7](https://arxiv.org/html/2607.20561#bib.bib31)\], GTSRB\[[8](https://arxiv.org/html/2607.20561#bib.bib32)\], MNIST\[[16](https://arxiv.org/html/2607.20561#bib.bib41)\], RESISC45\[[1](https://arxiv.org/html/2607.20561#bib.bib33)\], SUN397\[[32](https://arxiv.org/html/2607.20561#bib.bib35)\], and SVHN\[[21](https://arxiv.org/html/2607.20561#bib.bib34)\], while the 12\-task adds CIFAR100\[[15](https://arxiv.org/html/2607.20561#bib.bib37)\], Flowers102\[[22](https://arxiv.org/html/2607.20561#bib.bib38)\], OxfordIIITPet\[[24](https://arxiv.org/html/2607.20561#bib.bib36)\]and STL10\[[3](https://arxiv.org/html/2607.20561#bib.bib39)\]to the 8\-vision benchmark, and the 16\-tasks experiment add four datasets to the 12\-task configuration: FER2013\[[6](https://arxiv.org/html/2607.20561#bib.bib40)\], CIFAR10\[[15](https://arxiv.org/html/2607.20561#bib.bib37)\], FashionMNIST\[[31](https://arxiv.org/html/2607.20561#bib.bib42)\], and RenderedSST2\[[27](https://arxiv.org/html/2607.20561#bib.bib43)\]\.
We report normalized accuracy, computed as merged model per task accuracy divided by the individually fine tuned per task accuracy\. Baselines are Task Arithmetic \(TA\)\[[12](https://arxiv.org/html/2607.20561#bib.bib3)\], TIES Merging\[[33](https://arxiv.org/html/2607.20561#bib.bib9)\], KnOTS\-TIES\[[28](https://arxiv.org/html/2607.20561#bib.bib5)\], Iso\-CTS\[[19](https://arxiv.org/html/2607.20561#bib.bib1)\], and DC\-Merge\[[37](https://arxiv.org/html/2607.20561#bib.bib12)\]\. For baseline methods, we use the hyperparameters reported in their original papers\. For the global merge scale, we sweep the sameγ\\gammagrid for all methods and report the best value\. For CT\-Merging, the rank budget is set tok\+ntaskrres=128k\+n\_\{\\mathrm\{task\}\}r\_\{\\mathrm\{res\}\}=128unless otherwise specified\.
### 5\.2Main Results
Table[1](https://arxiv.org/html/2607.20561#S4.T1)reports average normalized accuracy on the DC\-Merge adapter benchmark\. We usedk=8k=8for Vit\-B backbones andk=16k=16for Vit\-L/14\. CT\-Merging gives the best result on eight of the nine backbone and task\-count settings\. It improves over DC\-Merge on all ViT\-B/16 and ViT\-L/14 settings, and also leads on ViT\-B/32 at twelve and sixteen tasks\. These results show that CT\-Merging improves SVD based LoRA recomposition across backbone sizes and task counts, and that coefficient assignment remains an important design choice even when strong directional alignment is used\.
Tables[2](https://arxiv.org/html/2607.20561#S4.T2)and[3](https://arxiv.org/html/2607.20561#S4.T3)evaluate the same method on the KnOTS checkpoints provided by\[[28](https://arxiv.org/html/2607.20561#bib.bib5)\]\. CT\-Merging gives larger gains in this setting, improving over DC\-Merge by 2\.56 points on ViT\-B/32 and 1\.51 points on ViT\-L/14\. These results suggest that our method remains effective regardless of the pretrained checkpoints\.
Figure 4:Effect of common rank k under fixed rank budget under the 8\-task ViT\-B/32 of KnOTS checkpoints\. The rank budget is set to 128\.Figure 5:Effect of the task vector coefficientγ\\gammaof the merged adapter\.
### 5\.3Ablation studies
The ablations evaluate the effect of hyperparameters and the role of consensus on fixed directions\.
#### 5\.3\.1Effect of hyperparameters
We sweep across a range ofγ\\gamma, and examine the effect ofkk, the common rank size under a fixed rank budget of 128\. The results are shown in figures[4](https://arxiv.org/html/2607.20561#S5.F4)and[5](https://arxiv.org/html/2607.20561#S5.F5)\. In the rank budget sweep, thek=0k=0setting allocates the full rank budget to task residual directions, and we observe that this setting performs worse on both models\. This also highlights the importance of the common direction structure under the same rank budget\. We observe that besidesk=0k=0, our method stays stable across moderatekkwith varying size of common rank in ViT\-L, and shows similar performance with moderate common rank on ViT B/32\. We suggest that largerkkreduces the residual rank per task, which hurts ViT B/32 and gives no consistent average gain on ViT L\.
#### 5\.3\.2Consensus Source Ablation
Table[4](https://arxiv.org/html/2607.20561#S5.T4)compares different sources for the common basis while keeping the remaining merge construction fixed\.Avg\-projectoruses the top eigenvectors of the average task subspace projector,T−1∑tUtUt⊤T^\{\-1\}\\sum\_\{t\}U\_\{t\}U\_\{t\}^\{\\top\}, which is our common\-basis source used in CT\-Merging\.Svd\-sumuses the left singular vectors of the summed task update,∑tΔt\\sum\_\{t\}\\Delta\_\{t\}, and tests whether a basis obtained directly from the mean update is sufficient\.Randomuses a random orthonormal basis with the same rank\. All variants use the same residual construction, polar projection, coefficient assignment, rank budget, and global merge scale\.
Avg\-projector gives the best average and worst\-task normalized accuracy on both backbones\. On ViT\-B/32, avg\-projector improves over summed\-update SVD by 2\.73 points in average normalized accuracy and 2\.61 points in worst\-task accuracy\. On ViT\-L/14, the corresponding gains are 1\.74 and 2\.96 points\. Avg\-projector also improves over random basis on both backbones\. These results show that repeated support across task SVD subspaces provides a stronger common basis than random directions or the SVD basis of the summed task update\.
Table 4:Consensus source ablation on the DC\-Merge CLIP ViT eight\-task setting\. The coefficient rule and global merge scale are fixed, and only the source of the common basis is varied\.
## 6Conclusion
We proposed CT\-Merging, a data free method for merging LoRA adapters through consensus directions and task\-level coefficient assignment\. The method addresses a coefficient transfer issue in SVD based LoRA merging, where component wise singular values are copied after the directions have been projected and paired into a new basis\. Our analysis shows that these inherited coefficients preserve component order but lose reliable magnitude calibration in the final basis\. CT\-Merging keeps shared and task residual directions, while replacing component wise magnitudes with task\-level RMS scales\.
Across the evaluated CLIP settings, CT\-Merging outperforms most baselines and gives larger gains on KnOTS checkpoints\. The ablations show that average projector directions provide a useful common basis\. These results support coefficient assignment and consensus construction as important design choices in LoRA adapter merging\.
## References
- \[1\]\(2017\)Remote sensing image scene classification: benchmark and state of the art\.Proceedings of the IEEE105\(10\),pp\. 1865–1883\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[2\]M\. Cimpoi, S\. Maji, I\. Kokkinos, S\. Mohamed, and A\. Vedaldi\(2014\-06\)Describing textures in the wild\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[3\]A\. Coates, A\. Ng, and H\. Lee\(2011\-11–13 Apr\)An analysis of single\-layer networks in unsupervised feature learning\.InProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics,G\. Gordon, D\. Dunson, and M\. Dudík \(Eds\.\),Proceedings of Machine Learning Research, Vol\.15,Fort Lauderdale, FL, USA,pp\. 215–223\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[4\]A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. Houlsby\(2021\)An image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=YicbFdNTTy)Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[5\]A\. A\. Gargiulo, D\. Crisostomi, M\. S\. Bucarelli, S\. Scardapane, F\. Silvestri, and E\. Rodolà\(2025\)Task singular vectors: reducing task interference in model merging\.In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Vol\.,pp\. 18695–18705\.Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p3.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px2.p1.1)\.
- \[6\]I\. J\. Goodfellow, D\. Erhan, P\. L\. Carrier, A\. Courville, M\. Mirza, B\. Hamner, W\. Cukierski, Y\. Tang, D\. Thaler, D\. Lee, Y\. Zhou, C\. Ramaiah, F\. Feng, R\. Li, X\. Wang, D\. Athanasakis, J\. Shawe\-Taylor, M\. Milakov, J\. Park, R\. Ionescu, M\. Popescu, C\. Grozea, J\. Bergstra, J\. Xie, L\. Romaszko, B\. Xu, Z\. Chuang, and Y\. Bengio\(2013\)Challenges in representation learning: a report on three machine learning contests\.External Links:1307\.0414,[Link](https://arxiv.org/abs/1307.0414)Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[7\]P\. Helber, B\. Bischke, A\. Dengel, and D\. Borth\(2019\)EuroSAT: a novel dataset and deep learning benchmark for land use and land cover classification\.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing12\(7\),pp\. 2217–2226\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[8\]S\. Houben, J\. Stallkamp, J\. Salmen, M\. Schlipsing, and C\. Igel\(2013\)Detection of traffic signs in real\-world images: the German Traffic Sign Detection Benchmark\.InInternational Joint Conference on Neural Networks,Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[9\]N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. De Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly\(2019\)Parameter\-efficient transfer learning for nlp\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p1.1)\.
- \[10\]E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p1.1)\.
- \[11\]C\. Huang, Q\. Liu, B\. Y\. Lin, T\. Pang, C\. Du, and M\. Lin\(2024\)LoraHub: efficient cross\-task generalization via dynamic lora composition\.InConference on Language Modeling \(COLM\),Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1)\.
- \[12\]G\. Ilharco, M\. T\. Ribeiro, M\. Wortsman, L\. Schmidt, H\. Hajishirzi, and A\. Farhadi\(2023\)Editing models with task arithmetic\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=6t0Kwf8-jrj)Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2607.20561#S4.SS1.p1.3),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p2.2)\.
- \[13\]X\. Jin, X\. Ren, D\. Preotiuc\-Pietro, and P\. Cheng\(2023\)Dataless knowledge fusion by merging weights of language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1)\.
- \[14\]J\. Krause, M\. Stark, J\. Deng, and L\. Fei\-Fei\(2013\-06\)3D object representations for fine\-grained categorization\.InProceedings of the IEEE International Conference on Computer Vision \(ICCV\) Workshops,Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[15\]A\. Krizhevsky, G\. Hinton,et al\.\(2009\)Learning multiple layers of features from tiny images\.pp\. 32–33\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[16\]Y\. Lecun, L\. Bottou, Y\. Bengio, and P\. Haffner\(1998\)Gradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[17\]C\. Lee, J\. Choi, C\. Lee, D\. Kim, and S\. Hong\(2026\)AdaRank: adaptive rank pruning for enhanced model merging\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p3.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px2.p1.1)\.
- \[18\]Y\. Li, Z\. Peng, J\. Zhang, J\. Guo, Y\. Duan, and Y\. Shi\(2026\)When shared knowledge hurts: spectral over\-accumulation in model merging\.InForty\-third International Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px2.p1.1)\.
- \[19\]D\. Marczak, S\. Magistri, S\. Cygert, B\. Twardowski, A\. D\. Bagdanov, and J\. van de Weijer\(2025\)No task left behind: isotropic model merging with common and task\-specific subspaces\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=RBZpAa27ls)Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p3.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p2.2)\.
- \[20\]M\. Matena and C\. Raffel\(2022\)Merging models with fisher\-weighted averaging\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1)\.
- \[21\]Y\. Netzer, T\. Wang, A\. Coates, A\. Bissacco, B\. Wu, and A\. Y\. Ng\(2011\)Reading digits in natural images with unsupervised feature learning\.InNIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011,External Links:[Link](http://ufldl.stanford.edu/housenumbers/nips2011_housenumbers.pdf)Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[22\]M\. Nilsback and A\. Zisserman\(2008\)Automated flower classification over a large number of classes\.In2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing,Vol\.,pp\. 722–729\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[23\]G\. Ortiz\-Jimenez, A\. Favero, and P\. Frossard\(2023\)Task arithmetic in the tangent space: improved editing of pre\-trained models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1)\.
- \[24\]O\. M\. Parkhi, A\. Vedaldi, A\. Zisserman, and C\. V\. Jawahar\(2012\)Cats and dogs\.InIEEE Conference on Computer Vision and Pattern Recognition,Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[25\]A\. Prabhakar, Y\. Li, K\. Narasimhan, S\. Kakade, E\. Malach, and S\. Jelassi\(2024\)LoRA soups: merging loras for practical skill composition tasks\.arXiv preprint arXiv:2410\.13025\.Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1)\.
- \[26\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. Sutskever\(2021\-18–24 Jul\)Learning transferable visual models from natural language supervision\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 8748–8763\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[27\]R\. Socher, A\. Perelygin, J\. Wu, J\. Chuang, C\. D\. Manning, A\. Ng, and C\. Potts\(2013\-10\)Recursive deep models for semantic compositionality over a sentiment treebank\.InProceedings of the 2013 Conference on Empirical Methods in Natural Language Processing,Seattle, Washington, USA,pp\. 1631–1642\.External Links:[Link](https://aclanthology.org/D13-1170/)Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[28\]G\. Stoica, P\. Ramesh, B\. Ecsedi, L\. Choshen, and J\. Hoffman\(2025\)Model merging with svd to tie the knots\.ICLR\.Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p3.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p2.2),[§5\.2](https://arxiv.org/html/2607.20561#S5.SS2.p2.1)\.
- \[29\]K\. Wang, N\. Dimitriadis, G\. Ortiz\-Jimenez, F\. Fleuret, and P\. Frossard\(2024\)Localizing task information for improved model merging and compression\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1)\.
- \[30\]M\. Wortsman, G\. Ilharco, S\. Y\. Gadre, R\. Roelofs, R\. Gontijo\-Lopes, A\. S\. Morcos, H\. Namkoong, A\. Farhadi, Y\. Carmon, S\. Kornblith, and L\. Schmidt\(2022\-17–23 Jul\)Model soups: averaging weights of multiple fine\-tuned models improves accuracy without increasing inference time\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 23965–23998\.Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1)\.
- \[31\]H\. Xiao, K\. Rasul, and R\. Vollgraf\(2017\)Fashion\-mnist: a novel image dataset for benchmarking machine learning algorithms\.External Links:1708\.07747,[Link](https://arxiv.org/abs/1708.07747)Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[32\]J\. Xiao, K\. A\. Ehinger, J\. Hays, A\. Torralba, and A\. Oliva\(2016\-08\)SUN database: exploring a large collection of scene categories\.International Journal of Computer Vision \(IJCV\)119\(1\),pp\. 3–22\.Cited by:[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1)\.
- \[33\]P\. Yadav, D\. Tam, L\. Choshen, C\. Raffel, and M\. Bansal\(2023\)TIES\-merging: resolving interference when merging models\.InThirty\-seventh Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p2.2)\.
- \[34\]E\. Yang, L\. Shen, G\. Guo, X\. Wang, X\. Cao, J\. Zhang, and D\. Tao\(2024\)Model merging in llms, mllms, and beyond: methods, theories, applications and opportunities\.arXiv preprint arXiv:2408\.07666\.Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1)\.
- \[35\]E\. Yang, Z\. Wang, L\. Shen, S\. Liu, G\. Guo, X\. Wang, and D\. Tao\(2024\)AdaMerging: adaptive model merging for multi\-task learning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1)\.
- \[36\]L\. Yu, B\. Yu, H\. Yu, F\. Huang, and Y\. Li\(2024\)Language models are super mario: absorbing abilities from homologous models as a free lunch\.InProceedings of the 41st International Conference on Machine Learning,pp\. 57755–57775\.Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p2.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px1.p1.1)\.
- \[37\]H\. Zhang, Z\. Zhou, M\. Luo, S\. Di, M\. Zhang, and T\. Wei\(2026\)DC\-merge: improving model merging with directional consistency\.InCVPR,Cited by:[§1](https://arxiv.org/html/2607.20561#S1.p3.1),[§2](https://arxiv.org/html/2607.20561#S2.SS0.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p1.1),[§5\.1](https://arxiv.org/html/2607.20561#S5.SS1.p2.2)\.Similar Articles
Crowded in B-Space: Calibrating Shared Directions for LoRA Merging
This paper introduces Pico, a data-free method that improves LoRA adapter merging by separately calibrating the output-side matrix B to reduce interference from shared directions while preserving task-specific information. Pico achieves 3.4–8.3 point accuracy improvements over existing merging methods across math, coding, finance, and medical benchmarks.
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
CoMerge is a conflict-driven preference optimization framework for merging multi-task LLMs, using self-supervised strategies to mitigate parameter interference and achieve high performance on benchmarks like MergeBench.
Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning
This paper diagnoses intra-adapter contention in MoE+LoRA fine-tuning and introduces SpawnLoRA to dynamically add sub-adapters, reducing negative transfer across domains.
Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks
The paper introduces Net Utility, a data-free metric for budget-aware LoRA merging that optimizes rank allocation across tasks, achieving +2.1% improvement on vision tasks and +2.2% on language tasks over uniform methods.
PermDoRA -- Understanding Adapter Interference in Language Models: Limits of Parameter-Space Geometry
This paper introduces DoRA-RBAC, a framework for composing LLM adapters, and tests whether geometry-aware merging improves multi-domain performance. Results show no consistent advantage over standard averaging, suggesting adapter interference is not primarily driven by parameter-space geometry.