Crowded in B-Space: Calibrating Shared Directions for LoRA Merging
Summary
This paper introduces Pico, a data-free method that improves LoRA adapter merging by separately calibrating the output-side matrix B to reduce interference from shared directions while preserving task-specific information. Pico achieves 3.4–8.3 point accuracy improvements over existing merging methods across math, coding, finance, and medical benchmarks.
View Cached Full Text
Cached at: 04/21/26, 07:21 AM
Paper page - Crowded in B-Space: Calibrating Shared Directions for LoRA Merging
Source: https://huggingface.co/papers/2604.16826 Published on Apr 18
·
Submitted byhttps://huggingface.co/yixuantt
yixuanon Apr 21
Abstract
LoRA adapter merging performance can be improved by separately calibrating the output-side matrix B to reduce interference from shared directions while preserving task-specific information.
Merging separately trainedLoRA adaptersis a practical alternative to joint multi-task training, but it often hurts performance. Existing methods usually treat theLoRA updateΔW = BA as a single object and do not distinguish the two LoRA matrices. We show that the main source of LoRA merge interference comes from the output-side matrix B. Across tasks, B repeatedly uses a small set ofshared directions, while A remains much more task-specific. As a result, the merged adapter overemphasizes theseshared directions, andtask-specific informationis lost. We propose Pico (Pre-merge interference calibrationinoutput-space), a data-free method that calibrates B before merge by downscaling over-shared directionsand then rescaling the merged update. Pico plugs directly into existing merging methods such asTask Arithmetic,TIES, andTSV-M. Across eight different benchmarks from math, coding, finance, and medical domains, Pico improves average accuracy by 3.4-8.3 points over the corresponding base method and achieves the best overall average performance. Pico also enablesmerged adaptersto outperform the LoRA trained with all task data. These results show that LoRA merging works better when the two LoRA matrices are treated separately.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2604\.16826
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.16826 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.16826 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.16826 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging
CT-Merging proposes a method to merge LoRA adapters by estimating consensus directions from task subspace projectors and assigning task-level RMS coefficient scales, achieving superior performance on the DC-Merge CLIP adapter benchmark.
From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging
This paper introduces Preference Delta Aggregation (PDA) and Geometric Alignment Merging (GAM) to aggregate multiple 'weak' preference signals from weaker model pairs via LoRA merging, improving strong LLMs on knowledge reasoning and agentic search tasks by over 6% on average.
Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning
This paper diagnoses intra-adapter contention in MoE+LoRA fine-tuning and introduces SpawnLoRA to dynamically add sub-adapters, reducing negative transfer across domains.
Beyond LoRA: Is Sparsity-Induced Adaptation Better?
This paper proposes sparsity-induced adaptations to LoRA, including Cheap LoRA (cLA) and a chained circulant variant (c³LA), and provides theoretical generalization bounds along with empirical evaluations showing up to 10% training time reduction and 15% peak GPU memory savings while maintaining competitive performance.
PermDoRA -- Understanding Adapter Interference in Language Models: Limits of Parameter-Space Geometry
This paper introduces DoRA-RBAC, a framework for composing LLM adapters, and tests whether geometry-aware merging improves multi-domain performance. Results show no consistent advantage over standard averaging, suggesting adapter interference is not primarily driven by parameter-space geometry.