Tag
SMAT introduces a simple and efficient merge-aware training method that jointly optimizes expert loss and expected loss at simulated merged parameters, improving merged model performance with minimal computational overhead.
This paper investigates whether routing drift in merged Mixture-of-Experts Large Language Models indicates failure, and proposes a routing analysis toolkit and Selective Router Repair method for assessment.
The paper investigates fine-tuning strategies for customer support LLMs, comparing multi-task training, sequential updates, and model merging across multiple model families. It concludes that multi-task full fine-tuning is the most robust default, while specialist models degrade off-task and require reliable routing.
The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.
This survey paper defines model fusion and organizes research into three levels—parameter, representation, and behavior—for integrating capabilities in large language models.
CoMerge is a conflict-driven preference optimization framework for merging multi-task LLMs, using self-supervised strategies to mitigate parameter interference and achieve high performance on benchmarks like MergeBench.
The paper introduces Net Utility, a data-free metric for budget-aware LoRA merging that optimizes rank allocation across tasks, achieving +2.1% improvement on vision tasks and +2.2% on language tasks over uniform methods.
The paper investigates the geometric challenges in merging differentially private task models and introduces DP-Merging, a framework to enhance mergeability while maintaining privacy guarantees.
This paper reveals a jailbreak risk in model merging even when constituent models are safety-aligned, and proposes Basin-Aware Jailbreak (BAJ) to generate transferable adversarial suffixes across merged model families.
GRIP introduces a reward-guided parameter interpolation framework to merge reasoning and instruction models, improving accuracy-efficiency trade-off for LLM reasoning without retraining.
CORAM introduces a coherent orthogonal rotation method for model merging that partitions weight matrices into row slices, uses SVD in the base model's frame, and merges updates on manifolds to improve accuracy over existing techniques.
This paper proposes Activation-Prune-Merge (APM), a training-free framework for cross-scale fusion that improves smaller language models using larger donors without semantic alignment, achieving performance gains on multiple benchmarks.
This paper presents CABS+, an enhanced model merging method that uses conflict-aware sparsification and adaptive weight allocation to improve multi-task model merging efficiency and performance, achieving significant gains over prior methods across 27 datasets.
HyperFix proposes a lightweight hypernetwork to predict nonlinear corrections for task vector merging across varying task subsets, reducing per-subset tuning costs and outperforming existing methods.
This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.
Introduces AgentPatch, a training-free coarse-to-fine repair framework for merging agentic multimodal large language models, addressing asymmetric capability preservation and behavior-critical forgetting.
A new sub-6B sparse activation AI model built with a fusion architecture combines weights from LFM2.5-2.6B and Qwen3.6-35B-A3B, achieving near-Qwen3.6-35B performance at a fraction of the size.
Presents B1ade, a minimalist RAG architecture with a 335M zero-training embedding model and a 1B SLM trained via GRPO on 723M tokens, showing emergent attribution behavior and competitive QA performance without large-scale pretraining.
This paper introduces SkillSmith, an LLM augmented to reason over both prefix weights and textual knowledge, enabling instruction-steered synthesis of new parametric skills that outperform text-only and weight-only baselines.
Liquid AI's head of post-training explains how to build a sub-1GB on-device model in 20 minutes using LFM2.5, on-policy preference alignment, agentic RL, curriculum training, and iterative model merging, achieving tool-calling reliability that beats much larger models.