model-merging

Tag

Cards List
#model-merging

SMAT: Simple and Efficient Merge-Aware Training

Hugging Face Daily Papers ↗ · 4d ago Cached

SMAT introduces a simple and efficient merge-aware training method that jointly optimizes expert loss and expected loss at simulated merged parameters, improving merged model performance with minimal computational overhead.

0 favorites 0 likes
#model-merging

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Hugging Face Daily Papers ↗ · 5d ago Cached

This paper investigates whether routing drift in merged Mixture-of-Experts Large Language Models indicates failure, and proposes a routing analysis toolkit and Selective Router Repair method for assessment.

0 favorites 0 likes
#model-merging

Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs

arXiv cs.CL ↗ · 2026-09-24 Cached

The paper investigates fine-tuning strategies for customer support LLMs, comparing multi-task training, sequential updates, and model merging across multiple model families. It concludes that multi-task full fine-tuning is the most robust default, while specialist models degrade off-task and require reliable routing.

0 favorites 0 likes
#model-merging

ACLArena: Agent Continue Learning in Multi-stage Post-training

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.

0 favorites 0 likes
#model-merging

From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models

arXiv cs.CL ↗ · 2026-09-18 Cached

This survey paper defines model fusion and organizes research into three levels—parameter, representation, and behavior—for integrating capabilities in large language models.

0 favorites 0 likes
#model-merging

CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

arXiv cs.AI ↗ · 2026-09-03 Cached

CoMerge is a conflict-driven preference optimization framework for merging multi-task LLMs, using self-supervised strategies to mitigate parameter interference and achieve high performance on benchmarks like MergeBench.

0 favorites 0 likes
#model-merging

Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

The paper introduces Net Utility, a data-free metric for budget-aware LoRA merging that optimizes rank allocation across tasks, achieving +2.1% improvement on vision tasks and +2.2% on language tasks over uniform methods.

0 favorites 0 likes
#model-merging

When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy

arXiv cs.LG ↗ · 2026-08-28 Cached

The paper investigates the geometric challenges in merging differentially private task models and introduces DP-Merging, a framework to enhance mergeability while maintaining privacy guarantees.

0 favorites 0 likes
#model-merging

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper reveals a jailbreak risk in model merging even when constituent models are safety-aligned, and proposes Basin-Aware Jailbreak (BAJ) to generate transferable adversarial suffixes across merged model families.

0 favorites 0 likes
#model-merging

GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning

arXiv cs.CL ↗ · 2026-08-27 Cached

GRIP introduces a reward-guided parameter interpolation framework to merge reasoning and instruction models, improving accuracy-efficiency trade-off for LLM reasoning without retraining.

0 favorites 0 likes
#model-merging

CORAM: Coherent Orthogonal Rotation for Model Merging

arXiv cs.LG ↗ · 2026-08-19 Cached

CORAM introduces a coherent orthogonal rotation method for model merging that partitions weight matrices into row slices, uses SVD in the base model's frame, and merges updates on manifolds to improve accuracy over existing techniques.

0 favorites 0 likes
#model-merging

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

arXiv cs.LG ↗ · 2026-08-17 Cached

This paper proposes Activation-Prune-Merge (APM), a training-free framework for cross-scale fusion that improves smaller language models using larger donors without semantic alignment, achieving performance gains on multiple benchmarks.

0 favorites 0 likes
#model-merging

CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper presents CABS+, an enhanced model merging method that uses conflict-aware sparsification and adaptive weight allocation to improve multi-task model merging efficiency and performance, achieving significant gains over prior methods across 27 datasets.

0 favorites 0 likes
#model-merging

HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging

arXiv cs.LG ↗ · 2026-08-13 Cached

HyperFix proposes a lightweight hypernetwork to predict nonlinear corrections for task vector merging across varying task subsets, reducing per-subset tuning costs and outperforming existing methods.

0 favorites 0 likes
#model-merging

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

arXiv cs.AI ↗ · 2026-08-10 Cached

This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.

0 favorites 0 likes
#model-merging

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

arXiv cs.AI ↗ · 2026-08-10 Cached

Introduces AgentPatch, a training-free coarse-to-fine repair framework for merging agentic multimodal large language models, addressing asymmetric capability preservation and behavior-critical forgetting.

0 favorites 0 likes
#model-merging

@maximelabonne: Wow, this gives me flashbacks of early model merging. Complete insanity, I love it!

X AI KOLs Following ↗ · 2026-08-06 Cached

A new sub-6B sparse activation AI model built with a fusion architecture combines weights from LFM2.5-2.6B and Qwen3.6-35B-A3B, achieving near-Qwen3.6-35B performance at a fraction of the size.

0 favorites 0 likes
#model-merging

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

arXiv cs.CL ↗ · 2026-07-31 Cached

Presents B1ade, a minimalist RAG architecture with a 335M zero-training embedding model and a 1B SLM trained via GRPO on 723M tokens, showing emergent attribution behavior and competitive QA performance without large-scale pretraining.

0 favorites 0 likes
#model-merging

SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper introduces SkillSmith, an LLM augmented to reason over both prefix weights and textual knowledge, enabling instruction-steered synthesis of new parametric skills that outperform text-only and weight-only baselines.

0 favorites 0 likes
#model-merging

@h100envy: Liquid AI's head of post-training explained how they built a small model that runs on-device under 1 GB in 20 minutes -…

X AI KOLs Following ↗ · 2026-07-30 Cached

Liquid AI's head of post-training explains how to build a sub-1GB on-device model in 20 minutes using LFM2.5, on-policy preference alignment, agentic RL, curriculum training, and iterative model merging, achieving tool-calling reliability that beats much larger models.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback