model-merging

Tag

Cards List
#model-merging

CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

arXiv cs.AI · yesterday Cached

This paper presents CABS+, an enhanced model merging method that uses conflict-aware sparsification and adaptive weight allocation to improve multi-task model merging efficiency and performance, achieving significant gains over prior methods across 27 datasets.

0 favorites 0 likes
#model-merging

HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging

arXiv cs.LG · 2d ago Cached

HyperFix proposes a lightweight hypernetwork to predict nonlinear corrections for task vector merging across varying task subsets, reducing per-subset tuning costs and outperforming existing methods.

0 favorites 0 likes
#model-merging

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

arXiv cs.AI · 5d ago Cached

This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.

0 favorites 0 likes
#model-merging

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

arXiv cs.AI · 5d ago Cached

Introduces AgentPatch, a training-free coarse-to-fine repair framework for merging agentic multimodal large language models, addressing asymmetric capability preservation and behavior-critical forgetting.

0 favorites 0 likes
#model-merging

@maximelabonne: Wow, this gives me flashbacks of early model merging. Complete insanity, I love it!

X AI KOLs Following · 2026-08-06 Cached

A new sub-6B sparse activation AI model built with a fusion architecture combines weights from LFM2.5-2.6B and Qwen3.6-35B-A3B, achieving near-Qwen3.6-35B performance at a fraction of the size.

0 favorites 0 likes
#model-merging

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

arXiv cs.CL · 2026-07-31 Cached

Presents B1ade, a minimalist RAG architecture with a 335M zero-training embedding model and a 1B SLM trained via GRPO on 723M tokens, showing emergent attribution behavior and competitive QA performance without large-scale pretraining.

0 favorites 0 likes
#model-merging

SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge

arXiv cs.CL · 2026-07-31 Cached

This paper introduces SkillSmith, an LLM augmented to reason over both prefix weights and textual knowledge, enabling instruction-steered synthesis of new parametric skills that outperform text-only and weight-only baselines.

0 favorites 0 likes
#model-merging

@h100envy: Liquid AI's head of post-training explained how they built a small model that runs on-device under 1 GB in 20 minutes -…

X AI KOLs Following · 2026-07-30 Cached

Liquid AI's head of post-training explains how to build a sub-1GB on-device model in 20 minutes using LFM2.5, on-policy preference alignment, agentic RL, curriculum training, and iterative model merging, achieving tool-calling reliability that beats much larger models.

0 favorites 0 likes
#model-merging

Making Open-Source Text LLM Watermarks Durable Against Merging

arXiv cs.CL · 2026-07-24 Cached

This paper proposes Merge-Adversarial Training to make text watermarks in open-source LLMs survive model merging, outperforming baselines while preserving downstream capabilities.

0 favorites 0 likes
#model-merging

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

arXiv cs.LG · 2026-07-15 Cached

This paper challenges the common assumption that domain experts for model merging should be trained to their optimal validation loss, showing that the optimal training duration depends strongly on the merging method. Simple averaging degrades with overfitting while sparsification-based methods benefit from training past the optimum, suggesting that training duration and merging method should be chosen jointly.

0 favorites 0 likes
#model-merging

Interference and Retention in Continual Learning

arXiv cs.LG · 2026-07-13 Cached

This paper proposes modeling forgetting in continual learning as interference between tasks and introduces Interference-Gated Functional Allocation (IGFA), a replay-free, Fisher-free method that shares or protects directions based on task geometry, achieving lossless retention when tasks are separable.

0 favorites 0 likes
#model-merging

Spectral Rewiring for Exploration, Purification, and Model Merging

arXiv cs.LG · 2026-07-07 Cached

The paper introduces Subspace-Aligned Rewiring (SAR), a post-hoc editing method that retains the spectral core of RL updates to preserve reasoning gains, remove interference, and enable model merging across experts, achieving strong performance with minimal parameters.

0 favorites 0 likes
#model-merging

Can Model Merging Improve Aggregation in DiLoCo?

arXiv cs.LG · 2026-07-07 Cached

This paper proposes using model merging techniques, specifically Iso-C aggregation, to improve the aggregation step in DiLoCo distributed training, resulting in a new method called IsoLoCo that outperforms DiLoCo on language model pre-training.

0 favorites 0 likes
#model-merging

Efficient Decentralized Multi-task Dataset Valuation via Model Merging

arXiv cs.CL · 2026-07-07 Cached

The paper introduces DMVM, a decentralized multi-task dataset valuation framework that leverages task arithmetic and model merging to estimate dataset contributions without retraining or data sharing, enabling scalable and privacy-preserving valuation for data marketplaces.

0 favorites 0 likes
#model-merging

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

arXiv cs.LG · 2026-07-03 Cached

This paper frames model merging as probabilistic inference under a product-of-experts scenario, showing that existing methods are special cases and proposing a heavy-tailed Cauchy expert design that better captures real residual behavior, achieving significant improvements over state-of-the-art baselines.

0 favorites 0 likes
#model-merging

Spectral Rewiring for Exploration, Purification, and Model Merging

Hugging Face Daily Papers · 2026-07-03 Cached

This paper introduces SAR, a training-free method that projects RL updates onto a compact reasoning core in spectral space, enabling purification, improved exploration, and stronger model merging.

0 favorites 0 likes
#model-merging

Optimizing Visual Generative Models via Distribution-wise Rewards

Hugging Face Daily Papers · 2026-07-02 Cached

This paper presents a reinforcement learning framework for visual generative models that uses distribution-wise rewards, with a subset-replace strategy for efficiency, improving image diversity and quality while addressing mode collapse and reward hacking.

0 favorites 0 likes
#model-merging

Sakana Fugu

Product Hunt · 2026-06-22

Sakana Fugu is a new tool that enables combining multiple AI models into one, inspired by the concept of 'one model to command them all'.

0 favorites 0 likes
#model-merging

Enhancing Multilingual Reasoning via Steerable Model Merging

arXiv cs.CL · 2026-06-18 Cached

This paper proposes ST-Merge, a steerable model merging framework that uses a gated cross-attention mechanism to adaptively modulate contributions of a multilingual model and a reasoning model, outperforming fixed merging approaches on multilingual reasoning benchmarks across 21 languages.

0 favorites 0 likes
#model-merging

PACT: Preserving Anchored Cores in Task-vectors for Model Merging

arXiv cs.LG · 2026-06-18 Cached

The paper identifies 'Load-Bearing Wall' dimensions in pre-trained models that retain task-specific knowledge not fully captured by task vectors in model merging, and proposes PACT (PreserveAnchoredCores) to preserve these cores, achieving state-of-the-art performance across benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback