knowledge-injection

Tag

Cards List
#knowledge-injection

Decoupled Mixture-of-Experts for Parametric Knowledge Injection

arXiv cs.CL · 2026-06-15 Cached

Decoupled Mixture-of-Experts (DMoE) proposes a modular architecture for parametric knowledge injection, decoupling experts and router from the base model to enable efficient auto-regressive inference and mitigate catastrophic forgetting.

0 favorites 0 likes
#knowledge-injection

MixSD: Mixed Contextual Self-Distillation for Knowledge Injection

Hugging Face Daily Papers · 2026-05-16 Cached

MixSD proposes a self-distillation method for knowledge injection in language models that aligns supervision with the model's native distribution, reducing catastrophic forgetting during fine-tuning. It achieves near-perfect memorization while retaining up to 100% of base capabilities, vastly outperforming standard SFT.

0 favorites 0 likes
← Back to home

Submit Feedback