knowledge-injection

Tag

Cards List
#knowledge-injection

KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training

arXiv cs.CL · 3d ago Cached

KItCAT introduces a lightweight training strategy for auto-regressive LLMs that uses input corruption to generate diverse training inputs, improving knowledge injection from niche documents without costly paraphrasing.

0 favorites 0 likes
#knowledge-injection

From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

arXiv cs.CL · 2026-08-27 Cached

This paper proposes Golden-GRPO Injection (GRIN), a mixed-policy reinforcement learning framework for continual knowledge injection in large language models, overcoming limitations of supervised fine-tuning. Experiments demonstrate superior performance on knowledge absorption benchmarks.

0 favorites 0 likes
#knowledge-injection

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

arXiv cs.AI · 2026-07-29 Cached

This paper introduces RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection into multimodal large language models. It improves injected knowledge accuracy while limiting drift in non-updated behaviors.

0 favorites 0 likes
#knowledge-injection

TokenMem: Faithful Knowledge Injection for Frozen LLMs

arXiv cs.AI · 2026-07-28 Cached

TokenMem injects knowledge into frozen LLMs via a dedicated cross-attention channel, training a thin gating adapter through two-phase curriculum to improve knowledge compliance under counterfactual knowledge, achieving 69-70% KC compared to 20-52% for vanilla RAG.

0 favorites 0 likes
#knowledge-injection

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

arXiv cs.CL · 2026-07-24 Cached

Presents a lightweight knowledge-injection framework for zero-shot ICU delirium prediction that augments structured EHR data summaries with external clinical knowledge at inference time, improving AUROC by up to 8.57 percentage points on LLaMA models without fine-tuning.

0 favorites 0 likes
#knowledge-injection

Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations

arXiv cs.CL · 2026-07-24 Cached

This paper investigates whether layer-wise differences exist in Mixture-of-Experts (MoE) models and proposes EAACD, an expert-aware adaptive contrast decoding method that leverages expert activation patterns in higher layers to reduce hallucinations in LLMs for QA tasks.

0 favorites 0 likes
#knowledge-injection

@rohanpaul_ai: New paper from @NaceAI shows a possible way to add large bodies of knowledge without rewriting the language model’s cor…

X AI KOLs Following · 2026-07-23 Cached

A new paper from NaceAI proposes a hypernetwork-based method for injecting knowledge into large language models without modifying their core parameters, potentially enabling efficient continual learning. The approach uses generated low-rank adapters to encode new facts while keeping the base model frozen.

0 favorites 0 likes
#knowledge-injection

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Hugging Face Daily Papers · 2026-07-21 Cached

The paper investigates scaling laws for hypernetwork-based knowledge injection into LLMs, finding predictive power law scaling and reliable out-of-distribution generalization, establishing hypernetworks as a scalable alternative to LoRA and full fine-tuning.

0 favorites 0 likes
#knowledge-injection

Decoupled Mixture-of-Experts for Parametric Knowledge Injection

arXiv cs.CL · 2026-06-15 Cached

Decoupled Mixture-of-Experts (DMoE) proposes a modular architecture for parametric knowledge injection, decoupling experts and router from the base model to enable efficient auto-regressive inference and mitigate catastrophic forgetting.

0 favorites 0 likes
#knowledge-injection

MixSD: Mixed Contextual Self-Distillation for Knowledge Injection

Hugging Face Daily Papers · 2026-05-16 Cached

MixSD proposes a self-distillation method for knowledge injection in language models that aligns supervision with the model's native distribution, reducing catastrophic forgetting during fine-tuning. It achieves near-perfect memorization while retaining up to 100% of base capabilities, vastly outperforming standard SFT.

0 favorites 0 likes
← Back to home

Submit Feedback