Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Hugging Face Daily Papers Papers

Summary

IAR is a three-stage post-training framework that internalizes structured document knowledge into language models for retrieval-free question answering, enhancing both domain-specific accuracy and general model performance.

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.
Original Article
View Cached Full Text

Cached at: 08/21/26, 04:07 AM

Paper page - Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Source: https://huggingface.co/papers/2608.20281

Abstract

IAR is a three-stage post-training framework that injects structured document knowledge into language models, aligns them for retrieval-free question answering, and recovers general capabilities, improving both domain accuracy and general performance.

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting asdocument knowledge internalization: converting a fixed corpus into usable parametric knowledge forretrieval-free question answering. We proposeIAR(Inject, Align, and Recover), a three-stage post-training framework that separatesstructured document knowledge injection,QA behavior alignment, andgeneral ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families,IARimproves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison,IARimproves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show thatLoRAandFAPMcan win individual general metrics, but among methods that also reach leading or near-leading domain internalization,IARretains one of the strongest general profiles.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2608\.20281

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.20281 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.20281 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.20281 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

Understanding the Behaviors of Environment-aware Information Retrieval

Hugging Face Daily Papers

This paper presents the first systematic analysis of how large language models can learn to adapt query formulation strategies for different retrievers using reinforcement learning, revealing distinct optimal query styles and introducing a branching-based rollout technique for multi-retrieval-step training stability.