Knowledge Pull Requests for Continual Document Authoring

Hugging Face Daily Papers Papers

Summary

Knowledge Pull Requests (KPRs) is a framework for continual document authoring that integrates new knowledge by extracting claims, filtering them, and producing a ChangeLog to make changes interpretable. It outperforms existing methods in preserving content and adding information, as evaluated on Wikipedia revisions and RAGTIME tasks.

We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or regenerate from scratch. A KPR integrates new knowledge into a document by extracting claims, filtering and routing them to sections, and flagging conflicts with existing content, producing a ChangeLog that separates what knowledge changes (claim proposal) from how the text changes (document diff). We evaluate KPRs on revising Wikipedia across languages and updating query-driven reports on RAGTIME. KPRs integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated. A KPR-revised article also grounds question answering better than a frontier model with search, which does not surface knowledge documented only in other languages.
Original Article
View Cached Full Text

Cached at: 09/24/26, 03:40 PM

Paper page - Knowledge Pull Requests for Continual Document Authoring

Source: https://huggingface.co/papers/2609.26634

Abstract

WeintroduceKnowledgePullRequests(KPRs),aframeworkforcontinualdocumentauthoringthatmakeseachchangeinterpretable.Documentsrequireongoingrevisionasnewknowledgesurfacesfromothersources,languages,ortimes,butexistingapproacheseithereditwithnoaccountofwhatknowledgechangedorregeneratefromscratch.AKPRintegratesnewknowledgeintoadocumentbyextractingclaims,filteringandroutingthemtosections,andflaggingconflictswithexistingcontent,producingaChangeLogthatseparateswhatknowledgechanges(claimproposal)fromhowthetextchanges(documentdiff).WeevaluateKPRsonrevisingWikipediaacrosslanguagesandupdatingquery-drivenreportsonRAGTIME.KPRsintegratemoreinformationandbetterpreserveexistingcontentthanrewritingfromsourcesorregeneratingfromscratch,whileaddingthemostinformationpertokengenerated.AKPR-revisedarticlealsogroundsquestionansweringbetterthanafrontiermodelwithsearch,whichdoesnotsurfaceknowledgedocumentedonlyinotherlanguages.

View arXiv pageView PDFGitHub0Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.26634 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.26634 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.26634 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

KARLA: Knowledge-base Augmented Retrieval for Language Models

arXiv cs.AI

KARLA proposes a method for LLMs to query a knowledge base during generation, enabling factual updates without retraining and improving transparency. Experiments show improved factual grounding in both short and long-form generation.

Tencent/WeKnora

GitHub Trending (daily)

WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, agents, and auto-wiki capabilities.

Note-Taking and Personal Knowledge Management

Hacker News Top

A critical response to Brennan Kenneth Brown's article questioning what note-taking PKMs have accomplished, arguing that tools like Obsidian enable people to contribute rather than directly adding to public knowledge.