Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
Summary
This paper presents a synthetic training pipeline that uses perturbed public documents to generate context-dependent training data, significantly improving large language models' performance on context learning tasks like CL-bench without human annotation.
View Cached Full Text
Cached at: 09/29/26, 04:08 AM
Paper page - Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
Source: https://huggingface.co/papers/2609.33642
Abstract
Real-worldtasksoftenrequirelargelanguagemodels(LLMs)tolearnfromcomplextask-specificcontextratherthanpretrainedparametricknowledge.ThiscapabilityremainsaweaknessofLLMs,whilehumanannotationforsuchtaskcontextsisexpensiveanddifficulttoscale.Publichigh-qualitydocumentsareanabundantalternative,butmuchofthepublicwebhasalreadybeenconsumedduringpretraining:trainingonsuchdocumentsnaivelywouldrewardmemorizationratherthancontextlearning.Inthiswork,weattempttomakeuseofhigh-qualitypublicdocumentswithsmallperturbationsandempiricallyfindthatLLMscansuccessfullygeneratecontext-dependentreasoningtracesandanswers,whicharethenusedtotrainastudentmodel.Specifically,weconstructasynthesispipelinethat(i)rewritessourcedocumentstoreducememorizationrisk,(ii)generatesquestionsandrubricsthatrequirereasoningoverthedocument,(iii)answersthequestionswiththedocumentascontext,and(iv)admitsonlysamplesthatgenuinelydependonthedocument.Withoutanyhumanannotators,ourpipelinegeneratesabout10ksamplesfrom3.5kdocuments,andtheresultingstudentmodelsubstantiallyimprovestheperformanceonCL-bench.SFTraisesaQwen3.6-35B-A3Bstudentfrom13.7%to22.8%,andasubsequentrubric-rewardRLstagereaches24.6%,onCL-benchcomparablewithafrontiermodelofoveratrillionparameters,Qwen3.8-2.4T(23.9%).Wealsoobserveabroadtransferofimprovementstolong-contextunderstanding,instructionfollowing,andreasoning,whilecodegenerationandknowledgeremainmostlyflat.WehopethisworkprovidesareproducibleandscalablewaytoimprovetheabilityofLLMstolearnfromcontext,andtofacilitatefurtherresearchoncontext-groundedreasoning.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.33642
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33642 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33642 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33642 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
When Context Misleads: In-context Learning with Jurisdiction in Large Language Models
This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.
Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation
This paper proposes a reinforcement learning approach to enable large language models to translate unseen languages by leveraging in-context linguistic knowledge, outperforming in-context learning and supervised fine-tuning.
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
The paper introduces ConceptGuard, a benchmark for evaluating context-sensitive unlearning in large language models using dual-use concepts, revealing that current unlearning techniques perform poorly under this practical evaluation framework.
A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance
This paper presents a framework for generating context-specific large language model benchmark datasets using expert guidance and synthetic data, improving validity and scalability over existing methods.
Sentence-Level Contextual Entrainment in Large Language Models
This paper extends contextual entrainment from token-level to sentence-level, showing that even counterfactual sentences in prompts increase their probability during inference. The effect decreases with model size and is driven by 2-4% of attention heads, which can be ablated without performance loss.