Learning to Learn from Context: Synthetic Training from Perturbed Public Documents

Hugging Face Daily Papers Papers

Summary

This paper presents a synthetic training pipeline that uses perturbed public documents to generate context-dependent training data, significantly improving large language models' performance on context learning tasks like CL-bench without human annotation.

Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:08 AM

Paper page - Learning to Learn from Context: Synthetic Training from Perturbed Public Documents

Source: https://huggingface.co/papers/2609.33642

Abstract

Real-worldtasksoftenrequirelargelanguagemodels(LLMs)tolearnfromcomplextask-specificcontextratherthanpretrainedparametricknowledge.ThiscapabilityremainsaweaknessofLLMs,whilehumanannotationforsuchtaskcontextsisexpensiveanddifficulttoscale.Publichigh-qualitydocumentsareanabundantalternative,butmuchofthepublicwebhasalreadybeenconsumedduringpretraining:trainingonsuchdocumentsnaivelywouldrewardmemorizationratherthancontextlearning.Inthiswork,weattempttomakeuseofhigh-qualitypublicdocumentswithsmallperturbationsandempiricallyfindthatLLMscansuccessfullygeneratecontext-dependentreasoningtracesandanswers,whicharethenusedtotrainastudentmodel.Specifically,weconstructasynthesispipelinethat(i)rewritessourcedocumentstoreducememorizationrisk,(ii)generatesquestionsandrubricsthatrequirereasoningoverthedocument,(iii)answersthequestionswiththedocumentascontext,and(iv)admitsonlysamplesthatgenuinelydependonthedocument.Withoutanyhumanannotators,ourpipelinegeneratesabout10ksamplesfrom3.5kdocuments,andtheresultingstudentmodelsubstantiallyimprovestheperformanceonCL-bench.SFTraisesaQwen3.6-35B-A3Bstudentfrom13.7%to22.8%,andasubsequentrubric-rewardRLstagereaches24.6%,onCL-benchcomparablewithafrontiermodelofoveratrillionparameters,Qwen3.8-2.4T(23.9%).Wealsoobserveabroadtransferofimprovementstolong-contextunderstanding,instructionfollowing,andreasoning,whilecodegenerationandknowledgeremainmostlyflat.WehopethisworkprovidesareproducibleandscalablewaytoimprovetheabilityofLLMstolearnfromcontext,andtofacilitatefurtherresearchoncontext-groundedreasoning.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.33642

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33642 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33642 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33642 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Sentence-Level Contextual Entrainment in Large Language Models

arXiv cs.CL

This paper extends contextual entrainment from token-level to sentence-level, showing that even counterfactual sentences in prompts increase their probability during inference. The effect decreases with model size and is driven by 2-4% of attention heads, which can be ablated without performance loss.