Sample-Efficient Learning from Agent Experience

Hugging Face Daily Papers Papers

Summary

Proposes Experience Distillation, a method that internalizes in-context learning gains from agent interaction histories into model weights without requiring additional environment interaction, achieving significant sample efficiency improvements on software engineering and text-adventure tasks.

Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interaction histories without sacrificing environment sample efficiency remains underexplored. We term this problem Experience Distillation and develop an implementation that requires no further environment interaction beyond the collected experience. Experiments on 749 curated software-engineering tasks and six text-adventure games show that it retains at least 64.8\% of the gains from in-context learning across both domains, whereas direct supervised fine-tuning on the collected experience recovers only 3.8\%. Compared with classical reinforcement-learning baselines, in-context learning from trial-and-error experience followed by Experience Distillation matches their performance with at least \(9.6\times\) fewer environment samples.
Original Article
View Cached Full Text

Cached at: 07/24/26, 09:00 AM

Paper page - Sample-Efficient Learning from Agent Experience

Source: https://huggingface.co/papers/2607.21051

Abstract

Real-worldagentlearningisoftenconstrainedbycostlyenvironmentinteractions,suchasrunningtime-consumingexperimentsorobtaininghumanfeedback.In-contextlearningoffersahighlysample-efficientwayforagentstolearnfromtheirowninteractionhistories,butitsgainsdisappearoncethatexperienceisremovedfromthecontext.Separately,contextdistillationprovidesamechanismforinternalizingcontextualinformationintomodelweights.However,applyingittoagents’interactionhistorieswithoutsacrificingenvironmentsampleefficiencyremainsunderexplored.WetermthisproblemExperienceDistillationanddevelopanimplementationthatrequiresnofurtherenvironmentinteractionbeyondthecollectedexperience.Experimentson749curatedsoftware-engineeringtasksandsixtext-adventuregamesshowthatitretainsatleast64.8\%ofthegainsfromin-contextlearningacrossbothdomains,whereasdirectsupervisedfine-tuningonthecollectedexperiencerecoversonly3.8\%.Comparedwithclassicalreinforcement-learningbaselines,in-contextlearningfromtrial-and-errorexperiencefollowedbyExperienceDistillationmatchestheirperformancewithatleast\(9.6\times\)fewerenvironmentsamples.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.21051 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.21051 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.21051 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Rethinking Experience Utilization in Self-Evolving Language Model Agents

arXiv cs.CL

This paper introduces ExpWeaver, a framework that optimizes how self-evolving language model agents utilize past experiences during runtime decision-making. It demonstrates that selectively invoking experience based on reasoning uncertainty improves performance across various environments and models.