Sample-Efficient Learning from Agent Experience
Summary
Proposes Experience Distillation, a method that internalizes in-context learning gains from agent interaction histories into model weights without requiring additional environment interaction, achieving significant sample efficiency improvements on software engineering and text-adventure tasks.
View Cached Full Text
Cached at: 07/24/26, 09:00 AM
Paper page - Sample-Efficient Learning from Agent Experience
Source: https://huggingface.co/papers/2607.21051
Abstract
Real-worldagentlearningisoftenconstrainedbycostlyenvironmentinteractions,suchasrunningtime-consumingexperimentsorobtaininghumanfeedback.In-contextlearningoffersahighlysample-efficientwayforagentstolearnfromtheirowninteractionhistories,butitsgainsdisappearoncethatexperienceisremovedfromthecontext.Separately,contextdistillationprovidesamechanismforinternalizingcontextualinformationintomodelweights.However,applyingittoagents’interactionhistorieswithoutsacrificingenvironmentsampleefficiencyremainsunderexplored.WetermthisproblemExperienceDistillationanddevelopanimplementationthatrequiresnofurtherenvironmentinteractionbeyondthecollectedexperience.Experimentson749curatedsoftware-engineeringtasksandsixtext-adventuregamesshowthatitretainsatleast64.8\%ofthegainsfromin-contextlearningacrossbothdomains,whereasdirectsupervisedfine-tuningonthecollectedexperiencerecoversonly3.8\%.Comparedwithclassicalreinforcement-learningbaselines,in-contextlearningfromtrial-and-errorexperiencefollowedbyExperienceDistillationmatchestheirperformancewithatleast\(9.6\times\)fewerenvironmentsamples.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.21051 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.21051 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.21051 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Proposes SEED, a self-evolving on-policy distillation framework that converts completed trajectories into hindsight skills to improve reinforcement learning for interactive agent tasks, achieving consistent performance gains and sample efficiency.
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
EDGE introduces a framework for guided exploration in agentic reinforcement learning by distilling experiences into policies, improving performance on tasks like ALFWorld and WebShop.
Rethinking Experience Utilization in Self-Evolving Language Model Agents
This paper introduces ExpWeaver, a framework that optimizes how self-evolving language model agents utilize past experiences during runtime decision-making. It demonstrates that selectively invoking experience based on reasoning uncertainty improves performance across various environments and models.
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
This paper proposes the EDV framework, which uses multiple heterogeneous agents in execute-distill-verify stages to build reliable experiences for LLM agents, preventing self-confirmatory errors and improving performance on long-horizon benchmarks.
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
The paper introduces Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent via hierarchical memory, improving tool-use benchmark performance by 3.4–27.2 percentage points.