EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

arXiv cs.CL Papers

Summary

EDGE introduces a framework for guided exploration in agentic reinforcement learning by distilling experiences into policies, improving performance on tasks like ALFWorld and WebShop.

arXiv:2608.21946v1 Announce Type: new Abstract: Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. On ALFWorld and WebShop, EDGE improves over GRPO by 8.3 and 12.5 success-rate points at the 7B scale and retains 96.0% of its scaffolded performance when external experiences are removed at inference time. The code is available at https://github.com/xvolcano02/EDGE.
Original Article
View Cached Full Text

Cached at: 08/25/26, 04:22 AM

# EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
Source: [https://arxiv.org/html/2608.21946](https://arxiv.org/html/2608.21946)
Can XieAffiliation:\[5pt\] School of Artificial Intelligence, University of Chinese Academy of SciencesAffiliation:Institute of Automation, Chinese Academy of SciencesYuyi ZhouAffiliation:Institute of Automation, Chinese Academy of SciencesAffiliation:School of Advanced Interdisciplinary Sciences, University of Chinese Academy of SciencesWen YangAffiliation:\[5pt\] School of Artificial Intelligence, University of Chinese Academy of SciencesAffiliation:Institute of Automation, Chinese Academy of SciencesZiyi ZhangAffiliation:Institute of Automation, Chinese Academy of SciencesAffiliation:School of Advanced Interdisciplinary Sciences, University of Chinese Academy of SciencesSiyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang11footnotemark:1Thanks:Corresponding authorsAffiliation:\[5pt\] School of Artificial Intelligence, University of Chinese Academy of SciencesAffiliation:\[5pt\] School of Artificial Intelligence, University of Chinese Academy of SciencesAffiliation:\[5pt\] School of Artificial Intelligence, University of Chinese Academy of SciencesAffiliation:Institute of Automation, Chinese Academy of SciencesAffiliation:Institute of Automation, Chinese Academy of SciencesAffiliation:Institute of Automation, Chinese Academy of SciencesAffiliation:Institute of Automation, Chinese Academy of SciencesAffiliation:Wuhan AI Research\{xiecan2024,shuo\.ren\}@ia\.ac\.cn,jjzhang@nlpr\.ia\.ac\.cn

###### Abstract

Reinforcement learning with outcome\-based objectives such as GRPO enables LLM\-based agents to solve complex, long\-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update\. Existing experience\-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy’s evolving capability and create persistent dependencies on external retrieval\. We proposeEDGE\(Experience\-Distillation forGuidedExploration\), a framework that treats retrieved experiences as temporary training\-time scaffolds and progressively internalizes their benefits into the parametric policy\. Concretely, EDGE partitions each rollout group into experience\-conditioned and experience\-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse\-KL objective on its own empirical support\. A co\-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves\. On ALFWorld and WebShop, EDGE improves over GRPO by 8\.3 and 12\.5 success\-rate points at the 7B scale and retains 96\.0% of its scaffolded performance when external experiences are removed at inference time\. The code is available at[https://github\.com/xvolcano02/EDGE](https://github.com/xvolcano02/EDGE)\.

![Refer to caption](https://arxiv.org/html/2608.21946v1/overview-v6.png)Figure 1:Overview of the EDGE framework\.\(a\) Experience\-guided exploration scaffolding uses contrastive rollouts to estimate the marginal gainΔe\\Delta\_\{e\}of a retrieved experience\. \(b\) Gain\-gated distillation internalizes beneficial scaffold behavior into the base policy only whenΔe\>0\\Delta\_\{e\}\>0\. \(c\) The experience bank co\-evolves with the policy via failure\-driven expansion and utility\-based pruning, yielding a retrieval\-free policy at deployment\.## 1Introduction

Reinforcement learning \(RL\) with outcome\-based objectives such as GRPO\([24](https://arxiv.org/html/2608.21946#bib.bib7)\)has become a standard post\-training paradigm for LLM\-based agents that reason, plan, and act through multi\-turn interactions\([38](https://arxiv.org/html/2608.21946#bib.bib12);[25](https://arxiv.org/html/2608.21946#bib.bib13);[37](https://arxiv.org/html/2608.21946#bib.bib5);[26](https://arxiv.org/html/2608.21946#bib.bib6);[9](https://arxiv.org/html/2608.21946#bib.bib1);[31](https://arxiv.org/html/2608.21946#bib.bib2);[11](https://arxiv.org/html/2608.21946#bib.bib3)\)\. Yet current agentic RL uses its own experience poorly\. A single rollout may contain reusable patterns—effective task decompositions, recovery strategies, dead\-end avoidance, but once a trajectory contributes a scalar advantage to one policy update, these patterns are largely discarded\. The agent must therefore rediscover them from scratch, a particularly costly failure mode in sparse\-reward, long\-horizon environments where agentic RL is most needed\.

A growing body of work addresses this inefficiency by augmenting agents with external experience at inference time, through episodic reflections\([25](https://arxiv.org/html/2608.21946#bib.bib13);[45](https://arxiv.org/html/2608.21946#bib.bib14)\), persistent memory\([4](https://arxiv.org/html/2608.21946#bib.bib17);[8](https://arxiv.org/html/2608.21946#bib.bib28)\), or skill libraries\([34](https://arxiv.org/html/2608.21946#bib.bib18);[13](https://arxiv.org/html/2608.21946#bib.bib19)\)\. While effective, these approaches share two structural limitations that become apparent as training progresses\. Experience utility is inherently policy\-dependent: guidance that accelerates an undertrained policy can become redundant\-\-\-or actively harmful\-\-\-once the corresponding behavior has been internalized, yet most methods apply experience unconditionally or filter it by static heuristics\. Moreover, if the agent must retrieve experience at deployment, part of its competence resides in the context window rather than in the model parameters, incurring persistent token overhead and sensitivity to retrieval noise\.111These are not hypothetical concerns: in our experiments, naive memory augmentation \(e\.g\., EvolveR, GRPO\+Mem0\) degrades performance well below vanilla GRPO, confirming that unfiltered experience injection is unreliable\.These observations suggest that experience reuse in agentic RL should be treated as a dynamic lifecycle rather than a static retrieval mechanism\. External experience could first guide exploration, as a temporary training\-time scaffold, then be exploited by consolidating its useful behavioral effect into the experience\-free policy, and finally be retired once its utility vanishes\. Under this view, experience reuse becomes a scaffold\-to\-parameter learning problem rather than a retrieve\-and\-prompt augmentation problem\.

In this paper, we instantiate this idea in EDGE \(Experience\-Distillation for Guided Exploration\), a framework that turns retrieved experience from persistent inference\-time memory into dynamically validated training\-time scaffolding\. EDGE uses online marginal\-gain estimation to decide when an experience should guide exploration, be distilled into the policy, or be retired as the policy evolves, via three core mechanisms\. First,experience\-guided exploration scaffoldingpartitions each GRPO rollout group into experience\-conditioned and experience\-free trajectories, estimating the marginal gain of a retrieved experience without additional environment sampling and retaining only positive\-gain instances\. Then,gain\-gated privileged distillationserves as the exploitation step, transferring the behavior induced by the privileged context into the standard policy via reverse\-KL divergence computed on the student’s own empirical support, activating only when the estimated gain is positive to avoid negative transfer\. Finally,experience bank evolution and managementenables the experience bank to co\-evolve with the policy throughout training: new experiences are synthesized from failure\-mode analysis, while obsolete ones are pruned based on tracked utility, keeping the scaffold aligned with the agent’s evolving capability\.

On ALFWorld\([26](https://arxiv.org/html/2608.21946#bib.bib6)\)and WebShop\([37](https://arxiv.org/html/2608.21946#bib.bib5)\),EDGEconsistently outperforms skill\-free RL and prior experience\-augmented methods, with the largest gains on exploration\-intensive subtasks \(e\.g\., Heat, Cool, Pick2\)\. When external experiences are withheld at inference time,EDGEretains 96\.0% of its scaffolded performance—compared with 82\.9% for SkillRL and 92\.3% for EMPO2—confirming that useful exploration priors have been absorbed into the parametric policy\.

Our contributions are as follows:

- •We identify two failure modes of experience\-augmented agentic RL—policy\-dependent experience utility and persistent inference\-time retrieval dependence—and reframe experience reuse as a dynamic scaffold\-to\-parameter transition\.
- •We introduce experience\-guided exploration scaffolding, which partitions rollout groups to estimate the marginal value of each retrieved experience under the current policy without requiring additional environment sampling\.
- •We propose gain\-gated privileged distillation, which internalizes scaffold\-induced behavior via reverse\-KL on the student’s own empirical support, gated by the estimated marginal gain to prevent negative transfer, together with a co\-evolutionary experience bank that expands and prunes in response to the policy’s evolving needs\.
- •Experiments on ALFWorld and WebShop demonstrate thatEDGEimproves task performance and training efficiency while enabling robust deployment without external experience retrieval\.

## 2Preliminaries

We formalize the multi\-turn decision\-making process of an LLM\-based agent as a Partially Observable Markov Decision Process \(POMDP\)\([12](https://arxiv.org/html/2608.21946#bib.bib9)\)⟨𝒮,𝒜,𝒪,𝒯,ℛ⟩\\langle\\mathcal\{S\},\\mathcal\{A\},\\mathcal\{O\},\\mathcal\{T\},\\mathcal\{R\}\\rangle\. The observation space𝒪\\mathcal\{O\}consists of natural\-language strings emitted by the environment, and the action space𝒜\\mathcal\{A\}comprises token sequences generated autoregressively by the policyπθ\\pi\_\{\\theta\}\. Given a task instructionxx, the agent interacts with the environment over a sequence of turns: at stepttit receives observationot∈𝒪o\_\{t\}\\in\\mathcal\{O\}and produces actionat∈𝒜a\_\{t\}\\in\\mathcal\{A\}according to

at∼πθ\(⋅∣x,ht,ot\),a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,h\_\{t\},o\_\{t\}\),\(1\)whereht=\(o1,a1,…,ot−1,at−1\)h\_\{t\}=\(o\_\{1\},a\_\{1\},\\ldots,o\_\{t\-1\},a\_\{t\-1\}\)is the interaction history\. An episode terminates upon task completion or at a maximum step limit, yielding a sparse binary rewardR⁡\(τ\)∈\{0,1\}R\(\\tau\)\\in\\\{0,1\\\}and a full trajectory

τ=\(x,o1,a1,…,oT,aT,R⁡\(τ\)\)\.\\tau=\(x,\\;o\_\{1\},a\_\{1\},\\;\\ldots,\\;o\_\{T\},a\_\{T\},\\;R\(\\tau\)\)\.\(2\)
We build on Group Relative Policy Optimization \(GRPO\)\([24](https://arxiv.org/html/2608.21946#bib.bib7)\), which samples a group ofGGparallel trajectories\{τi\}i=1G\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{G\}per task and computes group\-normalized advantages:

A^i=R⁡\(τi\)−mean⁡\(\{R⁡\(τj\)\}j=1G\)std⁡\(\{R⁡\(τj\)\}j=1G\)\.\\hat\{A\}\_\{i\}=\\frac\{R\(\\tau\_\{i\}\)\-\\mathrm\{mean\}\(\\\{R\(\\tau\_\{j\}\)\\\}\_\{j=1\}^\{G\}\)\}\{\\mathrm\{std\}\(\\\{R\(\\tau\_\{j\}\)\\\}\_\{j=1\}^\{G\}\)\}\.\(3\)The policy is updated by maximizing a clipped surrogate objective with KL regularization:

ℒRL\(θ\)=−𝔼\[1G∑i=1G1\|τi\|∑t=1\|τi\|\(min\(ρi,tA^i,\\displaystyle\\mathcal\{L\}\_\{\\text\{RL\}\}\(\\theta\)=\-\\mathbb\{E\}\\\!\\Bigg\[\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\frac\{1\}\{\|\\tau\_\{i\}\|\}\\sum\_\{t=1\}^\{\|\\tau\_\{i\}\|\}\\Big\(\\min\\\!\\big\(\\rho\_\{i,t\}\\,\\hat\{A\}\_\{i\},\\;clip\(ρi,t,−ϵ,\+ϵ\)A^i\)−βDKL\(πθ∥πref\)\)\],\\displaystyle\\mathrm\{clip\}\(\\rho\_\{i,t\},1\\\!\-\\\!\\epsilon,1\\\!\+\\\!\\epsilon\)\\,\\hat\{A\}\_\{i\}\\big\)\-\\beta\\,D\_\{\\mathrm\{KL\}\}\\big\(\\pi\_\{\\theta\}\\\|\\pi\_\{\\text\{ref\}\}\\big\)\\Big\)\\Bigg\],\(4\)whereρi,t=πθ​\(at∣x,ht,ot\)/πθold​\(at∣x,ht,ot\)\\rho\_\{i,t\}=\\pi\_\{\\theta\}\(a\_\{t\}\\mid x,h\_\{t\},o\_\{t\}\)\\,/\\,\\pi\_\{\\theta\_\{\\text\{old\}\}\}\(a\_\{t\}\\mid x,h\_\{t\},o\_\{t\}\)is the importance sampling ratio,πref\\pi\_\{\\text\{ref\}\}is the reference policy, andβ\\betacontrols KL penalty strength\. By contrasting outcomes within each group, GRPO steersθ\\thetatoward successful action sequences without a learned value function\. However, in sparse\-reward, partially observable environments, unguided exploration often yields groups in which few or no trajectories succeed, rendering the advantage estimate uninformative—a limitation we address in the following section\.

## 3Method: EDGE

We presentEDGE\(Experience\-Distillation forGuidedExploration\), a framework that iteratively strengthens the multi\-turn reasoning capability of LLM\-based agents \(Figure[1](https://arxiv.org/html/2608.21946#S0.F1)\)\. Central to our approach is treating retrieved external experience not as a static inference\-time prompt—which inflates the context window and creates persistent retrieval dependence—but as a*training\-time scaffold*that guides exploration and is discarded at deployment\. The agent first explores with privileged access to experience, then distills only the empirically beneficial behavior into its own parameters, so the deployed policy requires no external scaffold\. Section[3\.1](https://arxiv.org/html/2608.21946#S3.SS1)introduces the experience\-guided exploration scaffolding, Section[3\.2](https://arxiv.org/html/2608.21946#S3.SS2)details the gain\-gated privileged distillation mechanism, and Section[3\.3](https://arxiv.org/html/2608.21946#S3.SS3)describes the experience bank evolution and management strategy\.

### 3\.1Experience\-Guided Exploration Scaffolding

To provide exploratory guidance in the sparse\-reward, partially observable environments typical of agentic tasks, EDGE introduces a controlled information asymmetry within the standard GRPO rollout group\. Given a task instructionxxand an experience bankℰ\\mathcal\{E\}, we retrieve the top\-mmmost relevant experiences by embedding similarity, using the task instruction and initial environmental observations as the query, and select the highest\-scoring entrye∈ℰe\\in\\mathcal\{E\}\. Recall that GRPO samples a group ofGGparallel trajectories per task to estimate relative advantages \(Eq\. \([3](https://arxiv.org/html/2608.21946#S2.E3)\)\)\. We partition this group into two equal subsets without adding extra rollouts:

- •Teacher rollouts\(𝒯𝖳\\mathcal\{T\}^\{\\mathsf\{T\}\},G/2G/2trajectories\): conditioned on the*privileged context*c𝖳=x⊕ec^\{\\mathsf\{T\}\}=x\\oplus e, where⊕\\oplusdenotes concatenation under a unified chat template\.
- •Student rollouts\(𝒯𝖲\\mathcal\{T\}^\{\\mathsf\{S\}\},G/2G/2trajectories\): conditioned on the*standard context*c𝖲=xc^\{\\mathsf\{S\}\}=xalone\.

Because both subsets share the same policyπθ\\pi\_\{\\theta\}and differ only in whether the retrieved experience is visible, any performance gap can be attributed to the informational advantage provided byee\. We quantify this gap via the*instantaneous marginal gain*:

Δe=1\|𝒯𝖳\|​∑τ∈𝒯𝖳R⁡\(τ\)−1\|𝒯𝖲\|​∑τ∈𝒯𝖲R⁡\(τ\),\\Delta\_\{e\}\\;=\\;\\frac\{1\}\{\|\\mathcal\{T\}^\{\\mathsf\{T\}\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}^\{\\mathsf\{T\}\}\}R\(\\tau\)\\;\-\\;\\frac\{1\}\{\|\\mathcal\{T\}^\{\\mathsf\{S\}\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}^\{\\mathsf\{S\}\}\}R\(\\tau\),\(5\)whereR⁡\(τ\)R\(\\tau\)is the binary outcome reward defined in §[2](https://arxiv.org/html/2608.21946#S2)\. A positiveΔe\\Delta\_\{e\}signals that the experience provides useful guidance beyond the agent’s current capability; a non\-positive value indicates it is redundant or harmful\. This estimate gates the distillation objective \(§[3\.2](https://arxiv.org/html/2608.21946#S3.SS2)\) and drives experience bank updates \(§[3\.3](https://arxiv.org/html/2608.21946#S3.SS3)\)\.

Beyond gating distillation,Δe\\Delta\_\{e\}controls which rollouts enter the RL loss itself\. WhenΔe\>0\\Delta\_\{e\}\>0, advantages \(Eq\. \([3](https://arxiv.org/html/2608.21946#S2.E3)\)\) are computed over allGGtrajectories\. The pooled baseline therefore reflects the performance level achievable under the scaffold, giving student rollouts a stronger calibrated reference than a within\-subset baseline would provide\. WhenΔe≤0\\Delta\_\{e\}\\leq 0, teacher\-conditioned trajectories are excluded from the RL loss entirely: advantages are computed over𝒯𝖲\\mathcal\{T\}^\{\\mathsf\{S\}\}alone, reducing the update to a standard GRPO step\. Without this masking, non\-beneficial teacher rollouts would still shift the group baseline and distort student advantages—contaminating the policy gradient even though distillation is gated off\.

### 3\.2Gain\-Gated Privileged Distillation

The scaffolding in §[3\.1](https://arxiv.org/html/2608.21946#S3.SS1)exposes the agent to successful reasoning patterns it may not discover on its own, but because the retrieved experience is unavailable at deployment, these gains remain ephemeral unless consolidated into the policy itself\. This motivates a complementary mechanism: distilling the scaffold\-induced improvements into the base policy parameters so that the agent can reproduce them from the standard context alone\.

Our approach realizes this through asymmetric self\-distillation\. Rather than relying on a separate, unconditionally superior teacher, we use the same policyπθ\\pi\_\{\\theta\}under two informational conditions: the teacher isπθ\\pi\_\{\\theta\}evaluated with the privileged contextc𝖳c^\{\\mathsf\{T\}\}\(gradients stopped\), while the student operates underc𝖲c^\{\\mathsf\{S\}\}\. The asymmetry therefore lies in the information available to each role, not in model capacity\. The gain gate𝕀⁡\(Δe\>0\)\\mathbb\{I\}\(\\Delta\_\{e\}\>0\)further ensures that distillation is triggered only when the scaffold yields a verified performance advantage, allowing the policy to selectively internalize reusable reasoning patterns while avoiding negative transfer\.

To guard against covariate shift, we construct the distillation target on the student’s own empirical support\. For a gain\-gated task instance \(Δe\>0\\Delta\_\{e\}\>0\) and a student trajectoryτ∈𝒯𝖲\\tau\\in\\mathcal\{T\}^\{\\mathsf\{S\}\}with token sequence\(y1,…,ym\)\(y\_\{1\},\\ldots,y\_\{m\}\)generated underc𝖲c^\{\\mathsf\{S\}\}, we perform a no\-gradient forward pass ofπθ\\pi\_\{\\theta\}over the same tokens conditioned onc𝖳c^\{\\mathsf\{T\}\}\. Because both passes share identical actions, the comparison isolates the informational advantage ofeewithout introducing out\-of\-distribution transitions\.

We adopt the reverse KL divergenceDKL\(πstudent∥πteacher\)D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\text\{student\}\}\\\|\\pi\_\{\\text\{teacher\}\}\)as the distillation objective\. Unlike the forward KL, which compels the student to cover the full teacher support and can induce mode\-covering artifacts, the reverse KL is mode\-seeking: it encourages the policy to concentrate on the most effective reasoning mode under the scaffold\. The token\-level scaffold internalization loss is:

ℒdistill​\(θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{distill\}\}\(\\theta\)=𝔼τ∼𝒯𝖲\[𝕀\(Δe\>0\)\\displaystyle=\\mathbb\{E\}\_\{\\tau\\sim\\mathcal\{T\}^\{\\mathsf\{S\}\}\}\\\!\\Bigg\[\\mathbb\{I\}\(\\Delta\_\{e\}\\\!\>\\\!0\)∑t=1m∑y∈𝒱πθ𝖲\(y\)δt\(y\)\],\\displaystyle\\qquad\\sum\_\{t=1\}^\{m\}\\sum\_\{y\\in\\mathcal\{V\}\}\\pi\_\{\\theta\}^\{\\mathsf\{S\}\}\(y\)\\,\\delta\_\{t\}\(y\)\\Bigg\],δt​\(y\)\\displaystyle\\delta\_\{t\}\(y\)=log⁡πθ​\(y∣c<t𝖲\)πsg⁡\(θ\)​\(y∣c<t𝖳\),\\displaystyle=\\log\\frac\{\\pi\_\{\\theta\}\(y\\mid c^\{\\mathsf\{S\}\}\_\{<t\}\)\}\{\\pi\_\{\\mathrm\{sg\}\(\\theta\)\}\(y\\mid c^\{\\mathsf\{T\}\}\_\{<t\}\)\},\(6\)whereπθ𝖲​\(y\)≜πθ​\(y∣c<t𝖲\)\\pi\_\{\\theta\}^\{\\mathsf\{S\}\}\(y\)\\triangleq\\pi\_\{\\theta\}\(y\\mid c^\{\\mathsf\{S\}\}\_\{<t\}\)for brevity,δt​\(y\)\\delta\_\{t\}\(y\)is the per\-token log\-ratio between the student and the privileged teacher, andsg⁡\(⋅\)\\mathrm\{sg\}\(\\cdot\)denotes stop\-gradient\. This loss is combined with the GRPO objective \(Eq\. \([2](https://arxiv.org/html/2608.21946#S2.Ex1)\)\) to form the joint actor loss:

ℒactor​\(θ\)=ℒRL​\(θ\)\+λ​ℒdistill​\(θ\),\\mathcal\{L\}\_\{\\text\{actor\}\}\(\\theta\)=\\mathcal\{L\}\_\{\\text\{RL\}\}\(\\theta\)\+\\lambda\\,\\mathcal\{L\}\_\{\\text\{distill\}\}\(\\theta\),\(7\)whereλ\\lambdacontrols the relative weight of scaffold internalization versus the RL signal\. Through this joint optimization, the deployed agent reproduces privileged reasoning without any external scaffold at test time\.

### 3\.3Experience Bank Evolution and Management

A static experience library cannot adequately support the agent throughout training: asπθ\\pi\_\{\\theta\}improves, it encounters situations where existing experiences provide insufficient guidance; conversely, previously beneficial experiences may become redundant once the policy has internalized the corresponding behavior, or harmful if they conflict with newly discovered strategies\. EDGE therefore treatsℰ\\mathcal\{E\}as a living repository that co\-evolves with the policy through both expansion and pruning\.

Following[34](https://arxiv.org/html/2608.21946#bib.bib18), we generate new experiences from the agent’s own rollout trajectories\. After each training step, we identify task categories whose success rate falls below a thresholdξ\\xiand collect representative failed and successful trajectories from these categories\. An external reflector LLM then analyzes the contrast between them to synthesize new experiential guidance:

enew=freflect​\(τ\+,τ−\),e\_\{\\text\{new\}\}=f\_\{\\text\{reflect\}\}\\\!\\left\(\\tau^\{\+\},\\,\\tau^\{\-\}\\right\),\(8\)whereτ\+\\tau^\{\+\}andτ−\\tau^\{\-\}denote a successful and a failed trajectory from the same task category, respectively\. The generated experiences are deduplicated against existing entries and inserted intoℰ\\mathcal\{E\}for retrieval in subsequent training iterations\.

Meanwhile, the utility of existing experiences is continuously tracked via an Exponential Moving Average \(EMA\) score for each experienceee:

Ue\(t\)=\(1−μ\)​Ue\(t−1\)\+μ​Δe,U\_\{e\}^\{\(t\)\}=\(1\-\\mu\)\\,U\_\{e\}^\{\(t\-1\)\}\+\\mu\\,\\Delta\_\{e\},\(9\)whereμ∈\(0,1\)\\mu\\in\(0,1\)is the momentum coefficient andΔe\\Delta\_\{e\}is the instantaneous marginal gain from Eq\. \([5](https://arxiv.org/html/2608.21946#S3.E5)\)\. The EMA smooths over stochastic fluctuations while remaining responsive to genuine shifts in experience utility\. Experiences whose tracked utilityUe\(t\)U\_\{e\}^\{\(t\)\}falls below a thresholdη\\etaare removed fromℰ\\mathcal\{E\}, retiring scaffolds that have been absorbed into the policy and reducing overhead during retrieval and rollout\. This yields a co\-evolutionary dynamic: the policy improves by internalizing useful experiences, exposing new failure modes that drive bank expansion, while obsolete experiences are simultaneously pruned away\.

TypeMethodALFWorldWebShopPickLookCleanHeatCoolPick2AllScoreSucc\.Closed\-Source ModelPromptingGPT\-4o75\.360\.831\.256\.721\.649\.848\.031\.823\.7PromptingGemini\-2\.5\-Pro92\.863\.362\.169\.026\.658\.760\.342\.535\.9Qwen2\.5\-1\.5B\-InstructPromptingBase Model5\.95\.53\.39\.74\.20\.04\.123\.15\.2PromptingReAct17\.420\.515\.76\.27\.72\.012\.840\.111\.3PromptingReflexion35\.322\.221\.713\.619\.43\.721\.855\.821\.9Post\-TrainingGRPO85\.353\.784\.578\.259\.753\.572\.875\.856\.8Post\-TrainingSkillRL91\.2±4\.364\.3±4\.678\.1±5\.476\.9±6\.370\.8±6\.554\.6±6\.174\.2±4\.777\.3±3\.560\.9±3\.7Post\-TrainingEMPO286\.9±3\.366\.2±5\.279\.3±4\.979\.1±4\.575\.3±5\.764\.8±6\.676\.8±3\.778\.2±3\.763\.2±3\.3Post\-TrainingEDGE80\.6±2\.273\.7±5\.576\.0±4\.387\.3±4\.085\.6±6\.168\.4±5\.679\.7±2\.378\.8±2\.165\.6±3\.2Qwen2\.5\-7B\-InstructPromptingBase Model33\.421\.619\.36\.92\.83\.214\.826\.47\.8PromptingReAct48\.535\.434\.313\.218\.217\.631\.246\.219\.5PromptingReflexion62\.041\.644\.930\.936\.323\.842\.758\.128\.8Post\-TrainingEvolveR†64\.933\.346\.413\.333\.333\.343\.842\.517\.6Post\-TrainingGRPO\+Mem0†78\.154\.856\.131\.065\.026\.954\.758\.137\.5Post\-TrainingOPSD†50\.060\.022\.721\.417\.69\.532\.84\.52\.3Post\-TrainingGRPO92\.885\.789\.375\.774\.567\.782\.180\.370\.1Post\-TrainingGRPO\+OPSD†91\.461\.510087\.576\.552\.280\.486\.876\.5Post\-TrainingSkillRL94\.1±2\.483\.3±4\.388\.4±3\.685\.6±5\.190\.2±5\.678\.6±4\.988\.2±3\.786\.2±3\.176\.7±3\.3Post\-TrainingEMPO293\.3±3\.788\.9±3\.991\.2±4\.488\.5±5\.389\.4±4\.679\.8±5\.189\.1±2\.888\.3±2\.677\.1±4\.0Post\-TrainingEDGE96\.0±1\.985\.1±4\.893\.4±3\.290\.0±2\.692\.9±4\.581\.6±6\.090\.4±2\.289\.6±2\.882\.6±3\.8

Table 1:Performance on ALFWorld and WebShop\. We report the average success rate \(%\) per subtask and overall for ALFWorld, and both the average score and success rate \(%\) for WebShop\. Results ofEDGEand other experience\-augmented training methods are averaged over 3 random seeds \(mean±\\pmstd\)\.†\\daggerdenotes results replicated from[34](https://arxiv.org/html/2608.21946#bib.bib18)and[14](https://arxiv.org/html/2608.21946#bib.bib43)\.

## 4Experiments

We evaluateEDGEto examine whether experience scaffolding can improve agentic RL while being progressively internalized by the policy\. Our experiments address four questions: \(1\) DoesEDGEimprove over prompting, vanilla RL, and prior experience\-augmented methods, particularly on exploration\-intensive tasks? \(2\) Does the trained policy retain its performance when external experiences are removed at inference time? \(3\) How much does each component—gain gating, distillation, and pruning—contribute, and how do they interact? \(4\) Is experience utility truly non\-stationary, and does the co\-evolutionary bank adapt accordingly during training?

### 4\.1Experiment Setup

#### Environments\.

We evaluate on two interactive environments with sparse outcome\-level feedback\.ALFWorld\([26](https://arxiv.org/html/2608.21946#bib.bib6)\)is a text\-based household environment with 3,827 task instances spanning six types: Pick & Place \(Pick\), Examine in Light \(Look\), Clean & Place \(Clean\), Heat & Place \(Heat\), Cool & Place \(Cool\), and Pick Two & Place \(Pick2\)\.WebShop\([37](https://arxiv.org/html/2608.21946#bib.bib5)\)simulates an e\-commerce website with over 1\.1M products and 12k human\-written shopping instructions, requiring agents to search, browse, and purchase products matching detailed user specifications\.

#### Baselines\.

We compare against three families of methods: closed\-source LLM agents \(GPT\-4o\([17](https://arxiv.org/html/2608.21946#bib.bib20)\)and Gemini\-2\.5\-Pro\([6](https://arxiv.org/html/2608.21946#bib.bib21)\)\), prompting agents \(ReAct\([38](https://arxiv.org/html/2608.21946#bib.bib12)\)and Reflexion\([25](https://arxiv.org/html/2608.21946#bib.bib13)\)\), and training\-based agents\. The post\-training baselines include GRPO\([24](https://arxiv.org/html/2608.21946#bib.bib7)\), OPSD\([46](https://arxiv.org/html/2608.21946#bib.bib30)\), EvolveR\([33](https://arxiv.org/html/2608.21946#bib.bib15)\), GRPO augmented with Mem0\([4](https://arxiv.org/html/2608.21946#bib.bib17)\), SkillRL\([34](https://arxiv.org/html/2608.21946#bib.bib18)\), and EMPO2\([13](https://arxiv.org/html/2608.21946#bib.bib19)\)\. Detailed descriptions are provided in Appendix[B\.1](https://arxiv.org/html/2608.21946#A2.SS1)\.

#### Training details\.

We use Qwen2\.5\-1\.5B/7B\-Instruct\([21](https://arxiv.org/html/2608.21946#bib.bib16)\)as our base models\. For ALFWorld and WebShop, all Post\-Training methods use exactly the same hyperparameter configurations\. The rollout group sizeGGfor group\-based RL methods is set to 8\. For experience retrieval, we encode experiences with Qwen3\-Embedding\-0\.6B\([44](https://arxiv.org/html/2608.21946#bib.bib42)\)and rank candidates by cosine similarity against the task instruction and initial observations\. For experience bank expansion, we use GPT\-4o\([17](https://arxiv.org/html/2608.21946#bib.bib20)\)as the reflector LLM to contrast successful and failed trajectories and synthesize new experiences\. Full training settings and hyperparameter details are provided in Appendix[B\.2](https://arxiv.org/html/2608.21946#A2.SS2)\.

### 4\.2Main Results

Table[1](https://arxiv.org/html/2608.21946#S3.T1)presents results on ALFWorld and WebShop across two model scales\. At 7B,EDGEachieves 90\.4% success rate on ALFWorld and 82\.6% on WebShop, improving over GRPO by 8\.3 and 12\.5 points and over the strongest experience\-augmented baseline EMPO2by 1\.3 and 5\.5 points, with consistent gains at 1\.5B\.

The subtask breakdown reveals thatEDGE’s advantage concentrates on exploration\-heavy tasks—Heat, Cool, and Pick2—which demand longer action sequences with fewer intermediate rewards\. On WebShop, the success\-rate improvement over GRPO \(\+12\.5\) exceeds the score improvement \(\+9\.3\), indicating thatEDGEhelps agents complete full decision chains rather than merely accumulate partial credit\.

These results highlight that experience is most effective as selective exploration guidance rather than unconditional augmentation\. Naive memory approaches \(EvolveR, GRPO\+Mem0\) degrade well below vanilla GRPO, and standalone self\-distillation \(OPSD\) provides insufficient signal without effective exploration\. Combining the two \(GRPO\+OPSD\) recovers competitive aggregate performance but remains brittle across subtasks, whereasEDGE’s marginal\-gain gating ensures only beneficial experiences contribute, yielding both higher overall success and more uniform subtask coverage\.

Figure 2:Performance retention after scaffold removal at inference time\.EDGEpreserves 96\.0% of its scaffolded performance without external experiences, compared with 82\.9% for SkillRL and 92\.3% for EMPO2, indicating more effective internalization into the parametric policy at the 7B scale\.Figure 3:Training dynamics ofEDGEvs\. GRPO on ALFWorld with Qwen2\.5\-7B\-Instruct\.EDGEachieves higher validation success while reducing both environment steps and trajectory length more rapidly, indicating that scaffolded exploration is progressively internalized into more efficient experience\-free behavior\.#### Inference without External Scaffolds\.

Figure[2](https://arxiv.org/html/2608.21946#S4.F2)poses a stricter test: how much performance survives when all external experiences are withheld at inference time?EDGEpreserves 96\.0% of its scaffolded performance, compared with 92\.3% for EMPO2and 82\.9% for SkillRL, confirming that the reverse\-KL distillation stage successfully transfers scaffold\-induced behavior into the parametric policy\. The residual 3\.6\-point gap suggests that a small fraction of experience\-conditioned exploration strategies resist distillation, a direction we leave to future work\.

### 4\.3Analysis

#### Ablation Studies\.

Table 2:Ablation resultson ALFWorld with Qwen2\.5\-7B\-Instruct\. We report success rate \(%\) per subtask and overall\.Table[2](https://arxiv.org/html/2608.21946#S4.T2)isolates each component on ALFWorld with Qwen2\.5\-7B\-Instruct\. The most striking finding is that removing the gain gate drops overall success to 72\.3%—9\.8 points below vanilla GRPO—revealing that unfiltered experience injection does not merely fail to help but actively harms the policy, particularly on exploration\-heavy subtasks where misleading guidance compounds over long horizons\. This result also addresses a natural concern about theG/2\+G/2G/2\{\+\}G/2rollout partition: since all other ablated variants retain the same split yet outperform GRPO, the partition itself does not dilute the RL signal; the degradation is attributable entirely to distilling low\-quality experiences\.

The remaining two components contribute complementary benefits\. Without distillation, performance falls to 83\.6%, only 1\.5 points above GRPO, with losses concentrated on the same exploration\-heavy subtasks—confirming that scaffolded exploration provides transient guidance but does not durably reshape the policy without explicit behavioral transfer\. Without experience pruning, performance declines more modestly to 86\.7% with losses spread evenly across subtasks, indicating that pruning acts as a maintenance mechanism that keeps the experience bank aligned with the policy’s evolving capability rather than targeting any specific failure mode\.

#### Training Dynamics\.

Figure[3](https://arxiv.org/html/2608.21946#S4.F3)reveals a clear divergence between EDGE and GRPO after approximately step 100: GRPO plateaus and begins to regress, whereas EDGE continues to improve steadily toward 90% validation success\. We attribute GRPO’s decline to an exploration–exploitation collapse: once the policy commits to locally successful strategies, it loses the diversity needed to solve the remaining hard tasks, and further optimization erodes earlier gains\. EDGE’s experience scaffolds counteract this by continually injecting structured exploration guidance, while gain gating ensures that this guidance remains beneficial as the policy strengthens\. The efficiency panels corroborate this interpretation—EDGE reduces both environment steps and trajectory length more rapidly than GRPO, indicating that the policy internalizes increasingly direct action sequences rather than relying on extended trial\-and\-error, consistent with the scaffold\-removal results in Figure[2](https://arxiv.org/html/2608.21946#S4.F2)\.

Figure 4:Experience gain dynamics during training\.
#### Experience Gain Tracking\.

Figure[4](https://arxiv.org/html/2608.21946#S4.F4)tracks the mean marginal gainΔe\\Delta\_\{e\}\(Eq\. \([5](https://arxiv.org/html/2608.21946#S3.E5)\)\) across training to test a key premise of §[3\.3](https://arxiv.org/html/2608.21946#S3.SS3): that experience utility is non\-stationary\. With a static bank, the initially positive gain decays and turns negative after roughly step 100—confirming that once\-useful experiences become actively harmful as the policy outgrows them, and explaining why removing gain gating in Table[2](https://arxiv.org/html/2608.21946#S4.T2)degrades performance below vanilla GRPO\. Experience evolution delays this decay, but unchecked bank growth \(reaching over 650 entries\) introduces retrieval noise that keeps the gain signal volatile\. The full co\-evolutionary configuration sustains the most stable positive gain: pruning not only curbs bank growth but causes the bank to shrink after step 100, indicating that the policy absorbs existing experiences faster than new failure modes generate replacements—a direct signature of successful internalization\. Further details on bank evolution dynamics appear in Appendix[C\.2](https://arxiv.org/html/2608.21946#A3.SS2)\.

## 5Related Work

#### Reinforcement Learning for LLM Agents\.

RL has become a standard post\-training paradigm for LLM agents in multi\-turn environments\([37](https://arxiv.org/html/2608.21946#bib.bib5);[26](https://arxiv.org/html/2608.21946#bib.bib6);[9](https://arxiv.org/html/2608.21946#bib.bib1);[31](https://arxiv.org/html/2608.21946#bib.bib2);[11](https://arxiv.org/html/2608.21946#bib.bib3)\), progressing from PPO\-based methods\([23](https://arxiv.org/html/2608.21946#bib.bib4)\)to critic\-free objectives such as GRPO\([24](https://arxiv.org/html/2608.21946#bib.bib7)\)and RLOO\([2](https://arxiv.org/html/2608.21946#bib.bib8)\), with further improvements in multi\-turn credit assignment through turn\-level or stepwise signals\([9](https://arxiv.org/html/2608.21946#bib.bib1);[32](https://arxiv.org/html/2608.21946#bib.bib10);[29](https://arxiv.org/html/2608.21946#bib.bib11);[43](https://arxiv.org/html/2608.21946#bib.bib24)\)\. However, the reusable exploration patterns within trajectories are still consumed once and discarded\.EDGEretains the group\-sampling backbone of GRPO but repurposes it to estimate which retrieved experiences currently improve exploration and to transfer their effects into the policy\.

#### Experience\-Augmented LLM Agents\.

External memory and experience reuse have been explored through prompting\-based reflections\([25](https://arxiv.org/html/2608.21946#bib.bib13);[45](https://arxiv.org/html/2608.21946#bib.bib14);[36](https://arxiv.org/html/2608.21946#bib.bib22);[8](https://arxiv.org/html/2608.21946#bib.bib28);[18](https://arxiv.org/html/2608.21946#bib.bib29)\)and persistent retrieval during interaction\([4](https://arxiv.org/html/2608.21946#bib.bib17);[34](https://arxiv.org/html/2608.21946#bib.bib18);[33](https://arxiv.org/html/2608.21946#bib.bib15);[16](https://arxiv.org/html/2608.21946#bib.bib25);[30](https://arxiv.org/html/2608.21946#bib.bib27);[41](https://arxiv.org/html/2608.21946#bib.bib26);[15](https://arxiv.org/html/2608.21946#bib.bib46)\)\. Recent methods selectively replay past reasoning traces\([35](https://arxiv.org/html/2608.21946#bib.bib38);[42](https://arxiv.org/html/2608.21946#bib.bib39);[20](https://arxiv.org/html/2608.21946#bib.bib44);[40](https://arxiv.org/html/2608.21946#bib.bib23)\), but primarily target single\-turn tasks without the multi\-turn, partially observable interaction loops of agentic settings\. EMPO2\([13](https://arxiv.org/html/2608.21946#bib.bib19)\), the closest prior work, encourages memory\-free behavior but selects experiences by static heuristics and distills off\-policy, risking distribution mismatch with the student’s own visitation\.EDGEmeasures experience utility online via instantaneous marginal gain and transfers scaffold\-induced behavior through on\-policy self\-distillation on the student’s empirical support\.

#### Privileged Information and Self\-Distillation\.

Leveraging training\-time information unavailable at deployment spans LUPI\([28](https://arxiv.org/html/2608.21946#bib.bib33)\), asymmetric actor\-critic methods\([19](https://arxiv.org/html/2608.21946#bib.bib37)\), and context distillation in LLMs\([27](https://arxiv.org/html/2608.21946#bib.bib34);[5](https://arxiv.org/html/2608.21946#bib.bib31);[13](https://arxiv.org/html/2608.21946#bib.bib19);[3](https://arxiv.org/html/2608.21946#bib.bib45)\), though these approaches are largely off\-policy and suffer from distribution mismatch with the student’s own visitation\. On\-policy distillation\([1](https://arxiv.org/html/2608.21946#bib.bib36);[39](https://arxiv.org/html/2608.21946#bib.bib35)\)addresses this by supervising the student on its own sequences, and on\-policy self\-distillation \(OPSD\)\([46](https://arxiv.org/html/2608.21946#bib.bib30);[10](https://arxiv.org/html/2608.21946#bib.bib32);[14](https://arxiv.org/html/2608.21946#bib.bib43)\)further removes the need for a separate teacher\.EDGEshares this structure but adds gain\-gating to activate distillation only under verified positive gain and co\-evolves the experience bank through utility\-driven expansion and pruning\.

## 6Conclusion

We presentedEDGE, a framework for using retrieved experience as a temporary training\-time scaffold rather than a persistent inference\-time dependency\. EDGE estimates the marginal utility of retrieved experiences under the current policy, admits only positive\-gain scaffolds, and distills their behavioral effect into the standard experience\-free policy\. Across ALFWorld and WebShop, EDGE improves over GRPO and prior experience\-augmented agents, with especially large gains on exploration\-intensive subtasks and strong retention after scaffold removal\. These results suggest a practical principle for agentic RL: external experience is most useful when it is dynamically validated, selectively applied, and ultimately internalized\.

## Limitations

WhileEDGErequires no extra environment rollouts, it introduces training\-time overhead for maintaining the experience bank, computing teacher–student comparisons, and invoking the reflector LLM\. The effectiveness of gain\-gating also depends on reward quality; as with all outcome\-based RL methods, noisy or misspecified rewards would reduce the reliability of the marginal\-gain estimates\. In terms of empirical scope, our evaluation covers two text\-based benchmarks with Qwen2\.5 models at the 1\.5B and 7B scales; broader validation across environments, model families, and larger scales remains future work\. More broadly,EDGEassumes experiences are discrete textual artifacts; extending the scaffold\-to\-parameter principle to latent memory representations is an open problem\.

## References

- R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. R\. Garea, M\. Geist, and O\. BachemOn\-policy distillation of language models: learning from self\-generated mistakes\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=3zKtaqxLhW)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Ahmadianet al\.\(2024\)A\. Ahmadian, C\. Cremer, M\. Gallé, M\. Fadaee, J\. Kreutzer, O\. Pietquin, A\. Üstün, and S\. HookerBack to basics: revisiting reinforce\-style optimization for learning from human feedback in llms\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 12248–12267\.External Links:[Link](https://doi.org/10.18653/v1/2024.acl-long.662),[Document](https://dx.doi.org/10.18653/V1/2024.ACL-LONG.662)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Chenget al\.\(2026\)Z\. Cheng, Z\. Liu, Y\. Shan, X\. Wang, X\. Zhu, Y\. Ma, H\. Wang, Y\. Guo, W\. Lin, and Y\. WangMem2\{\}^\{2\}evolve: towards self\-evolving agents via co\-evolutionary capability expansion and experience distillation\.External Links:2604\.10923,[Link](https://arxiv.org/abs/2604.10923)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Chhikaraet al\.\(2025\)P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready ai agents with scalable long\-term memory\.External Links:2504\.19413,[Link](https://arxiv.org/abs/2504.19413)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix9.p1.1.1),[§1](https://arxiv.org/html/2608.21946#S1.p2.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Choudhury and Sodhi \(2025\)S\. Choudhury and P\. SodhiBetter than your teacher: LLM agents that learn from privileged AI feedback\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,External Links:[Link](https://openreview.net/forum?id=st7XqFgbAH)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Comaniciet al\.\(2025\)G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen, L\. Marris, S\. Petulla, C\. Gaffney, A\. Aharoni, N\. Lintz, T\. C\. Pais, H\. Jacobsson, I\. Szpektor, N\. Jiang, K\. Haridasan, A\. Omran, N\. Saunshi, D\. Bahri, G\. Mishra, E\. Chu, T\. Boyd, B\. Hekman, A\. Parisi, C\. Zhang, K\. Kawintiranon, T\. Bedrax\-Weiss, O\. Wang, Y\. Xu, O\. Purkiss, U\. Mendlovic, I\. Deutel, N\. Nguyen, A\. Langley, F\. Korn, L\. Rossazza, A\. Ramé, S\. Waghmare, H\. Miller, N\. Byrd, A\. Sheshan, R\. Hadsell, S\. Bhardwaj, P\. Janus, T\. Rissa, D\. Horgan, A\. Abdagic, L\. Belenki, J\. Allingham, A\. Singh, T\. Guidroz, S\. Srinivasan, H\. Schmit, K\. Chiafullo, A\. Elisseeff, N\. Jha, P\. Kolhar, L\. Berrada, F\. Ding, X\. Si, S\. B\. Mallick, F\. Och, S\. Erell, E\. Ni, T\. Latkar, S\. Yang, P\. Sirkovic, Z\. Feng, R\. Leland, R\. Hornung, G\. Wu, C\. Blundell, H\. Alvari, P\. Huang, C\. Yip, S\. Deur, L\. Liu, G\. Surita, P\. Duque, D\. Damen, J\. Jia, A\. Guez, M\. Mircea, A\. Sinha, A\. Magni, P\. Stradomski, T\. Marian, V\. Galić, W\. Chen, H\. Husain, A\. Singhal, D\. Grewe, F\. Aubet, S\. Song, L\. Blanco, L\. Rechis, L\. Ho, R\. Munoz, K\. Zheng, J\. Hamrick, K\. Mather, H\. Taitelbaum, E\. Rutherford, Y\. Lei, K\. Chen, A\. Shukla, E\. Moreira, E\. Doi, B\. Isik, N\. Shabat, D\. Rogozińska, K\. Kolipaka, J\. Chang, E\. Vušak, S\. Venkatachary, S\. Noghabi, T\. Bharti, Y\. Jun, A\. Zaks, S\. Green, J\. Challagundla, W\. Wong, M\. Mohammad, D\. Hirsch, Y\. Cheng, I\. Naim, L\. Proleev, D\. Vincent, A\. Singh, M\. Krikun, D\. Krishnan, Z\. Ghahramani, A\. Atias, R\. Aggarwal, C\. Kirov, D\. Vytiniotis, C\. Koh, A\. Chronopoulou, P\. Dogra, V\. Ion, G\. Tyen, J\. Lee, F\. Weissenberger, T\. Strohman, A\. Balakrishna, J\. Rae, M\. Velic, R\. de Liedekerke, O\. Elyada, W\. Yuan, C\. Liu, L\. Shani, S\. Kishchenko, B\. Alessio, Y\. Li, R\. Song, S\. Kwei, O\. Jankowski, A\. Pappu, Y\. Namiki, Y\. Ma, N\. Tripuraneni, C\. Cherry, M\. Ikonomidis, Y\. Ling, C\. Ji, B\. Westberg, A\. Wright, D\. Yu, D\. Parkinson, S\. Ramaswamy, J\. Connor, S\. H\. Yeganeh, S\. Grover, G\. Kenwright, L\. Litchev, C\. Apps, A\. Tomala, F\. Halim, A\. Castro\-Ros, Z\. Li, A\. Boral, P\. Sho, M\. Yarom, E\. Malmi, D\. Klinghoffer, R\. Lin, A\. Ansell, P\. K\. S, S\. Zhao, S\. Zuo, A\. Santoro, H\. Cheng, S\. Demmessie, Y\. Liu, N\. Brichtova, A\. Culp, N\. Braun, D\. Graur, W\. Ng, N\. Mehta, A\. Phillips, P\. Sundberg, V\. Godbole, F\. Liu, Y\. Katariya, D\. Rim, M\. Seyedhosseini, S\. Ammirati, J\. Valfridsson, M\. Malihi, T\. Knight, A\. Toor, T\. Lampe, A\. Ittycheriah, L\. Chiang, C\. Yeung, A\. Fréchette, J\. Rao, H\. Wang, H\. Srivastava, R\. Zhang, R\. Rhodes, A\. Brand, D\. Weesner, I\. Figotin, F\. Gimeno, R\. Fellinger, P\. Marcenac, J\. Leal, E\. Marcus, V\. Cotruta, R\. Cabrera, S\. Luo, D\. Garrette, V\. Axelrod, S\. Baltateanu, D\. Barker, D\. Chen, H\. Toma, B\. Ingram, J\. Riesa, C\. Kulkarni, Y\. Zhang, H\. Liu, C\. Wang, M\. Polacek, W\. Wu, K\. Hui, A\. N\. Reyes, Y\. Su, M\. Barnes, I\. Malhi, A\. Siddiqui, Q\. Feng, M\. Damaschin, D\. Pighin, A\. Steiner, S\. Yang, R\. S\. Boppana, S\. Ivanov, A\. Kandoor, A\. Shah, A\. Mujika, D\. Huang, C\. A\. Choquette\-Choo, M\. Patel, T\. Yu, T\. Creswell, Jerry, Liu, C\. Barros, Y\. Razeghi, A\. Roy, P\. Culliton, B\. Xiong, J\. Pan, T\. Strohmann, T\. Powell, B\. Seal, D\. DeCarlo, P\. Shyam, K\. Katircioglu, X\. Wang, C\. Hardin, I\. Odisho, J\. Broder, O\. Chang, A\. Nair, A\. Shtefan, M\. O’Brien, M\. Agarwal, S\. Potluri, S\. Goyal, A\. Jhindal, S\. Thakur, Y\. Stuken, J\. Lyon, K\. Toutanova, F\. Feng, A\. Wu, B\. Horn, A\. Wang, A\. Cullum, G\. Taubman, D\. Shrivastava, C\. Shi, H\. Tomlinson, R\. Patel, T\. Tu, A\. M\. Oflazer, F\. Pongetti, M\. Yang, A\. A\. Taïga, V\. Perot, N\. W\. Pierse, F\. Han, Y\. Drori, I\. Iturrate, A\. Chakrabarti, L\. Yeung, D\. Dopson, Y\. Chen, A\. Kulshreshtha, T\. Guo, P\. Pham, T\. Schuster, J\. Chen, A\. Polozov, J\. Xing, H\. Zhou, P\. Kacham, D\. Kukliansky, A\. Miech, S\. Yaroshenko, E\. Chi, S\. Douglas, H\. Fei, M\. Blondel, P\. Myla, L\. Madmoni, X\. Wu, D\. Keysers, K\. Kjems, I\. Albuquerque, L\. Yu, J\. D’sa, M\. Plantan, V\. Ionescu, J\. S\. Elias, A\. Gupta, M\. R\. Vuyyuru, F\. Alcober, T\. Zhou, K\. Ji, F\. Hartmann, S\. Puttagunta, H\. Song, E\. Amid, A\. Stefanoiu, A\. Lee, P\. Pucciarelli, E\. Wang, A\. Raul, S\. Petrov, I\. Tian, V\. Anklin, N\. Nti, V\. Gomes, M\. Schumacher, G\. Vesom, A\. Panagopoulos, K\. Bousmalis, D\. Andor, J\. Jacob, Y\. Zhang, B\. Rosgen, M\. Kecman, M\. Tung, A\. Belias, N\. Goodman, P\. Covington, B\. Wieder, N\. Saxena, E\. Davoodi, M\. Huang, S\. Maddineni, V\. Roulet, F\. Campbell\-Ajala, P\. G\. Sessa, Xintian, Wu, G\. Lai, P\. Collins, A\. Haig, V\. Sakenas, X\. Xu, M\. Giustina, L\. E\. Shafey, P\. Charoenpanit, S\. Garg, J\. Ainslie, B\. Severson, M\. G\. Arenas, S\. Pathak, S\. Rajayogam, J\. Feng, M\. Bakker, S\. Li, N\. Wichers, J\. Rogers, X\. Geng, Y\. Li, R\. Jagerman, C\. Jia, N\. Olmert, D\. Sharon, M\. Mauger, S\. Mariserla, H\. Ma, M\. Mohabey, K\. Kim, A\. Andreev, S\. Pollom, J\. Love, V\. Jain, P\. Agrawal, Y\. Schroecker, A\. Fortin, M\. Warmuth, J\. Liu, A\. Leach, I\. Blok, G\. P\. Girirajan, R\. Aharoni, B\. Uria, A\. Sozanschi, D\. Goldberg, L\. Ionita, M\. T\. Ribeiro, M\. Zlocha, V\. Birodkar, S\. Lachgar, L\. Yuan, H\. Choudhury, M\. Ginsberg, F\. Zheng, G\. Dibb, E\. Graves, S\. Lokhande, G\. Rasskin, G\. Muraru, C\. Quick, S\. Tata, P\. Sermanet, A\. Chawla, I\. Karo, Y\. Wang, S\. Zhang, O\. Keller, A\. Dragan, G\. Su, I\. Chou, X\. Liu, Y\. Tao, S\. Prabhakara, M\. Wilson, R\. Liu, S\. Wang, G\. Evans, D\. Du, A\. Castaño, G\. Prasad, M\. E\. Mahdy, S\. Gerlach, M\. Reid, J\. Kahn, A\. Zait, T\. S\. Pillai, T\. Ulrich, G\. Wang, J\. Wassenberg, E\. Farkash, K\. Yalasangi, C\. Wang, M\. Bauza, S\. Bucher, T\. Liu, J\. Yan, G\. Leung, V\. Sindhwani, P\. Barnes, A\. Singh, I\. Jurin, J\. Chang, N\. K\. Bhumihar, S\. Eiger, G\. Citovsky, B\. Withbroe, Z\. Li, S\. Xue, N\. D\. Santo, G\. Stoyanov, Y\. Raimond, S\. Zheng, Y\. Gao, V\. Listík, S\. Kwasiborski, R\. Saputro, A\. Ozturel, G\. Mallya, K\. Majmundar, R\. West, P\. Caron, J\. Wei, L\. Castrejon, S\. Vikram, D\. Ramachandran, N\. Dhawan, J\. Park, S\. Smoot, G\. van den Driessche, Y\. Blau, C\. Malik, W\. Liang, R\. Hirsch, C\. N\. dos Santos, E\. Weinstein, A\. van den Oord, S\. Lall, N\. FitzGerald, Z\. Jiang, X\. Yang, D\. Webster, A\. Elqursh, A\. Pope, G\. Rotival, D\. Raposo, W\. Zhu, J\. Dean, S\. Alabed, D\. Tran, A\. Gupta, Z\. Gleicher, J\. Austin, E\. Rosseel, M\. Umekar, D\. Das, Y\. Sun, K\. Chen, K\. Misiunas, X\. Zhou, Y\. Di, A\. Loo, J\. Newlan, B\. Li, V\. Ramasesh, Y\. Xu, A\. Chen, S\. Gandhe, R\. Soricut, N\. Gupta, S\. Hu, S\. El\-Sayed, X\. Garcia, I\. Brusilovsky, P\. Chen, A\. Bolt, L\. Huang, A\. Gurney, Z\. Zhang, A\. Pritzel, J\. Wilkiewicz, B\. Seybold, B\. K\. Shamanna, F\. Fischer, J\. Dean, K\. Gill, R\. Mcilroy, A\. Bhowmick, J\. Selier, A\. Yang, D\. Cheng, V\. Magay, J\. Tan, D\. Varma, C\. Walder, T\. Kocisky, R\. Nakashima, P\. Natsev, M\. Kwong, I\. Gog, C\. Zhang, S\. Dieleman, T\. Jimma, A\. Ryabtsev, S\. Brahma, D\. Steiner, D\. Du, A\. Žužul, M\. Žanić, M\. Raghavachari, W\. Gierke, Z\. Zheng, D\. Petrova, Y\. Dauphin, Y\. Liu, I\. Kessler, S\. Hand, C\. Duvarney, S\. Kim, H\. Lee, L\. Hussenot, J\. Hui, J\. Smith, D\. Jain, J\. Xia, G\. S\. Tomar, K\. Amiri, D\. Phan, F\. Fuchs, T\. Weyand, N\. Tomasev, A\. Cordell, X\. Liu, J\. Mallinson, P\. Joshi, A\. Crawford, A\. Suggala, S\. Chien, N\. Fernando, M\. Sanchez\-Vargas, D\. Williams, P\. Crone, X\. Luo, I\. Karpov, J\. Shan, T\. Thurk, R\. Strudel, P\. Voigtlaender, P\. Patil, T\. Dozat, A\. Khodaei, S\. Singla, P\. Ambroszczyk, Q\. Wu, Y\. Chang, B\. Roark, C\. Hegde, T\. Ding, A\. Filos, Z\. Wu, A\. S\. Pinto, S\. Liu, S\. Khanna, A\. Pandey, S\. Mcloughlin, Q\. Li, S\. Haves, A\. Zhou, E\. Buchatskaya, I\. Leal, P\. de Boursac, N\. Akazawa, N\. Anderson, T\. Chen, K\. Somandepalli, C\. Liang, S\. Goenka, S\. Winkler, A\. Grushetsky, Y\. Ding, J\. Smith, F\. Ye, J\. Pont\-Tuset, E\. Li, R\. Li, T\. Golany, D\. Wegner, T\. Jiang, O\. Barak, Y\. Shangguan, E\. Vértes, R\. Wong, J\. Bornschein, A\. Tudor, M\. Bevilacqua, T\. Schaul, A\. S\. Rawat, Y\. Zhao, K\. Axiotis, L\. Meng, C\. McLean, J\. Lai, J\. Beattie, N\. Kushman, Y\. Liu, B\. Kutzman, F\. Lang, J\. Ye, P\. Netrapalli, P\. Mishra, M\. Khan, M\. Goel, R\. Willoughby, D\. Tian, H\. Zhuang, J\. Chen, Z\. Tsai, T\. Kementsietsidis, A\. Khare, J\. Keeling, K\. Xu, N\. Waters, F\. Altché, A\. Popat, B\. Mittal, D\. Saxton, D\. E\. Badawy, M\. Mathieu, Z\. Zheng, H\. Zhou, N\. Ranka, R\. Shin, Q\. Duan, T\. Salimans, I\. Mihailescu, U\. Shaham, M\. Chang, Y\. Assael, N\. Dikkala, M\. Izzard, V\. Cohen\-Addad, C\. Graves, V\. Feinberg, G\. Chung, D\. Strouse, D\. Karmon, S\. Sharifzadeh, Z\. Ashwood, K\. Pham, J\. Blanton, A\. Vasiloff, J\. Barber, M\. Geller, A\. Zhou, F\. Zubach, T\. Huang, L\. Zhang, H\. Gupta, M\. Young, J\. Proskurnia, R\. Votel, V\. Gabeur, G\. Barcik, A\. Tripathi, H\. Yu, G\. Yan, B\. Changpinyo, F\. Pavetić, A\. Coyle, Y\. Fujii, J\. G\. Mendez, T\. Zhou, H\. Rajamani, B\. Hechtman, E\. Cao, D\. Juan, Y\. Tan, V\. Dalibard, Y\. Du, N\. Clay, K\. Yao, W\. Jia, D\. Vijaykumar, Y\. Zhou, X\. Bai, W\. Hung, S\. Pecht, G\. Todorov, N\. Khadke, P\. Gupta, P\. Lahoti, A\. Autef, K\. Duddu, J\. Lee\-Thorp, A\. Bykovsky, T\. Misiunas, S\. Flennerhag, S\. Thangaraj, J\. McGiffin, Z\. Nado, M\. Kunesch, A\. Noever, A\. Hertz, M\. Liang, V\. Stone, E\. Palmer, S\. Daruki, A\. Pramanik, S\. Põder, A\. Kyker, M\. Khan, E\. Sluzhaev, M\. Ritter, A\. Ruderman, W\. Zhou, C\. Nagpal, K\. Vodrahalli, G\. Necula, P\. Barham, E\. Pavlick, J\. Hartford, I\. Shafran, L\. Zhao, M\. Mikuła, T\. Eccles, H\. Shimokawa, K\. Garg, L\. Vilnis, H\. Chen, I\. Shumailov, K\. Lee, A\. Abdelhamed, M\. Xie, V\. Cohen, E\. Hlavnova, D\. Malkin, C\. Sitawarin, J\. Lottes, P\. Coquinot, T\. Yu, S\. Kumar, J\. Zhang, A\. Mahendru, Z\. Ahmed, J\. Martens, T\. Chen, A\. Boag, D\. Peng, C\. Devin, A\. Klimovskiy, M\. Phuong, D\. Vainstein, J\. Xie, B\. Ramabhadran, N\. Howard, X\. Yu, G\. Goswami, J\. Cui, S\. Shleifer, M\. Pinto, C\. Yeh, M\. Yang, S\. Javanmardi, D\. Ethier, C\. Lee, J\. Orbay, S\. Kotecha, C\. Bromberg, P\. Shaw, J\. Thornton, A\. G\. Rosenthal, S\. Gu, M\. Thomas, I\. Gemp, A\. Ayyar, A\. Ushio, A\. Selvan, J\. Wee, C\. Liu, M\. Majzoubi, W\. Yu, J\. Abernethy, T\. Liechty, R\. Pan, H\. Nguyen, Qiong, Hu, S\. Perrin, A\. Arora, E\. Pitler, W\. Wang, K\. Shivakumar, F\. Prost, B\. Limonchik, J\. Wang, Y\. Gao, T\. Cour, S\. Buch, H\. Gui, M\. Ivanova, P\. Neubeck, K\. Chan, L\. Kim, H\. Chen, N\. Goyal, D\. Chung, L\. Liu, Y\. Su, A\. Petrushkina, J\. Shen, A\. Joulin, Y\. Xu, S\. X\. Lin, Y\. Kulizhskaya, C\. Chelba, S\. Vasudevan, E\. Collins, V\. Bashlovkina, T\. Lu, D\. Fritz, J\. Park, Y\. Zhou, C\. Su, R\. Tanburn, M\. Sushkov, M\. Rasquinha, J\. Li, J\. Prendki, Y\. Li, P\. LV, S\. Sharma, H\. Fitoussi, H\. Huang, A\. Dai, P\. Dao, M\. Burrows, H\. Prior, D\. Qin, G\. Pundak, L\. L\. Sjoesund, A\. Khurshudov, Z\. Zhu, A\. Webson, E\. Kemp, T\. Tan, S\. Agrawal, S\. Sargsyan, L\. Cheng, J\. Stephan, T\. Kwiatkowski, D\. Reid, A\. Byravan, A\. H\. Michaely, N\. Heess, L\. Zhou, S\. Goenka, V\. Carpenter, A\. Levskaya, B\. Wang, R\. Roberts, R\. Leblond, S\. Chikkerur, S\. Ginzburg, M\. Chang, R\. Riachi, Chuqiao, Xu, Z\. Borsos, M\. Pliskin, J\. Pawar, M\. Lustman, H\. Kirkwood, A\. Anand, A\. Chaudhary, N\. Kalb, K\. Milan, S\. Augenstein, A\. Goldie, L\. Prince, K\. Raman, Y\. Sun, V\. Xia, A\. Cohen, Z\. Huo, J\. Camp, S\. Ellis, L\. Zilka, D\. V\. Torres, L\. Patel, S\. Arora, B\. Chan, J\. Adler, K\. Ayoub, J\. Liang, F\. Jamil, J\. Jiang, S\. Baumgartner, H\. Sun, Y\. Karov, Y\. Akulov, H\. Zheng, I\. Cai, C\. Fantacci, J\. Rubin, A\. R\. Acha, M\. Wang, N\. D’Souza, R\. Sathyanarayana, S\. Dai, S\. Rowe, A\. Simanovsky, O\. Goldman, Y\. Kuang, X\. Pan, A\. Rosenberg, T\. Rojas\-Esponda, P\. Dutta, A\. Zeng, I\. Jurenka, G\. Farquhar, Y\. Bansal, S\. Iqbal, B\. Roelofs, G\. Joung, P\. Beak, C\. Ryu, R\. Poplin, Y\. Wu, J\. Alayrac, S\. Buthpitiya, O\. Ronneberger, C\. Habtegebriel, W\. Li, P\. Cavallaro, A\. Wei, G\. Bensky, T\. Denk, H\. Ganapathy, J\. Stanway, P\. Joshi, F\. Bertolini, J\. Lo, O\. Ma, Z\. Charles, G\. Sampemane, H\. Sahni, X\. Chen, H\. Askham, D\. Gaddy, P\. Young, J\. Tan, M\. Eyal, A\. Bražinskas, L\. Zhong, Z\. Wu, M\. Epstein, K\. Bailey, A\. Hard, K\. Lee, S\. Goldshtein, A\. Ruiz, M\. Badawi, M\. Lochbrunner, J\. Kearns, A\. Brown, F\. Pardo, T\. Weber, H\. Yang, P\. Jiang, B\. Akin, Z\. Fu, M\. Wainwright, C\. Zou, M\. Gaba, P\. Manzagol, W\. Kan, Y\. Song, K\. Zainullina, R\. Lin, J\. Ko, S\. Deshmukh, A\. Jindal, J\. Svensson, D\. Tyam, H\. Zhao, C\. Kaeser\-Chen, S\. Baird, P\. Moradi, J\. Hall, Q\. Guo, V\. Tsang, B\. Liang, F\. Pereira, S\. Ganesh, I\. Korotkov, J\. Adamek, S\. Thiagarajan, V\. Tran, C\. Chen, C\. Tar, S\. Jain, I\. Dasgupta, T\. Bilal, D\. Reitter, K\. Zhao, G\. Vezzani, Y\. Gehman, P\. Mehta, L\. Beltrone, X\. Dotiwalla, S\. Guadarrama, Z\. Abbas, S\. Karp, P\. Georgiev, C\. Ferng, M\. Brockschmidt, L\. Peng, C\. Hirnschall, V\. Verma, Y\. Bi, Y\. Xiao, A\. Dabush, K\. Xu, P\. Wallis, R\. Parker, Q\. Wang, Y\. Xu, I\. Safarli, D\. Tewari, Y\. Zhang, S\. Kim, A\. Gesmundo, M\. Thomas, S\. Levi, A\. Chowdhury, K\. Rao, P\. Garst, S\. Conway\-Rahman, H\. Ran, K\. McKinney, Z\. Xiao, W\. Yu, R\. Agrawal, A\. Stjerngren, C\. Ionescu, J\. Chen, V\. Sharma, J\. Chiu, F\. Liu, K\. Franko, C\. Sanford, X\. Cai, P\. Michel, S\. Ganapathy, J\. Labanowski, Z\. Garrett, B\. Vargas, S\. Sun, B\. Gale, T\. Buschmann, G\. Desjardins, N\. Ghelani, P\. Jain, M\. Verma, C\. Asawaroengchai, J\. Eisenschlos, J\. Harlalka, H\. Kazawa, D\. Metzler, J\. Howland, Y\. Jian, J\. Ades, V\. Shah, T\. Gangwani, S\. Lee, R\. Ring, S\. M\. Hernandez, D\. Reich, A\. Sinha, A\. Sathe, J\. Kovac, A\. Gill, A\. Kannan, A\. D’olimpio, M\. Sevenich, J\. Whang, B\. Kim, K\. C\. Sim, J\. Chen, J\. Zhang, S\. Lall, Y\. Matias, B\. Jia, A\. Friesen, S\. Nasso, A\. Thapliyal, B\. Perozzi, T\. Yu, A\. Shekhawat, S\. Huda, P\. Grabowski, E\. Wang, A\. Sreevatsa, H\. Dib, M\. Hassen, P\. Schuh, V\. Milutinovic, C\. Welty, M\. Quinn, A\. Shah, B\. Wang, G\. Barth\-Maron, J\. Frye, N\. Axelsson, T\. Zhu, Y\. Ma, I\. Giannoumis, H\. Sedghi, C\. Ye, Y\. Luan, K\. Aydin, B\. Chandra, V\. Sampathkumar, R\. Huang, V\. Lavrenko, A\. Eleryan, Z\. Hong, S\. Hansen, S\. M\. Carthy, B\. Samanta, D\. Ćevid, X\. Wang, F\. Li, M\. Voznesensky, M\. Hoffman, A\. Terzis, V\. Sehwag, G\. Fidel, L\. He, M\. Cai, Y\. He, A\. Feng, M\. Nikoltchev, S\. Phatale, J\. Chase, R\. Lawton, M\. Zhang, T\. Ouyang, M\. Tragut, M\. H\. Manshadi, A\. Narayanan, J\. Shen, X\. Gao, T\. Bolukbasi, N\. Roy, X\. Li, D\. Golovin, L\. Panait, Z\. Qin, G\. Han, T\. Anthony, S\. Kudugunta, V\. Patraucean, A\. Ray, X\. Chen, X\. Yang, T\. Bhatia, P\. Talluri, A\. Morris, A\. Ražnatović, B\. Brownfield, J\. An, S\. Peng, P\. Kane, C\. Zheng, N\. Duduta, J\. Kessinger, J\. Noraky, S\. Liu, K\. Rong, P\. Veličković, K\. Rush, A\. Goldin, F\. Wei, S\. M\. R\. Garlapati, C\. Pantofaru, O\. Kwon, J\. Ni, E\. Noland, J\. D\. Trapani, F\. Beaufays, A\. G\. Roy, Y\. Chow, A\. Turker, G\. Cideron, L\. Mei, J\. Clark, Q\. Dou, M\. Bošnjak, R\. Leith, Y\. Du, A\. Yazdanbakhsh, M\. Nasr, C\. Kwak, S\. S\. Sheth, A\. Kaskasoli, A\. Anand, B\. Lakshminarayanan, S\. Jerome, D\. Bieber, C\. Chu, A\. Senges, T\. Shen, M\. Sridhar, N\. Ndebele, B\. Beyret, S\. Mohamed, M\. Chen, M\. Freitag, J\. Guo, L\. Liu, P\. Roit, H\. Chen, S\. Yan, T\. Stone, J\. Co\-Reyes, J\. Cole, S\. Scellato, S\. Azizi, H\. Hashemi, A\. Jin, A\. Iyer, M\. Valentine, A\. György, A\. Ahuja, D\. H\. Diaz, C\. Lee, N\. Clement, W\. Kong, D\. Garmon, I\. Watts, K\. Bhatia, K\. Gupta, M\. Miecnikowski, H\. Vallet, A\. Taly, E\. Loper, S\. Joshi, J\. Atwood, J\. Chick, M\. Collier, F\. Iliopoulos, R\. Trostle, B\. Gunel, R\. Leal\-Cavazos, A\. M\. Hrafnkelsson, M\. Guzman, X\. Ju, A\. Forbes, J\. Emond, K\. Chauhan, B\. Caine, L\. Xiao, W\. Zeng, A\. Moufarek, D\. Murphy, M\. Meng, N\. Gupta, F\. Riedel, A\. Das, E\. Lawal, S\. Narayan, T\. Sosea, J\. Swirhun, L\. Friso, B\. Neyshabur, J\. Lu, S\. Girgin, M\. Wunder, E\. Yvinec, A\. Pyne, V\. Carbune, S\. Rijhwani, Y\. Guo, T\. Doshi, A\. Briukhov, M\. Bain, A\. Hitron, X\. Wang, A\. Gupta, K\. Chen, C\. Du, W\. Zhang, D\. Shah, A\. Akula, M\. Dylla, A\. Kachra, W\. Kuo, T\. Zou, L\. Wang, L\. Xu, J\. Zhu, J\. Snyder, S\. Menon, O\. Firat, I\. Mordatch, Y\. Yuan, N\. Ponomareva, R\. Blevins, L\. Moore, W\. Wang, P\. Chen, M\. Scholz, A\. Dwornik, J\. Lin, S\. Li, D\. Antognini, T\. I, X\. Song, M\. Miller, U\. Kalra, A\. Raveret, O\. Akerlund, F\. Wu, A\. Nystrom, N\. Godbole, T\. Liu, H\. DeBalsi, J\. Zhao, B\. Liu, A\. Caciularu, L\. Lax, U\. Khandelwal, V\. Langston, E\. Bailey, S\. Lattanzi, Y\. Wang, N\. Kovelamudi, S\. Mondal, G\. Guruganesh, N\. Hua, O\. Roval, P\. Wesołowski, R\. Ingale, J\. Halcrow, T\. Sohn, C\. Angermueller, B\. Raad, E\. Stickgold, E\. Lu, A\. Kosik, J\. Xie, T\. Lillicrap, A\. Huang, L\. L\. Zhang, D\. Paulus, C\. Farabet, A\. Wertheim, B\. Wang, R\. Joshi, C\. Ko, Y\. Wu, S\. Agrawal, L\. Lin, X\. Sheng, P\. Sung, T\. Breland\-King, C\. Butterfield, S\. Gawde, S\. Singh, Q\. Zhang, R\. Apte, S\. Shetty, A\. Hutter, T\. Li, E\. Salesky, F\. Lebron, J\. Kanerva, M\. Paganini, A\. Nguyen, R\. Vallu, J\. Peter, S\. Velury, D\. Kao, J\. Hoover, A\. Bortsova, C\. Bishop, S\. Jakobovits, A\. Agostini, A\. Agarwal, C\. Liu, C\. Kwong, S\. Tavakkol, I\. Bica, A\. Greve, A\. GP, J\. Marcus, L\. Hou, T\. Duerig, R\. Moroshko, D\. Lacey, A\. Davis, J\. Amelot, G\. Wang, F\. Kim, T\. Strinopoulos, H\. Wan, C\. L\. Lan, S\. Krishnan, H\. Tang, P\. Humphreys, J\. Bai, I\. H\. Shtacher, D\. Machado, C\. Pang, K\. Burke, D\. Liu, R\. Aravamudhan, Y\. Song, E\. Hirst, A\. Singh, B\. Jou, L\. Bai, F\. Piccinno, C\. K\. Fu, R\. Alazard, B\. Meiri, D\. Winter, C\. Chen, M\. Zhang, J\. Heitkaemper, J\. Lambert, J\. Lee, A\. Frömmgen, S\. Rogulenko, P\. Nair, P\. Niemczyk, A\. Bulyenov, B\. Xu, H\. Shemtov, M\. Zadimoghaddam, S\. Toropov, M\. Wirth, H\. Dai, S\. Gollapudi, D\. Zheng, A\. Kurakin, C\. Lee, K\. Bullard, N\. Serrano, I\. Balazevic, Y\. Li, J\. Schalkwyk, M\. Murphy, M\. Zhang, K\. Sequeira, R\. Datta, N\. Agrawal, C\. Sutton, N\. Attaluri, M\. Chiang, W\. Farhan, G\. Thornton, K\. Lin, T\. Choma, H\. Nguyen, K\. Dasgupta, D\. Robinson, I\. Comşa, M\. Riley, A\. Pillai, B\. Mustafa, B\. Golan, A\. Zandieh, J\. Lespiau, B\. Porter, D\. Ross, S\. Rajayogam, M\. Agarwal, S\. Venugopalan, B\. Shahriari, Q\. Yan, H\. Xu, T\. Tobin, P\. Dubov, H\. Shi, A\. Recasens, A\. Kovsharov, S\. Borgeaud, L\. Dery, S\. Vasanth, E\. Gribovskaya, L\. Qiu, M\. Mahdieh, W\. Skut, E\. Nielsen, C\. Zheng, A\. Yu, C\. G\. Bostock, S\. Gupta, A\. Archer, C\. Rawles, E\. Davies, A\. Svyatkovskiy, T\. Tsai, Y\. Halpern, C\. Reisswig, B\. Wydrowski, B\. Chang, J\. Puigcerver, M\. H\. Taege, J\. Li, E\. Schnider, X\. Li, D\. Dena, Y\. Xu, U\. Telang, T\. Shi, H\. Zen, K\. Kastner, Y\. Ko, N\. Subramaniam, A\. Kumar, P\. Blois, Z\. Dai, J\. Wieting, Y\. Lu, Y\. Zeldes, T\. Xie, A\. Hauth, A\. Ţifrea, Y\. Li, S\. El\-Husseini, D\. Abolafia, H\. Zhou, W\. Ding, S\. Ghalebikesabi, C\. Guía, A\. Maksai, Á\. Weisz, S\. Arik, N\. Sukhanov, A\. Świetlik, X\. Jia, L\. Yu, W\. Wang, M\. Brand, D\. Bloxwich, S\. Kirmani, Z\. Chen, A\. Go, P\. Sprechmann, N\. Kannen, A\. Carin, P\. Sandhu, I\. Edkins, L\. Nooteboom, J\. Gupta, L\. Maggiore, J\. Azizi, Y\. Pritch, P\. Yin, M\. Gupta, D\. Tarlow, D\. Smith, D\. Ivanov, M\. Babaeizadeh, A\. Goel, S\. Kambala, G\. Chu, M\. Kastelic, M\. Liu, H\. Soltau, A\. Stone, S\. Agrawal, M\. Kim, K\. Soparkar, S\. Tadepalli, O\. Bunyan, R\. Soh, A\. Kannan, D\. Kim, B\. J\. Chen, A\. Halumi, S\. Roy, Y\. Wang, O\. Sercinoglu, G\. Gibson, S\. Bhatnagar, M\. Sano, D\. von Dincklage, Q\. Ren, B\. Mitrevski, M\. Olšák, J\. She, C\. Doersch, Jilei, Wang, B\. Liu, Q\. Tan, T\. Yakar, T\. Warkentin, A\. Ramirez, C\. Lebsack, J\. Dillon, R\. Mathews, T\. Cobley, Z\. Wu, Z\. Chen, J\. Simon, S\. Nath, T\. Sainath, A\. Bendebury, R\. Julian, B\. Mankalale, D\. Ćurko, P\. Zacchello, A\. R\. Brown, K\. Sodhia, H\. Howard, S\. Caelles, A\. Gupta, G\. Evans, A\. Bulanova, L\. Katzen, R\. Goldenberg, A\. Tsitsulin, J\. Stanton, B\. Schillings, V\. Kovalev, C\. Fry, R\. Shah, K\. Lin, S\. Upadhyay, C\. Li, S\. Radpour, M\. Maggioni, J\. Xiong, L\. Haas, J\. Brennan, A\. Kamath, N\. Savinov, A\. Nagrani, T\. Yacovone, R\. Kappedal, K\. Andriopoulos, L\. Lao, Y\. Li, G\. Rozhdestvenskiy, K\. Hashimoto, A\. Audibert, S\. Austin, D\. Rodriguez, A\. Ruoss, G\. Honke, D\. Karkhanis, X\. Xiong, Q\. Wei, J\. Huang, Z\. Leng, V\. Premachandran, S\. Bileschi, G\. Evangelopoulos, T\. Mensink, J\. Pavagadhi, D\. Teplyashin, P\. Chang, L\. Xue, G\. Tanzer, S\. Goldman, K\. Patel, S\. Li, J\. Wiesner, I\. Zheng, I\. Stewart\-Binks, J\. Han, Z\. Li, L\. Luo, K\. Lenc, M\. Lučić, F\. Xue, R\. Mullins, A\. Guseynov, C\. Chang, I\. Galatzer\-Levy, A\. Zhang, G\. Bingham, G\. Hu, A\. Hartman, Y\. Ma, J\. Griffith, A\. Irpan, C\. Radebaugh, S\. Yue, L\. Fan, V\. Ungureanu, C\. Sorokin, H\. Teufel, P\. Li, R\. Anil, D\. Paparas, T\. Wang, C\. Lin, H\. Peng, M\. Shum, G\. Petrovic, D\. Brady, R\. Nguyen, K\. Macherey, Z\. Li, H\. Singh, M\. Yenugula, M\. Iinuma, X\. Chen, K\. Kopparapu, A\. Stern, S\. Dave, C\. Thekkath, F\. Perot, A\. Kumar, F\. Li, Y\. Xiao, M\. Bilotti, M\. H\. Bateni, I\. Noble, L\. Lee, A\. Vázquez\-Reina, J\. Salazar, X\. Yang, B\. Wang, E\. Gruzewska, A\. Rao, S\. Raghuram, Z\. Xu, E\. Ben\-David, J\. Mei, S\. Dalmia, Z\. Zhang, Y\. Liu, G\. Bansal, H\. Pankov, S\. Schwarcz, A\. Burns, C\. Chan, S\. Sanghai, R\. Liang, E\. Liang, A\. He, A\. Stuart, A\. Narayanan, Y\. Zhu, C\. Frank, B\. Fatemi, A\. Sabne, O\. Lang, I\. Bhattacharya, S\. Settle, M\. Wang, B\. McMahan, A\. Tacchetti, L\. B\. Soares, M\. Hadian, S\. Cabi, T\. Chung, N\. Putikhin, G\. Li, J\. Chen, A\. Tarango, H\. Michalewski, M\. Kazemi, H\. Masoom, H\. Sheftel, R\. Shivanna, A\. Vadali, R\. Comanescu, D\. Reid, J\. Moore, A\. Neelakantan, M\. Sander, J\. Herzig, A\. Rosenberg, M\. Dehghani, J\. Choi, M\. Fink, R\. Hayes, E\. Ge, S\. Weng, C\. Ho, J\. Karro, K\. Krishna, L\. N\. Thiet, A\. Skerry\-Ryan, D\. Eppens, M\. Andreetto, N\. Sarma, S\. Bonacina, B\. K\. Ayan, M\. Nawhal, Z\. Shan, M\. Dusenberry, S\. Thakoor, S\. Gubbi, D\. D\. Nguyen, R\. Tsarfaty, S\. Albanie, J\. Mitrović, M\. Gandhi, B\. Chen, A\. Epasto, G\. Stephanov, Y\. Jin, S\. Gehman, A\. Amini, J\. Weber, F\. Behbahani, S\. Xu, M\. Allamanis, X\. Chen, M\. Ott, C\. Sha, M\. Jastrzebski, H\. Qi, D\. Greene, X\. Wu, A\. Toki, D\. Vlasic, J\. Shapiro, R\. Kotikalapudi, Z\. Shen, T\. Saeki, S\. Xie, A\. Cassirer, S\. Bharadwaj, T\. Kiyono, S\. Bhojanapalli, E\. Rosenfeld, S\. Ritter, J\. Mao, J\. G\. Oliveira, Z\. Egyed, B\. Bandemer, E\. Parisotto, K\. Kinoshita, J\. Pluto, P\. Maniatis, S\. Li, Y\. Guo, G\. Ghiasi, J\. Tarbouriech, S\. Chatterjee, J\. Jin, Katrina, Xu, J\. Palomaki, S\. Arnold, M\. Sewak, F\. Piccinini, M\. Sharma, B\. Albrecht, S\. Purser\-haskell, A\. Vaswani, C\. Chen, M\. Wisniewski, Q\. Cao, J\. Aslanides, N\. M\. Phu, M\. Sieb, L\. Agubuzu, A\. Zheng, D\. Sohn, M\. Selvi, A\. Andreassen, K\. Subudhi, P\. Eruvbetine, O\. Woodman, T\. Mery, S\. Krause, X\. Ren, X\. Ma, J\. Luo, D\. Chen, W\. Fan, H\. Griffiths, C\. Schuler, A\. Li, S\. Zhang, J\. Sarr, S\. Luo, R\. Patana, M\. Watson, D\. Naboulsi, M\. Collins, S\. Sidhwani, E\. Hoogeboom, S\. Silver, E\. Caveness, X\. Zhao, M\. Rodriguez, M\. Deines, L\. Bai, P\. Griffin, M\. Tagliasacchi, E\. Xue, S\. R\. Babbula, B\. Pang, N\. Ding, G\. Shen, E\. Peake, R\. Crocker, S\. S\. Raghvendra, D\. Swisher, W\. Han, R\. Singh, L\. Wu, V\. Pchelin, T\. Munkhdalai, D\. Alon, G\. Bacon, E\. Robles, J\. Bulian, M\. Johnson, G\. Powell, F\. T\. Ferreira, Y\. Li, F\. Benzing, M\. Velimirović, H\. Soyer, W\. Kong, Tony, Nguyên, Z\. Yang, J\. Liu, J\. van Amersfoort, D\. Gillick, B\. Sun, N\. Rauschmayr, K\. Zhang, S\. Zhan, T\. Zhou, A\. Frolov, C\. Yang, D\. Vnukov, L\. Rouillard, H\. Li, A\. Mandhane, N\. Fallen, R\. Venkataraman, C\. H\. Hu, J\. Brennan, J\. Lee, J\. Chang, M\. Sundermeyer, Z\. Pan, R\. Ke, S\. Tong, A\. Fabrikant, W\. Bono, J\. Gu, R\. Foley, Y\. Mao, M\. Delakis, D\. Bhaswar, R\. Frostig, N\. Li, A\. Zipori, C\. Hope, O\. Kozlova, S\. Mishra, J\. Djolonga, C\. Schiff, M\. A\. Merey, E\. Briakou, P\. Morgan, A\. Wan, A\. Hassidim, R\. Skerry\-Ryan, K\. Sengupta, M\. Jasarevic, P\. Kallakuri, P\. Kunkle, H\. Brennan, T\. Lieber, H\. Mansoor, J\. Walker, B\. Zhang, A\. Xie, G\. Žužić, A\. Chukwuka, A\. Druinsky, D\. Cho, R\. Yao, F\. Naeem, S\. Butt, E\. Kim, Z\. Jia, M\. Jordan, A\. Lelkes, M\. Kurzeja, S\. Wang, J\. Zhao, A\. Over, A\. Chakladar, M\. Prasetya, N\. Jha, S\. Ganapathy, Y\. Cong, P\. Shroff, C\. Saroufim, S\. Miryoosefi, M\. Hammad, T\. Nasir, W\. Xi, Y\. Gao, Y\. Maeng, B\. Hora, C\. Cheng, P\. Haghani, Y\. Lewenberg, C\. Lu, M\. Matysiak, N\. Raisinghani, H\. Wang, L\. Baugher, R\. Sukthankar, M\. Giang, J\. Schultz, N\. Fiedel, M\. Chen, C\. Lee, T\. Dey, H\. Zheng, S\. Paul, C\. Smith, A\. Ly, Y\. Wang, R\. Bansal, B\. Perz, S\. Ricco, S\. Blank, V\. Keshava, D\. Sharma, M\. Chow, K\. Lad, K\. Jalan, S\. Osindero, C\. Swanson, J\. Scott, A\. Ilić, X\. Li, S\. R\. Jonnalagadda, A\. S\. Soudagar, Y\. Xiong, B\. Batsaikhan, D\. Jarrett, N\. Kumar, M\. Shah, M\. Lawlor, A\. Waters, M\. Graham, R\. May, S\. Ramos, S\. Lefdal, Z\. Cankara, N\. Cano, B\. O’Donoghue, J\. Borovik, F\. Liu, J\. Grimstad, M\. Alnahlawi, K\. Tsihlas, T\. Hudson, N\. Grigorev, Y\. Jia, T\. Huang, T\. P\. Igwe, S\. Lebedev, X\. Tang, I\. Krivokon, F\. Garcia, M\. Tan, E\. Jia, P\. Stys, S\. Vashishth, Y\. Liang, B\. Venkatraman, C\. Gu, A\. Kementsietsidis, C\. Zhu, J\. Jung, Y\. Bai, M\. J\. Hosseini, F\. Ahmed, A\. Gupta, X\. Yuan, S\. Ashraf, S\. Nigam, G\. Vasudevan, P\. Awasthi, A\. M\. Gilady, Z\. Mariet, R\. Eskander, H\. Li, H\. Hu, G\. Garrido, P\. Schlattner, G\. Zhang, R\. Saxena, P\. Dević, K\. Muralidharan, A\. Murthy, Y\. Zhou, M\. Choi, A\. Wongpanich, Z\. Wang, P\. Shah, Y\. Xu, Y\. Huang, S\. Spencer, A\. Chen, J\. Cohan, J\. Wang, J\. Tompson, J\. Wu, R\. Haroun, H\. Li, B\. Huergo, F\. Yang, T\. Yin, J\. Wendt, M\. Bendersky, R\. Chaabouni, J\. Snaider, J\. Ferret, A\. Jindal, T\. Thompson, A\. Xue, W\. Bishop, S\. M\. Phal, A\. Sharma, Y\. Sung, P\. Radhakrishnan, M\. Shomrat, R\. Ingle, R\. Vij, J\. Gilmer, M\. D\. Istin, S\. Sobell, Y\. Lu, E\. Nottage, D\. Sadigh, J\. Willcock, T\. Zhang, S\. Xu, S\. Brown, K\. Lee, G\. Wang, Y\. Zhu, Y\. Tay, C\. Kim, A\. Gutierrez, A\. Sharma, Y\. Xian, S\. Seo, C\. Cui, E\. Pochernina, C\. Baetu, K\. Jastrzębski, M\. Ly, M\. Elhawaty, D\. Suh, E\. Sezener, P\. Wang, N\. Yuen, G\. Tucker, J\. Cai, Z\. Yang, C\. Wang, A\. Muzio, H\. Qian, J\. Yoo, D\. Lockhart, K\. R\. McKee, M\. Guo, M\. Mehrotra, A\. Mendonça, S\. V\. Mehta, S\. Ben, C\. Tekur, J\. Mu, M\. Zhu, V\. Krakovna, H\. Lee, A\. Maschinot, S\. Cevey, H\. Choe, A\. Bai, H\. Srinivasan, D\. Gasaway, N\. Young, P\. Siegler, D\. Holtmann\-Rice, V\. Piratla, K\. Baumli, R\. Yogev, A\. Hofer, H\. van Hasselt, S\. Grant, Y\. Chervonyi, D\. Silver, A\. Hogue, A\. Agarwal, K\. Wang, P\. Singh, F\. Flynn, J\. Lipschultz, R\. David, L\. Bellot, Y\. Yang, L\. Le, F\. Graziano, K\. Olszewska, K\. Hui, A\. Maurya, N\. Parotsidis, W\. Chen, T\. Oguntebi, J\. Kelley, A\. Baddepudi, J\. Mauerer, G\. Shaw, A\. Siegman, L\. Yang, S\. Shetty, S\. Roy, Y\. Song, W\. Stokowiec, R\. Burnell, O\. Savant, R\. Busa\-Fekete, J\. Miao, S\. Ghosh, L\. MacDermed, P\. Lippe, M\. Dektiarev, Z\. Behrman, F\. Mentzer, K\. Nguyen, M\. Wei, S\. Verma, C\. Knutsen, S\. Dasari, Z\. Yan, P\. Mitrichev, X\. Wang, V\. Shejwalkar, J\. Austin, S\. Sunkara, N\. Potti, Y\. Virin, C\. Wright, G\. Liu, O\. Riva, E\. Pot, G\. Kochanski, Q\. Le, G\. Balasubramaniam, A\. Dhar, Y\. Liao, A\. Bloniarz, D\. Shukla, E\. Cole, J\. Lee, S\. Zhang, S\. Kafle, S\. Vashishtha, P\. Mahmoudieh, G\. Chen, R\. Hoffmann, P\. Srinivasan, A\. D\. Lago, Y\. B\. Shalom, Z\. Wang, M\. Elabd, A\. Sharma, J\. Oh, S\. Kothawade, M\. Le, M\. Monteiro, S\. Yang, K\. Alarakyia, R\. Geirhos, D\. Mincu, H\. Garnes, H\. Kobayashi, S\. Mariooryad, K\. Krasowiak, Zhixin, Lai, S\. Mourad, M\. Wang, F\. Bu, O\. Aharoni, G\. Chen, A\. Goyal, V\. Zubov, A\. Bapna, E\. Dabir, N\. Kothari, K\. Lamerigts, N\. D\. Cao, J\. Shar, C\. Yew, N\. Kulkarni, D\. Mahaarachchi, M\. Joshi, Z\. Zhu, J\. Lichtarge, Y\. Zhou, H\. Muckenhirn, V\. Selo, O\. Vinyals, P\. Chen, A\. Brohan, V\. Mehta, S\. Cogan, R\. Wang, T\. Geri, W\. Ko, W\. Chen, F\. Viola, K\. Shivam, L\. Wang, M\. C\. Elish, R\. A\. Popa, S\. Pereira, J\. Liu, R\. Koster, D\. Kim, G\. Zhang, S\. Ebrahimi, P\. Talukdar, Y\. Zheng, P\. Poklukar, A\. Mikhalap, D\. Johnson, A\. Vijayakumar, M\. Omernick, M\. Dibb, A\. Dubey, Q\. Hu, A\. Suman, V\. Aggarwal, I\. Kornakov, F\. Xia, W\. Lowe, A\. Kolganov, T\. Xiao, V\. Nikolaev, S\. Hemingray, B\. Li, J\. Iljazi, M\. Rybiński, B\. Sandhu, P\. Lu, T\. Luong, R\. Jenatton, V\. Govindaraj, Hui, Li, G\. Dulac\-Arnold, W\. Park, H\. Wang, A\. Modi, J\. Pouget\-Abadie, K\. Greller, R\. Gupta, R\. Berry, P\. Ramachandran, J\. Xie, L\. McCafferty, J\. Wang, K\. Gupta, H\. Lim, B\. Bratanič, A\. Brock, I\. Akolzin, J\. Sproch, D\. Karliner, D\. Kim, A\. Goedeckemeyer, N\. Shazeer, C\. Schmid, D\. Calandriello, P\. Bhatia, K\. Choromanski, C\. Montgomery, D\. Dua, A\. Ramalho, H\. King, Y\. Gao, L\. Nguyen, D\. Lindner, D\. Pitta, O\. Johnson, K\. Salama, D\. Ardila, M\. Han, E\. Farnese, S\. Odoom, Z\. Wang, X\. Ding, N\. Rink, R\. Smith, H\. T\. Lehri, E\. Cohen, N\. Vats, T\. He, P\. Gopavarapu, A\. Paszke, M\. Patel, W\. V\. Gansbeke, L\. Loher, L\. Castro, M\. Voitovich, T\. von Glehn, N\. George, S\. Niklaus, Z\. Eaton\-Rosen, N\. Rakićević, E\. Jue, S\. Perel, C\. Zhang, Y\. Bahat, A\. Pouget, Z\. Xing, F\. Huot, A\. Shenoy, T\. Bos, V\. Coriou, B\. Richter, N\. Noy, Y\. Wang, S\. Ontanon, S\. Qin, G\. Makarchuk, D\. Hassabis, Z\. Li, M\. Sharma, K\. Venkatesan, I\. Kemaev, R\. Daniel, S\. Huang, S\. Shah, O\. Ponce, Warren, Chen, M\. Faruqui, J\. Wu, S\. Andačić, S\. Payrits, D\. McDuff, T\. Hume, Y\. Cao, M\. Tessler, Q\. Wang, Y\. Wang, I\. Rendulic, E\. Agustsson, M\. Johnson, T\. Lando, A\. Howard, S\. G\. S\. Padmanabhan, M\. Daswani, A\. Banino, M\. Kilgore, J\. Heek, Z\. Ji, A\. Caceres, C\. Li, N\. Kassner, A\. Vlaskin, Z\. Liu, A\. Grills, Y\. Hou, R\. Sukkerd, G\. Cheon, N\. Shetty, L\. Markeeva, P\. Stanczyk, T\. Iyer, Y\. Gong, S\. Gao, K\. Gopalakrishnan, T\. Blyth, M\. Reynolds, A\. Bhoopchand, M\. Bilenko, D\. Gharibian, V\. Zayats, A\. Faust, A\. Singh, M\. Ma, H\. Jiao, S\. Vijayanarasimhan, L\. Aroyo, V\. Yadav, S\. Chakera, A\. Kakarla, V\. Meshram, K\. Gregor, G\. Botea, E\. Senter, D\. Jia, G\. Kovacs, N\. Sharma, S\. Baur, K\. Kang, Y\. He, L\. Zhuo, M\. Kostelac, I\. Laish, S\. Peng, L\. O’Bryan, D\. Kasenberg, G\. R\. Rao, E\. Leurent, B\. Zhang, S\. Stevens, A\. Salazar, Y\. Zhang, I\. Lobov, J\. Walker, A\. Porter, M\. Redshaw, H\. Ke, A\. Rao, A\. Lee, H\. Lam, M\. Moffitt, J\. Kim, S\. Qiao, T\. Koo, R\. Dadashi, X\. Song, M\. Sundararajan, P\. Xu, C\. Kawamoto, Y\. Zhong, C\. Barbu, A\. Reddy, M\. Verzetti, L\. Li, G\. Papamakarios, H\. Klimczak\-Plucińska, M\. Cassin, K\. Kavukcuoglu, R\. Swavely, A\. Vaucher, J\. Zhao, R\. Hemsley, M\. Tschannen, H\. Ge, G\. Menghani, Y\. Yu, N\. Ha, W\. He, X\. Wu, M\. Song, R\. Sterneck, S\. Zinke, D\. A\. Calian, A\. Marsden, A\. C\. Ruiz, M\. Hessel, A\. Gueta, B\. Lee, B\. Farris, M\. Gupta, Y\. Li, M\. Saleh, V\. Misra, K\. Xiao, P\. Mendolicchio, G\. Buttimore, V\. Krayvanova, N\. Nayakanti, M\. Wiethoff, Y\. Pande, A\. Mirhoseini, N\. Lao, J\. Liu, Y\. Hua, A\. Chen, Y\. Malkov, D\. Kalashnikov, S\. Gupta, K\. Audhkhasi, Y\. Zhai, S\. Kopalle, P\. Jain, E\. Ofek, C\. Meyer, K\. Baatarsukh, H\. Strejček, J\. Qian, J\. Freedman, R\. Figueira, M\. Sokolik, O\. Bachem, R\. Lin, D\. Kharrat, C\. Hidey, P\. Xu, D\. Duan, Y\. Li, M\. Ersoy, R\. Everett, K\. Cen, R\. Santamaria\-Fernandez, A\. Taubenfeld, I\. Mackinnon, L\. Deng, P\. Zablotskaia, S\. Viswanadha, S\. Goel, D\. Yates, Y\. Deng, P\. Choy, M\. Chen, A\. Sinha, A\. Mossin, Y\. Wang, A\. Szlam, S\. Hao, P\. K\. Rubenstein, M\. Toksoz\-Exley, M\. Aperghis, Y\. Zhong, J\. Ahn, M\. Isard, O\. Lacombe, F\. Luisier, C\. Anastasiou, Y\. Kalley, U\. Prabhu, E\. Dunleavy, S\. Bijwadia, J\. Mao\-Jones, K\. Chen, R\. Pasumarthi, E\. Wood, A\. Dostmohamed, N\. Hurley, J\. Simsa, A\. Parrish, M\. Pajarskas, M\. Harvey, O\. Skopek, Y\. Kochinski, J\. Rey, V\. Rieser, D\. Zhou, S\. J\. Lee, T\. Acharya, G\. Li, J\. Jiang, X\. Zhang, B\. Gipson, E\. Mahintorabi, M\. Gelmi, N\. Khajehnouri, A\. Yeh, K\. Lee, L\. Matthey, L\. Baker, T\. Pham, H\. Fu, A\. Pak, P\. Gupta, C\. Vasconcelos, A\. Sadovsky, B\. Walker, S\. Hsiao, P\. Zochbauer, A\. Marzoca, N\. Velan, J\. Zeng, G\. Baechler, D\. Driess, D\. Jain, Y\. Huang, L\. Tao, J\. Maggs, N\. Levine, J\. Schneider, E\. Gemzer, S\. Petit, S\. Han, Z\. Fisher, D\. Zelle, C\. Biles, E\. Ie, A\. Fadeeva, C\. Liu, J\. V\. Franco, A\. Collister, H\. Zhang, R\. Wang, R\. Zhao, L\. Kieliger, K\. Shuster, R\. Zhu, B\. Gong, L\. Chan, R\. Sun, S\. Basu, R\. Zimmermann, J\. Hayes, A\. Bapna, J\. Snoek, W\. Yang, P\. Datta, J\. A\. Abdallah, K\. Kilgour, L\. Li, S\. Mah, Y\. Jun, M\. Rivière, A\. Karmarkar, T\. Spalink, T\. Huang, L\. Gonzalez, D\. Tran, A\. Nowak, J\. Palowitch, M\. Chadwick, E\. Talius, H\. Mehta, T\. Sellam, P\. Fränken, M\. Nicosia, K\. He, A\. Kini, D\. Amos, S\. Basu, H\. Jobe, E\. Shaw, Q\. Xu, C\. Evans, D\. Ikeda, C\. Yan, L\. Jin, L\. Wang, S\. Yadav, I\. Labzovsky, R\. Sampath, A\. Ma, C\. Schumann, A\. Siddhant, R\. Shah, J\. Youssef, R\. Agarwal, N\. Dabney, A\. Tonioni, M\. Ambar, J\. Li, I\. Guyon, B\. Li, D\. Soergel, B\. Fang, G\. Karadzhov, C\. Udrescu, T\. Trinh, V\. Raunak, S\. Noury, D\. Guo, S\. Gupta, M\. Finkelstein, D\. Petek, L\. Liang, G\. Billock, P\. Sun, D\. Wood, Y\. Song, X\. Yu, T\. Matejovicova, R\. Cohen, K\. Andra, D\. D’Ambrosio, Z\. Deng, V\. Nallatamby, E\. Songhori, R\. Dangovski, A\. Lampinen, P\. Botadra, A\. Hillier, J\. Cao, N\. Baddi, A\. Kuncoro, T\. Yoshino, A\. Bhagatwala, M\. Ranzato, R\. Schaeffer, T\. Liu, S\. Ye, O\. Sarvana, J\. Nham, C\. Kuang, I\. Gao, J\. Baek, S\. Mittal, A\. Wahid, A\. Gergely, B\. Ni, J\. Feldman, C\. Muir, P\. Lamblin, W\. Macherey, E\. Dyer, L\. Kilpatrick, V\. Campos, M\. Bhutani, S\. Fort, Y\. Ahmad, A\. Severyn, K\. Chatziprimou, O\. Ferludin, M\. Dimarco, A\. Kusupati, J\. Heyward, D\. Bahir, K\. Villela, K\. Millican, D\. Marcus, S\. Bahargam, C\. Unlu, N\. Roth, Z\. Wei, S\. Gopal, D\. Ghoshal, E\. Lee, S\. Lin, J\. Lees, D\. Lee, A\. Hosseini, C\. Fan, S\. Neel, M\. Wu, Y\. Altun, H\. Cai, E\. Piqueras, J\. Woodward, A\. Bissacco, S\. Haykal, M\. Bordbar, P\. Sundaram, S\. Hodkinson, D\. Toyama, G\. Polovets, A\. Myers, A\. Sinha, T\. Levinboim, K\. Krishnakumar, R\. Chhaparia, T\. Sholokhova, N\. B\. Gundavarapu, G\. Jawahar, H\. Qureshi, J\. Hu, N\. Momchev, M\. Rahtz, R\. Wu, A\. P\. S, K\. Dhamdhere, M\. Guo, U\. Gupta, A\. Eslami, M\. Schain, M\. Blokzijl, D\. Welling, D\. Orr, L\. Bolelli, N\. Perez\-Nieves, M\. Sirotenko, A\. Prasad, A\. Kar, B\. D\. B\. Pigem, T\. Terzi, G\. Weisz, D\. Ghosh, A\. Mavalankar, D\. Madeka, K\. Daugaard, H\. Adam, V\. Shah, D\. Berman, M\. Tran, S\. Baker, E\. Andrejczuk, G\. Chole, G\. Raboshchuk, M\. Mirzazadeh, T\. Kagohara, S\. Wu, C\. Schallhart, B\. Orlando, C\. Wang, A\. Rrustemi, H\. Xiong, H\. Liu, A\. Vezer, N\. Ramsden, S\. Chang, S\. Mudgal, Y\. Li, N\. Vieillard, Y\. Hoshen, F\. Ahmad, A\. Slone, A\. Hua, N\. Potikha, M\. Rossini, J\. Stritar, S\. Prakash, Z\. Wang, X\. Dong, A\. Nazari, E\. Nehoran, K\. Tekelioglu, Y\. Li, K\. Badola, T\. Funkhouser, Y\. Li, V\. Yerram, R\. Ganeshan, D\. Formoso, K\. Langner, T\. Shi, H\. Li, Y\. Yamamori, A\. Panda, A\. Saade, A\. S\. Scarpati, C\. Breaux, C\. Carey, Z\. Zhou, C\. Hsieh, S\. Bridgers, A\. Butryna, N\. Gupta, V\. Tulsyan, S\. Woo, E\. Eltyshev, W\. Grathwohl, C\. Parks, S\. Benjamin, R\. Panigrahy, S\. Dodhia, D\. D\. Freitas, C\. Sauer, W\. Song, F\. Alet, J\. Tolins, C\. Paduraru, X\. Zhou, B\. Albert, Z\. Zhang, L\. Shu, M\. Bansal, S\. Nguyen, A\. Globerson, O\. Xiao, J\. Manyika, T\. Hennigan, R\. Rong, J\. Matak, A\. Bakalov, A\. Sharma, D\. Sinopalnikov, A\. Pierson, S\. Roller, G\. Brown, M\. Gao, T\. Fukuzawa, A\. Ghafouri, K\. Vassigh, I\. Barr, Z\. Wang, A\. Korsun, R\. Jayaram, L\. Ren, T\. Zaman, S\. Khan, Y\. Lunts, D\. Deutsch, D\. Uthus, N\. Katz, M\. Samsikova, A\. Khalifa, N\. Sethi, J\. Sun, L\. Tang, U\. Alon, X\. Luo, D\. Yu, A\. Nayyar, B\. Petrini, W\. Truong, V\. Hellendoorn, N\. Chinaev, C\. Alberti, W\. Wang, J\. Hu, V\. Mirrokni, A\. Balashankar, A\. Aharon, A\. Mehta, A\. Iscen, J\. Kready, L\. Manning, A\. Mohananey, Y\. Chen, A\. Tripathi, A\. Wu, I\. Petrovski, D\. Hwang, M\. Baeuml, S\. Chandrakaladharan, Y\. Liu, R\. Coaguila, M\. Chen, S\. Ma, P\. Tafti, S\. Tatineni, T\. Spitz, J\. Ye, P\. Vicol, M\. Rosca, A\. Puigdomènech, Z\. Yahav, S\. Ghemawat, H\. Lin, P\. Kirk, Z\. Nabulsi, S\. Brin, B\. Bohnet, K\. Caluwaerts, A\. S\. Veerubhotla, D\. Zheng, Z\. Dai, P\. Petrov, Y\. Xu, R\. Mehran, Z\. Xu, L\. Zintgraf, J\. Choi, S\. A\. Hombaiah, R\. Thoppilan, S\. Reddi, L\. Lew, L\. Li, K\. Webster, K\. Sawhney, L\. Lamprou, S\. Shakeri, M\. Lunayach, J\. Chen, S\. Bagri, A\. Salcianu, Y\. Chen, Y\. Donchev, C\. Magister, S\. Nørly, V\. Rodrigues, T\. Izo, H\. Noga, J\. Zou, T\. Köppe, W\. Zhou, K\. Lee, X\. Long, D\. Eisenbud, A\. Chen, C\. Schenck, C\. M\. To, P\. Zhong, E\. Taropa, M\. Truong, O\. Levy, D\. Martins, Z\. Zhang, C\. Semturs, K\. Zhang, A\. Yakubovich, P\. Moreno, L\. McConnaughey, D\. Lu, S\. Redmond, L\. Weerts, Y\. Bitton, T\. Refice, N\. Lacasse, A\. Conmy, C\. Tallec, J\. Odell, H\. Forbes\-Pollard, A\. Socala, J\. Hoech, P\. Kohli, A\. Walton, R\. Wang, M\. Sazanovich, K\. Zhu, A\. Kapishnikov, R\. Galt, M\. Denton, B\. Murdoch, C\. Sikora, K\. Mohamed, W\. Wei, U\. First, T\. McConnell, L\. C\. Cobo, J\. Qin, T\. Avrahami, D\. Balle, Y\. Watanabe, A\. Louis, A\. Kraft, S\. Ariafar, Y\. Gu, E\. Rives, C\. Yoon, A\. Rusu, J\. Cobon\-Kerr, C\. Hahn, J\. Luo, Yuvein, Zhu, N\. Ahuja, R\. Benenson, R\. L\. Kaufman, H\. Yu, L\. Hightower, J\. Zhang, D\. Ni, L\. A\. Hendricks, G\. Wang, G\. Yona, L\. Jain, P\. Barrio, S\. Bhupatiraju, S\. Velusamy, A\. Dafoe, S\. Riedel, T\. Thomas, Z\. Yuan, M\. Bellaiche, S\. Panthaplackel, K\. Kloboves, S\. Jauhari, C\. Akbulut, T\. Davchev, E\. Gladchenko, D\. Madras, A\. Chuklin, T\. Hill, Q\. Yuan, M\. Madhavan, L\. Leonhard, D\. Scandinaro, Q\. Chen, N\. Niu, A\. Douillard, B\. Damoc, Y\. Onoe, F\. Pedregosa, F\. Bertsch, C\. Leichner, J\. Pagadora, J\. Malmaud, S\. Ponda, A\. Twigg, O\. Duzhyi, J\. Shen, M\. Wang, R\. Garg, J\. Chen, U\. Evci, J\. Lee, L\. Liu, K\. Kojima, M\. Yamaguchi, A\. Rajendran, A\. Piergiovanni, V\. K\. Rajendran, M\. Fornoni, G\. Ibagon, H\. Ragan, S\. M\. Khan, J\. Blitzer, A\. Bunner, G\. Sun, T\. Kosakai, S\. Lundberg, N\. Elue, K\. Guu, S\. Park, J\. Park, A\. Narayanaswamy, C\. Wu, J\. Mudigonda, T\. Cohn, H\. Mu, R\. Kumar, L\. Graesser, Y\. Zhang, R\. Killam, V\. Zhuang, M\. Giménez, W\. A\. Jishi, R\. Ley\-Wild, A\. Zhai, K\. Osawa, D\. Cedillo, J\. Liu, M\. Upadhyay, M\. Sieniek, R\. Sharma, T\. Paine, A\. Angelova, S\. Addepalli, C\. Parada, K\. Majumder, A\. Lamp, S\. Kumar, X\. Deng, A\. Myaskovsky, T\. Sabolić, J\. Dudek, S\. York, F\. de Chaumont Quitry, J\. Nie, D\. Cattle, A\. Gunjan, B\. Piot, W\. Khawaja, S\. Bang, S\. Wang, S\. Khodadadeh, R\. R, P\. Rawlani, R\. Powell, K\. Lee, J\. Griesser, G\. Oh, C\. Magalhaes, Y\. Li, S\. Tokumine, H\. N\. Vogel, D\. Hsu, A\. BC, D\. Jindal, M\. Cohen, Z\. Yang, J\. Yuan, D\. de Cesare, T\. Bruguier, J\. Xu, M\. Roy, A\. Jacovi, D\. Belov, R\. Arya, P\. Meadowlark, S\. Cohen\-Ganor, W\. Ye, P\. Morris\-Suzuki, P\. Banzal, G\. Song, P\. Ponnuramu, F\. Zhang, G\. Scrivener, S\. Zaiem, A\. R\. Rochman, K\. Han, B\. Ghazi, K\. Lee, S\. Drath, D\. Suo, A\. Girgis, P\. Shenoy, D\. Nguyen, D\. Eck, S\. Gupta, L\. Yan, J\. Carreira, A\. Gulati, R\. Sang, D\. Mirylenka, E\. Cooney, E\. Chou, M\. Ling, C\. Fan, B\. Coleman, G\. Tubone, R\. Kumar, J\. Baldridge, F\. Hernandez\-Campos, A\. Lazaridou, J\. Besley, I\. Yona, N\. Bulut, Q\. Wellens, A\. Pierigiovanni, J\. George, R\. Green, P\. Han, C\. Tao, G\. Clark, C\. You, A\. Abdolmaleki, J\. Fu, T\. Chen, A\. Chaugule, A\. Chandorkar, A\. Rahman, W\. Thompson, P\. Koanantakool, M\. Bernico, J\. Ren, A\. Vlasov, S\. Vassilvitskii, M\. Kula, Y\. Liang, D\. Kim, Y\. Huang, C\. Ye, D\. Lepikhin, and W\. HelmholzGemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix2.p1.1.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1)\.
- de Haanet al\.\(2019\)P\. de Haan, D\. Jayaraman, and S\. LevineCausal confusion in imitation learning\.InAdvances in Neural Information Processing Systems,H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(Eds\.\),Vol\.32,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2019/file/947018640bf36a2bb609d3557a285329-Paper.pdf)Cited by:[§A\.4](https://arxiv.org/html/2608.21946#A1.SS4.p1.1)\.
- Fanget al\.\(2026\)R\. Fang, Y\. Liang, X\. Wang, J\. Wu, S\. Qiao, P\. Xie, F\. Huang, H\. Chen, and N\. ZhangMemp: exploring agent procedural memory\.External Links:2508\.06433,[Link](https://arxiv.org/abs/2508.06433)Cited by:[§1](https://arxiv.org/html/2608.21946#S1.p2.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Fenget al\.\(2025\)L\. Feng, Z\. Xue, T\. Liu, and B\. AnGroup\-in\-group policy optimization for llm agent training\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 46375–46408\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/420c9f777c0b4f78d515e53cf74d58b2-Paper-Conference.pdf)Cited by:[§B\.2](https://arxiv.org/html/2608.21946#A2.SS2.p1.1),[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Hübotteret al\.\(2026\)J\. Hübotter, F\. Lübeck, L\. Behric, A\. Baumann, M\. Bagatella, D\. Marta, I\. Hakimi, I\. Shenfeld, T\. K\. Buening, C\. Guestrin, and A\. KrauseReinforcement learning via self\-distillation\.External Links:2601\.20802,[Link](https://arxiv.org/abs/2601.20802)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Jinet al\.\(2025\)B\. Jin, H\. Zeng, Z\. Yue, J\. Yoon, S\. Arik, D\. Wang, H\. Zamani, and J\. HanSearch\-r1: training llms to reason and leverage search engines with reinforcement learning\.External Links:2503\.09516,[Link](https://arxiv.org/abs/2503.09516)Cited by:[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Kaelblinget al\.\(1998\)L\. P\. Kaelbling, M\. L\. Littman, and A\. R\. CassandraPlanning and acting in partially observable stochastic domains\.Artificial Intelligence101\(1–2\),pp\. 99–134\.Cited by:[§2](https://arxiv.org/html/2608.21946#S2.p1.1)\.
- Liuet al\.\(2026\)Z\. Liu, J\. Kim, X\. Luo, D\. Li, and Y\. YangExploratory memory\-augmented llm agent via hybrid on\- and off\-policy optimization\.External Links:2602\.23008,[Link](https://arxiv.org/abs/2602.23008)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix11.p1.1.1),[§1](https://arxiv.org/html/2608.21946#S1.p2.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Luet al\.\(2026a\)Z\. Lu, Z\. Yao, Z\. Han, Z\. Wang, J\. Wu, Q\. Gu, X\. Cai, W\. Lu, J\. Xiao, Y\. Zhuang, and Y\. ShenSelf\-distilled agentic reinforcement learning\.External Links:2605\.15155,[Link](https://arxiv.org/abs/2605.15155)Cited by:[Table 1](https://arxiv.org/html/2608.21946#S3.T1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Luet al\.\(2026b\)Z\. Lu, Z\. Yao, J\. Wu, C\. Han, Q\. Gu, X\. Cai, W\. Lu, J\. Xiao, Y\. Zhuang, and Y\. ShenSKILL0: in\-context agentic reinforcement learning for skill internalization\.External Links:2604\.02268,[Link](https://arxiv.org/abs/2604.02268)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Maet al\.\(2026\)W\. Ma, Y\. Zeng, Y\. Song, X\. Cui, J\. Zhao, X\. Liu, and M\. ElhoseinyFreshness\-aware prioritized experience replay for llm/vlm reinforcement learning\.External Links:2604\.16918,[Link](https://arxiv.org/abs/2604.16918)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- OpenAIet al\.\(2024\)OpenAI, :, A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford, A\. Mądry, A\. Baker\-Whitcomb, A\. Beutel, A\. Borzunov, A\. Carney, A\. Chow, A\. Kirillov, A\. Nichol, A\. Paino, A\. Renzin, A\. T\. Passos, A\. Kirillov, A\. Christakis, A\. Conneau, A\. Kamali, A\. Jabri, A\. Moyer, A\. Tam, A\. Crookes, A\. Tootoochian, A\. Tootoonchian, A\. Kumar, A\. Vallone, A\. Karpathy, A\. Braunstein, A\. Cann, A\. Codispoti, A\. Galu, A\. Kondrich, A\. Tulloch, A\. Mishchenko, A\. Baek, A\. Jiang, A\. Pelisse, A\. Woodford, A\. Gosalia, A\. Dhar, A\. Pantuliano, A\. Nayak, A\. Oliver, B\. Zoph, B\. Ghorbani, B\. Leimberger, B\. Rossen, B\. Sokolowsky, B\. Wang, B\. Zweig, B\. Hoover, B\. Samic, B\. McGrew, B\. Spero, B\. Giertler, B\. Cheng, B\. Lightcap, B\. Walkin, B\. Quinn, B\. Guarraci, B\. Hsu, B\. Kellogg, B\. Eastman, C\. Lugaresi, C\. Wainwright, C\. Bassin, C\. Hudson, C\. Chu, C\. Nelson, C\. Li, C\. J\. Shern, C\. Conger, C\. Barette, C\. Voss, C\. Ding, C\. Lu, C\. Zhang, C\. Beaumont, C\. Hallacy, C\. Koch, C\. Gibson, C\. Kim, C\. Choi, C\. McLeavey, C\. Hesse, C\. Fischer, C\. Winter, C\. Czarnecki, C\. Jarvis, C\. Wei, C\. Koumouzelis, D\. Sherburn, D\. Kappler, D\. Levin, D\. Levy, D\. Carr, D\. Farhi, D\. Mely, D\. Robinson, D\. Sasaki, D\. Jin, D\. Valladares, D\. Tsipras, D\. Li, D\. P\. Nguyen, D\. Findlay, E\. Oiwoh, E\. Wong, E\. Asdar, E\. Proehl, E\. Yang, E\. Antonow, E\. Kramer, E\. Peterson, E\. Sigler, E\. Wallace, E\. Brevdo, E\. Mays, F\. Khorasani, F\. P\. Such, F\. Raso, F\. Zhang, F\. von Lohmann, F\. Sulit, G\. Goh, G\. Oden, G\. Salmon, G\. Starace, G\. Brockman, H\. Salman, H\. Bao, H\. Hu, H\. Wong, H\. Wang, H\. Schmidt, H\. Whitney, H\. Jun, H\. Kirchner, H\. P\. de Oliveira Pinto, H\. Ren, H\. Chang, H\. W\. Chung, I\. Kivlichan, I\. O’Connell, I\. O’Connell, I\. Osband, I\. Silber, I\. Sohl, I\. Okuyucu, I\. Lan, I\. Kostrikov, I\. Sutskever, I\. Kanitscheider, I\. Gulrajani, J\. Coxon, J\. Menick, J\. Pachocki, J\. Aung, J\. Betker, J\. Crooks, J\. Lennon, J\. Kiros, J\. Leike, J\. Park, J\. Kwon, J\. Phang, J\. Teplitz, J\. Wei, J\. Wolfe, J\. Chen, J\. Harris, J\. Varavva, J\. G\. Lee, J\. Shieh, J\. Lin, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Jang, J\. Q\. Candela, J\. Beutler, J\. Landers, J\. Parish, J\. Heidecke, J\. Schulman, J\. Lachman, J\. McKay, J\. Uesato, J\. Ward, J\. W\. Kim, J\. Huizinga, J\. Sitkin, J\. Kraaijeveld, J\. Gross, J\. Kaplan, J\. Snyder, J\. Achiam, J\. Jiao, J\. Lee, J\. Zhuang, J\. Harriman, K\. Fricke, K\. Hayashi, K\. Singhal, K\. Shi, K\. Karthik, K\. Wood, K\. Rimbach, K\. Hsu, K\. Nguyen, K\. Gu\-Lemberg, K\. Button, K\. Liu, K\. Howe, K\. Muthukumar, K\. Luther, L\. Ahmad, L\. Kai, L\. Itow, L\. Workman, L\. Pathak, L\. Chen, L\. Jing, L\. Guy, L\. Fedus, L\. Zhou, L\. Mamitsuka, L\. Weng, L\. McCallum, L\. Held, L\. Ouyang, L\. Feuvrier, L\. Zhang, L\. Kondraciuk, L\. Kaiser, L\. Hewitt, L\. Metz, L\. Doshi, M\. Aflak, M\. Simens, M\. Boyd, M\. Thompson, M\. Dukhan, M\. Chen, M\. Gray, M\. Hudnall, M\. Zhang, M\. Aljubeh, M\. Litwin, M\. Zeng, M\. Johnson, M\. Shetty, M\. Gupta, M\. Shah, M\. Yatbaz, M\. J\. Yang, M\. Zhong, M\. Glaese, M\. Chen, M\. Janner, M\. Lampe, M\. Petrov, M\. Wu, M\. Wang, M\. Fradin, M\. Pokrass, M\. Castro, M\. O\. T\. de Castro, M\. Pavlov, M\. Brundage, M\. Wang, M\. Khan, M\. Murati, M\. Bavarian, M\. Lin, M\. Yesildal, N\. Soto, N\. Gimelshein, N\. Cone, N\. Staudacher, N\. Summers, N\. LaFontaine, N\. Chowdhury, N\. Ryder, N\. Stathas, N\. Turley, N\. Tezak, N\. Felix, N\. Kudige, N\. Keskar, N\. Deutsch, N\. Bundick, N\. Puckett, O\. Nachum, O\. Okelola, O\. Boiko, O\. Murk, O\. Jaffe, O\. Watkins, O\. Godement, O\. Campbell\-Moore, P\. Chao, P\. McMillan, P\. Belov, P\. Su, P\. Bak, P\. Bakkum, P\. Deng, P\. Dolan, P\. Hoeschele, P\. Welinder, P\. Tillet, P\. Pronin, P\. Tillet, P\. Dhariwal, Q\. Yuan, R\. Dias, R\. Lim, R\. Arora, R\. Troll, R\. Lin, R\. G\. Lopes, R\. Puri, R\. Miyara, R\. Leike, R\. Gaubert, R\. Zamani, R\. Wang, R\. Donnelly, R\. Honsby, R\. Smith, R\. Sahai, R\. Ramchandani, R\. Huet, R\. Carmichael, R\. Zellers, R\. Chen, R\. Chen, R\. Nigmatullin, R\. Cheu, S\. Jain, S\. Altman, S\. Schoenholz, S\. Toizer, S\. Miserendino, S\. Agarwal, S\. Culver, S\. Ethersmith, S\. Gray, S\. Grove, S\. Metzger, S\. Hermani, S\. Jain, S\. Zhao, S\. Wu, S\. Jomoto, S\. Wu, Shuaiqi, Xia, S\. Phene, S\. Papay, S\. Narayanan, S\. Coffey, S\. Lee, S\. Hall, S\. Balaji, T\. Broda, T\. Stramer, T\. Xu, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Cunninghman, T\. Degry, T\. Dimson, T\. Raoux, T\. Shadwell, T\. Zheng, T\. Underwood, T\. Markov, T\. Sherbakov, T\. Rubin, T\. Stasi, T\. Kaftan, T\. Heywood, T\. Peterson, T\. Walters, T\. Eloundou, V\. Qi, V\. Moeller, V\. Monaco, V\. Kuo, V\. Fomenko, W\. Chang, W\. Zheng, W\. Zhou, W\. Manassra, W\. Sheu, W\. Zaremba, Y\. Patil, Y\. Qian, Y\. Kim, Y\. Cheng, Y\. Zhang, Y\. He, Y\. Zhang, Y\. Jin, Y\. Dai, and Y\. MalkovGPT\-4o system card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix1.p1.1.1),[§B\.2](https://arxiv.org/html/2608.21946#A2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px3.p1.1)\.
- Ouyanget al\.\(2026\)S\. Ouyang, J\. Yan, I\. Hsu, Y\. Chen, K\. Jiang, Z\. Wang, R\. Han, L\. T\. Le, S\. Daruki, X\. Tang, V\. Tirumalashetty, G\. Lee, M\. Rofouei, H\. Lin, J\. Han, C\. Lee, and T\. PfisterReasoningBank: scaling agent self\-evolving with reasoning memory\.External Links:2509\.25140,[Link](https://arxiv.org/abs/2509.25140)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Pintoet al\.\(2017\)L\. Pinto, M\. Andrychowicz, P\. Welinder, W\. Zaremba, and P\. AbbeelAsymmetric actor critic for image\-based robot learning\.External Links:1710\.06542,[Link](https://arxiv.org/abs/1710.06542)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Qinet al\.\(2025\)Y\. Qin, X\. Tan, Z\. He, G\. Li, H\. Lin, Z\. Li, Z\. Xu, Y\. Shi, S\. Cai, R\. Rui, S\. Cai, Y\. Cai, X\. Zhang, S\. Ye, K\. Li, and X\. SunLearn the ropes, then trust the wins: self\-imitation with progressive exploration for agentic reinforcement learning\.External Links:2509\.22601,[Link](https://arxiv.org/abs/2509.22601)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Qwenet al\.\(2025\)Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px3.p1.1)\.
- Rosset al\.\(2011\)S\. Ross, G\. Gordon, and D\. BagnellA reduction of imitation learning and structured prediction to no\-regret online learning\.InProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics,G\. Gordon, D\. Dunson, and M\. Dudík \(Eds\.\),Proceedings of Machine Learning Research, Vol\.15,Fort Lauderdale, FL, USA,pp\. 627–635\.External Links:[Link](https://proceedings.mlr.press/v15/ross11a.html)Cited by:[§A\.4](https://arxiv.org/html/2608.21946#A1.SS4.p1.1)\.
- Schulmanet al\.\(2017\)J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. KlimovProximal policy optimization algorithms\.External Links:1707\.06347,[Link](https://arxiv.org/abs/1707.06347)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.External Links:2402\.03300,[Link](https://arxiv.org/abs/2402.03300)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix5.p1.1.1),[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§2](https://arxiv.org/html/2608.21946#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix4.p1.1.1),[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§1](https://arxiv.org/html/2608.21946#S1.p2.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Shridharet al\.\(2021\)M\. Shridhar, X\. Yuan, M\. Côté, Y\. Bisk, A\. Trischler, and M\. J\. HausknechtALFWorld: aligning text and embodied environments for interactive learning\.In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\-7, 2021,External Links:[Link](https://openreview.net/forum?id=0IOX0YcCdTn)Cited by:[Appendix F](https://arxiv.org/html/2608.21946#A6.p1.1),[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§1](https://arxiv.org/html/2608.21946#S1.p4.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Snellet al\.\(2022\)C\. Snell, D\. Klein, and R\. ZhongLearning by distilling context\.External Links:2209\.15189,[Link](https://arxiv.org/abs/2209.15189)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Vapnik and Vashist \(2009\)V\. Vapnik and A\. VashistA new learning paradigm: learning using privileged information\.Neural Networks22\(5\),pp\. 544–557\.Note:Advances in Neural Networks Research: IJCNN2009External Links:ISSN 0893\-6080,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.neunet.2009.06.042),[Link](https://www.sciencedirect.com/science/article/pii/S0893608009001130)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2025a\)H\. Wang, C\. T\. Leong, J\. Wang, J\. Wang, and W\. LiSPA\-rl: reinforcing llm agents via stepwise progress attribution\.External Links:2505\.20732,[Link](https://arxiv.org/abs/2505.20732)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2026\)J\. Wang, Q\. Yan, Y\. Wang, Y\. Tian, S\. S\. Mishra, Z\. Xu, M\. Gandhi, P\. Xu, and L\. L\. CheongReinforcement learning for self\-improving agent with skill library\.External Links:2512\.17102,[Link](https://arxiv.org/abs/2512.17102)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2025b\)Z\. Wang, K\. Wang, Q\. Wang, P\. Zhang, L\. Li, Z\. Yang, X\. Jin, K\. Yu, M\. N\. Nguyen, L\. Liu, E\. Gottlieb, Y\. Lu, K\. Cho, J\. Wu, L\. Fei\-Fei, L\. Wang, Y\. Choi, and M\. LiRAGEN: understanding self\-evolution in llm agents via multi\-turn reinforcement learning\.External Links:2504\.20073,[Link](https://arxiv.org/abs/2504.20073)Cited by:[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Weiet al\.\(2025\)Q\. Wei, S\. Zeng, C\. Li, W\. Brown, O\. Frunza, W\. Deng, A\. Schneider, Y\. Nevmyvaka, Y\. K\. Zhao, A\. Garcia, and M\. HongReinforcing multi\-turn reasoning in llm agents via turn\-level reward design\.External Links:2505\.11821,[Link](https://arxiv.org/abs/2505.11821)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2026\)R\. Wu, X\. Wang, J\. Mei, P\. Cai, D\. Fu, C\. Yang, L\. Wen, X\. Yang, Y\. Shen, Y\. Wang, and B\. ShiEvolveR: self\-evolving llm agents through an experience\-driven lifecycle\.External Links:2510\.16079,[Link](https://arxiv.org/abs/2510.16079)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix8.p1.1.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Xiaet al\.\(2026\)P\. Xia, J\. Chen, H\. Wang, J\. Liu, K\. Zeng, Y\. Wang, S\. Han, Y\. Zhou, X\. Zhao, H\. Chen, Z\. Zheng, C\. Xie, and H\. YaoSkillRL: evolving agents via recursive skill\-augmented reinforcement learning\.External Links:2602\.08234,[Link](https://arxiv.org/abs/2602.08234)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix10.p1.1.1),[§1](https://arxiv.org/html/2608.21946#S1.p2.1),[§3\.3](https://arxiv.org/html/2608.21946#S3.SS3.p2.1),[Table 1](https://arxiv.org/html/2608.21946#S3.T1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Yanet al\.\(2025\)J\. Yan, Y\. Li, Z\. Hu, Z\. Wang, G\. Cui, X\. Qu, Y\. Cheng, and Y\. ZhangLearning to reason under off\-policy guidance\.External Links:2504\.14945,[Link](https://arxiv.org/abs/2504.14945)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2024\)L\. Yang, Z\. Yu, T\. Zhang, S\. Cao, M\. Xu, W\. Zhang, J\. E\. Gonzalez, and B\. CuiBuffer of thoughts: thought\-augmented reasoning with large language models\.InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 \- 15, 2024,A\. Globersons, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. M\. Tomczak, and C\. Zhang \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2024/hash/cde328b7bf6358f5ebb91fe9c539745e-Abstract-Conference.html)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Yaoet al\.\(2022\)S\. Yao, H\. Chen, J\. Yang, and K\. NarasimhanWebShop: towards scalable real\-world web interaction with grounded language agents\.InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 \- December 9, 2022,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2022/hash/82ad13ec01f9fe44c01cb91814fd7b8c-Abstract-Conference.html)Cited by:[Appendix F](https://arxiv.org/html/2608.21946#A6.p1.1),[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§1](https://arxiv.org/html/2608.21946#S1.p4.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1\-5, 2023,External Links:[Link](https://openreview.net/forum?id=WE\_vluYUL-X)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix3.p1.1.1),[§1](https://arxiv.org/html/2608.21946#S1.p1.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1)\.
- Yeet al\.\(2026\)T\. Ye, L\. Dong, X\. Wu, S\. Huang, and F\. WeiOn\-policy context distillation for language models\.External Links:2602\.12275,[Link](https://arxiv.org/abs/2602.12275)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.
- Zhanet al\.\(2026\)R\. Zhan, Y\. Li, Z\. Wang, X\. Qu, D\. Liu, J\. Shao, D\. F\. Wong, and Y\. ChengExGRPO: learning to reason from experience\.External Links:2510\.02245,[Link](https://arxiv.org/abs/2510.02245)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2026a\)H\. Zhang, Q\. Long, J\. Bao, T\. Feng, W\. Zhang, H\. Yue, and W\. WangMemSkill: learning and evolving memory skills for self\-evolving agents\.External Links:2602\.02474,[Link](https://arxiv.org/abs/2602.02474)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025a\)K\. Zhang, A\. Lv, J\. Li, Y\. Wang, F\. Wang, H\. Hu, and R\. YanStepHint: multi\-level stepwise hints enhance reinforcement learning to reason\.External Links:2507\.02841,[Link](https://arxiv.org/abs/2507.02841)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2026b\)S\. Zhang, Y\. Xiong, X\. Chen, Z\. Jia, R\. Huang, J\. Xu, and J\. ZhangRAPO: expanding exploration for llm agents via retrieval\-augmented policy optimization\.External Links:2603\.03078,[Link](https://arxiv.org/abs/2603.03078)Cited by:[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025b\)Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin, F\. Huang, and J\. ZhouQwen3 embedding: advancing text embedding and reranking through foundation models\.External Links:2506\.05176,[Link](https://arxiv.org/abs/2506.05176)Cited by:[§B\.2](https://arxiv.org/html/2608.21946#A2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px3.p1.1)\.
- Zhaoet al\.\(2024\)A\. Zhao, D\. Huang, Q\. Xu, M\. Lin, Y\. Liu, and G\. HuangExpeL: LLM agents are experiential learners\.InThirty\-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty\-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20\-27, 2024, Vancouver, Canada,M\. J\. Wooldridge, J\. G\. Dy, and S\. Natarajan \(Eds\.\),pp\. 19632–19642\.External Links:[Link](https://doi.org/10.1609/aaai.v38i17.29936),[Document](https://dx.doi.org/10.1609/AAAI.V38I17.29936)Cited by:[§1](https://arxiv.org/html/2608.21946#S1.p2.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2026\)S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. GroverSelf\-distilled reasoner: on\-policy self\-distillation for large language models\.External Links:2601\.18734,[Link](https://arxiv.org/abs/2601.18734)Cited by:[item •](https://arxiv.org/html/2608.21946#A2.I1.ix6.p1.1.1),[§4\.1](https://arxiv.org/html/2608.21946#S4.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.21946#S5.SS0.SSS0.Px3.p1.1)\.

## Appendix ATheoretical Analysis ofEDGE

This section provides an analytical justification for the four core components ofEDGE, following the pipeline through which a retrieved experienceeeinfluences the policy update: its marginal utility is estimated \(§[A\.1](https://arxiv.org/html/2608.21946#A1.SS1)\), smoothed across training steps \(§[A\.2](https://arxiv.org/html/2608.21946#A1.SS2)\), used to gate and calibrate the RL advantage \(§[A\.3](https://arxiv.org/html/2608.21946#A1.SS3)\), and channeled through on\-support distillation \(§[A\.4](https://arxiv.org/html/2608.21946#A1.SS4)\)\.

Throughout, letc𝖲c^\{\\mathsf\{S\}\}denote the standard context,c𝖳=c𝖲⊕ec^\{\\mathsf\{T\}\}=c^\{\\mathsf\{S\}\}\\oplus ethe privileged context,πθ\\pi\_\{\\theta\}the current policy, andR⁡\(τ\)∈\{0,1\}R\(\\tau\)\\in\\\{0,1\\\}the binary outcome reward\. Teacher and student rollout subsets𝒯𝖳,𝒯𝖲\\mathcal\{T\}^\{\\mathsf\{T\}\},\\mathcal\{T\}^\{\\mathsf\{S\}\}each have sizeK=G/2K=G/2\. All expectations and probabilities condition on fixedπθ\\pi\_\{\\theta\}andee\.

### A\.1Marginal\-Gain Estimation and Gating

Define the*value of privileged information*as the expected return gap between the privileged and standard contexts:

𝒱e\(θ\)=𝔼τ∼πθ\(⋅∣c𝖳\)\[R\(τ\)\]−𝔼τ∼πθ\(⋅∣c𝖲\)\[R\(τ\)\]\.\\mathcal\{V\}\_\{e\}\(\\theta\)\\;=\\;\\mathbb\{E\}\_\{\\tau\\sim\\pi\_\{\\theta\}\(\\cdot\\mid c^\{\\mathsf\{T\}\}\)\}\\\!\\bigl\[R\(\\tau\)\\bigr\]\\;\-\\;\\mathbb\{E\}\_\{\\tau\\sim\\pi\_\{\\theta\}\(\\cdot\\mid c^\{\\mathsf\{S\}\}\)\}\\\!\\bigl\[R\(\\tau\)\\bigr\]\.\(10\)The empirical marginal gain is

Δe=1K​∑τ∈𝒯𝖳R⁡\(τ\)−1K​∑τ∈𝒯𝖲R⁡\(τ\)\.\\Delta\_\{e\}\\;=\\;\\frac\{1\}\{K\}\\\!\\sum\_\{\\tau\\in\\mathcal\{T\}^\{\\mathsf\{T\}\}\}\\\!R\(\\tau\)\\;\-\\;\\frac\{1\}\{K\}\\\!\\sum\_\{\\tau\\in\\mathcal\{T\}^\{\\mathsf\{S\}\}\}\\\!R\(\\tau\)\.\(11\)Under the assumption that the two subsets are conditionally independent and identically distributed up to the presence ofee,Δe\\Delta\_\{e\}is unbiased for𝒱e​\(θ\)\\mathcal\{V\}\_\{e\}\(\\theta\)with variance\(pT​\(1−pT\)\+pS​\(1−pS\)\)/K\(p\_\{T\}\(1\{\-\}p\_\{T\}\)\+p\_\{S\}\(1\{\-\}p\_\{S\}\)\)/K, wherepTp\_\{T\}andpSp\_\{S\}are the respective success probabilities\. DecomposingΔe\\Delta\_\{e\}as a sum of2​K2Kindependent terms each bounded in an interval of length1/K1/K, Hoeffding’s inequality yields

Pr⁡\(\|Δe−𝒱e​\(θ\)\|≥ϵ\)≤2​exp⁡\(−K​ϵ2\)\.\\Pr\\\!\\bigl\(\|\\Delta\_\{e\}\-\\mathcal\{V\}\_\{e\}\(\\theta\)\|\\geq\\epsilon\\bigr\)\\;\\leq\\;2\\exp\\\!\\bigl\(\-K\\epsilon^\{2\}\\bigr\)\.\(12\)
EDGEconverts this estimate into a binary gate:

Me=𝕀⁡\(Δe\>0\)\.M\_\{e\}\\;=\\;\\mathbb\{I\}\(\\Delta\_\{e\}\>0\)\.\(13\)The gate serves as a probabilistic risk\-control mechanism\. By the one\-sided form of \([12](https://arxiv.org/html/2608.21946#A1.E12)\), if an experience is truly harmful with marginγ\>0\\gamma\>0\(𝒱e​\(θ\)≤−γ\\mathcal\{V\}\_\{e\}\(\\theta\)\\leq\-\\gamma\), thenPr⁡\(Me=1\)≤exp⁡\(−K​γ2\)\\Pr\(M\_\{e\}\{=\}1\)\\leq\\exp\(\-K\\gamma^\{2\}\); the symmetric bound holds for missed activations when𝒱e​\(θ\)≥γ\\mathcal\{V\}\_\{e\}\(\\theta\)\\geq\\gamma\. WhenMe=0M\_\{e\}=0, teacher rollouts are masked from both the RL and distillation losses, and the update falls back to standard GRPO on the student subset\.

### A\.2EMA Smoothing

BecauseΔe\\Delta\_\{e\}is high\-variance for smallKKand the true utility𝒱e​\(θ\)\\mathcal\{V\}\_\{e\}\(\\theta\)drifts as the policy evolves,EDGEtracks each experience’s utility with an exponential moving average:

Ue\(t\)=\(1−μ\)​Ue\(t−1\)\+μ​Δe\(t\),μ∈\(0,1\)\.U\_\{e\}^\{\(t\)\}\\;=\\;\(1\-\\mu\)\\,U\_\{e\}^\{\(t\-1\)\}\\;\+\\;\\mu\\,\\Delta\_\{e\}^\{\(t\)\},\\qquad\\mu\\in\(0,1\)\.\(14\)Under a local\-stationarity approximation \(successiveΔe\(t\)\\Delta\_\{e\}^\{\(t\)\}treated as i\.i\.d\. with varianceσe2\\sigma\_\{e\}^\{2\}\), the steady\-state variance isVar⁡\(Ue\)=μ2−μ​σe2\\mathrm\{Var\}\(U\_\{e\}\)=\\frac\{\\mu\}\{2\-\\mu\}\\,\\sigma\_\{e\}^\{2\}, which is strictly less thanσe2\\sigma\_\{e\}^\{2\}for anyμ<1\\mu<1\. Smallerμ\\muyields greater smoothing at the cost of slower adaptation to genuine utility shifts\. The pruning thresholdη\\etathus operates on a smoothed estimate of recent marginal utility rather than a single noisy contrast\.

### A\.3Pooled Advantage Calibration

When the gate activates, teacher rollouts enter the RL update and alter the advantage baseline\. LetbS,bTb\_\{S\},b\_\{T\}be the student and teacher mean returns, soΔe=bT−bS\\Delta\_\{e\}=b\_\{T\}\-b\_\{S\}\. The active set is

𝒜=\{𝒯𝖲∪𝒯𝖳,Me=1,𝒯𝖲,Me=0,\\mathcal\{A\}\\;=\\;\\begin\{cases\}\\mathcal\{T\}^\{\\mathsf\{S\}\}\\cup\\mathcal\{T\}^\{\\mathsf\{T\}\},&M\_\{e\}=1,\\\\\[2\.0pt\] \\mathcal\{T\}^\{\\mathsf\{S\}\},&M\_\{e\}=0,\\end\{cases\}\(15\)with baselineb𝒜=meanj∈𝒜​\(Rj\)b\_\{\\mathcal\{A\}\}=\\mathrm\{mean\}\_\{j\\in\\mathcal\{A\}\}\(R\_\{j\}\)\. WhenMe=1M\_\{e\}=1, the pooled baselinebpool=12​\(bS\+bT\)b\_\{\\mathrm\{pool\}\}=\\frac\{1\}\{2\}\(b\_\{S\}\+b\_\{T\}\)shifts the raw advantage numerators \(prior to GRPO’s standard\-deviation normalization, which rescales uniformly without changing signs\) as follows:

A~SEDGE\\displaystyle\\widetilde\{A\}\_\{S\}^\{\\textsc\{EDGE\}\{\}\}=A~Swithin−12​Δe,\\displaystyle=\\widetilde\{A\}\_\{S\}^\{\\mathrm\{within\}\}\-\\tfrac\{1\}\{2\}\\Delta\_\{e\},\(16\)A~TEDGE\\displaystyle\\widetilde\{A\}\_\{T\}^\{\\textsc\{EDGE\}\{\}\}=A~Twithin\+12​Δe,\\displaystyle=\\widetilde\{A\}\_\{T\}^\{\\mathrm\{within\}\}\+\\tfrac\{1\}\{2\}\\Delta\_\{e\},\(17\)where the superscript “within” denotes baselines computed from each subset alone\. SinceMe=1M\_\{e\}=1impliesΔe\>0\\Delta\_\{e\}\>0, student advantages are shifted*downward*and teacher advantages*upward*: privileged successes receive stronger reinforcement, while unguided successes are tempered when the scaffold demonstrates superior performance\. WhenMe=0M\_\{e\}=0, the teacher subset is excluded and no shift is applied\.

### A\.4On\-Support Reverse\-KL Distillation

Directly imitating teacher\-generated trajectories risks covariate shift\([22](https://arxiv.org/html/2608.21946#bib.bib40)\), as the teacher may visit prefixes whose rationale depends on information absent from the student context \(causal misidentification;[7](https://arxiv.org/html/2608.21946#bib.bib41)\)\.EDGEmitigates this by evaluating the teacher only on student\-generated prefixes\. For a student trajectoryτ=\(y1,…,ym\)∈𝒯𝖲\\tau=\(y\_\{1\},\\ldots,y\_\{m\}\)\\in\\mathcal\{T\}^\{\\mathsf\{S\}\}, letht𝖲=\(c𝖲,y<t\)h\_\{t\}^\{\\mathsf\{S\}\}=\(c^\{\\mathsf\{S\}\},y\_\{<t\}\)andht𝖳=\(c𝖳,y<t\)h\_\{t\}^\{\\mathsf\{T\}\}=\(c^\{\\mathsf\{T\}\},y\_\{<t\}\)\. The on\-support distillation loss is

ℒ^distill​\(θ\)=Me∑τ∈𝒯𝖲∑t=1\|τ\|DKL\(πθ\(⋅∣ht𝖲\)∥πsg⁡\(θ\)\(⋅∣ht𝖳\)\),\\begin\{split\}\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{distill\}\}\(\\theta\)\\;=\\;\\;&M\_\{e\}\\\!\\sum\_\{\\tau\\in\\mathcal\{T\}^\{\\mathsf\{S\}\}\}\\sum\_\{t=1\}^\{\|\\tau\|\}\\\\ &D\_\{\\mathrm\{KL\}\}\\\!\\bigl\(\\pi\_\{\\theta\}\(\\cdot\\mid h\_\{t\}^\{\\mathsf\{S\}\}\)\\;\\big\\\|\\;\\pi\_\{\\mathrm\{sg\}\(\\theta\)\}\(\\cdot\\mid h\_\{t\}^\{\\mathsf\{T\}\}\)\\bigr\),\\end\{split\}\(18\)wheresg⁡\(⋅\)\\mathrm\{sg\}\(\\cdot\)denotes stop\-gradient\. Unlike forward\-KL imitation on teacher rollouts, this objective compares distributions only at prefixes the student has actually reached, avoiding training on teacher\-only states\. The reverse\-KL direction is mode\-seeking with respect to the teacher, concentrating student mass on teacher\-preferred actions rather than spreading to cover the full teacher distribution, and thereby keeping the update conservative on the student support\.

The final actor loss combines both pathways:

ℒactor​\(θ\)=ℒRL​\(θ,𝒜\)\+λ​ℒ^distill​\(θ\)\.\\mathcal\{L\}\_\{\\mathrm\{actor\}\}\(\\theta\)\\;=\\;\\mathcal\{L\}\_\{\\mathrm\{RL\}\}\(\\theta;\\mathcal\{A\}\)\\;\+\\;\\lambda\\,\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{distill\}\}\(\\theta\)\.\(19\)WhenMe=0M\_\{e\}=0, both terms reduce to unprivileged\-only updates \(𝒜=𝒯𝖲\\mathcal\{A\}=\\mathcal\{T\}^\{\\mathsf\{S\}\},ℒ^distill=0\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{distill\}\}=0\)\. WhenMe=1M\_\{e\}=1, the teacher subset raises the RL baseline \(§[A\.3](https://arxiv.org/html/2608.21946#A1.SS3)\) and the reverse KL transfers privileged behavior into the standard\-context policy\.

## Appendix BImplementation Details

### B\.1Baselines

- •GPT\-4o\([17](https://arxiv.org/html/2608.21946#bib.bib20)\): A closed\-source multimodal LLM from OpenAI, used as a strong proprietary agent baseline with standard prompting\.
- •Gemini\-2\.5\-Pro\([6](https://arxiv.org/html/2608.21946#bib.bib21)\): A closed\-source reasoning model from Google DeepMind, serving as another proprietary agent baseline with standard prompting\.
- •ReAct\([38](https://arxiv.org/html/2608.21946#bib.bib12)\): A prompting framework that interleaves chain\-of\-thought reasoning with environment actions, enabling LLMs to plan and act in a synergistic loop\.
- •Reflexion\([25](https://arxiv.org/html/2608.21946#bib.bib13)\): Extends ReAct by appending verbal self\-reflection after task failures, allowing the agent to refine its strategy across successive trials without weight updates\.
- •GRPO\([24](https://arxiv.org/html/2608.21946#bib.bib7)\): A group\-relative policy optimization algorithm that estimates advantages from a group of sampled rollouts, eliminating the need for a separate critic network\.
- •OPSD\([46](https://arxiv.org/html/2608.21946#bib.bib30)\): An on\-policy self\-distillation method that distills the model’s own high\-quality rollouts back into itself to improve reasoning without external supervision\.
- •GRPO\+OPSD: A hybrid baseline that combines GRPO’s group\-relative advantage estimation with OPSD’s on\-policy self\-distillation objective\.
- •EvolveR\([33](https://arxiv.org/html/2608.21946#bib.bib15)\): An experience\-augmented training method that iteratively evolves a retrieval\-augmented memory of past trajectories to guide policy learning\.
- •GRPO\+Mem0\([4](https://arxiv.org/html/2608.21946#bib.bib17)\): Augments GRPO with Mem0, a memory module that stores and retrieves past interaction experiences as additional context during both training and inference\.
- •SkillRL\([34](https://arxiv.org/html/2608.21946#bib.bib18)\): A skill\-based RL framework that extracts reusable skills from successful trajectories and conditions policy optimization on retrieved skill demonstrations\.
- •EMPO2\([13](https://arxiv.org/html/2608.21946#bib.bib19)\): An exploratory memory\-augmented policy optimization method that leverages curated past experiences to enhance exploration during RL training\.

These baselines span closed\-source LLM agents, prompting\-based reasoning frameworks, standard RL algorithms, self\-distillation methods, and experience\-augmented training approaches, enabling a comprehensive evaluation ofEDGEfrom multiple perspectives\.

### B\.2RL\-Training Configuration

We implementEDGEon top of the verl\-agent framework\([9](https://arxiv.org/html/2608.21946#bib.bib1)\)and train the model with the joint optimization objective in Eq\.[19](https://arxiv.org/html/2608.21946#A1.E19)\. To reduce the computational cost of reverse\-KL distillation, we approximate the vocabulary\-level loss using only the top\-kkstudent tokens with the highest log probabilities\. The experience bank is initialized as empty, and each newly inserted experience is assigned an initial utility score of 0\. For retrieval, we use the task instruction and initial environment observations as the query, encode experiences with Qwen3\-Embedding\-0\.6B\([44](https://arxiv.org/html/2608.21946#bib.bib42)\), and rank candidates by cosine similarity\. We retrieve a top\-mmcandidate pool and use the highest\-scoring experience as the scaffold for the current rollout\.

After each training step, we update the experience bank through both expansion and pruning\. For expansion, if the success rate of a task category falls below the expansion thresholdξ\\xi, GPT\-4o\([17](https://arxiv.org/html/2608.21946#bib.bib20)\)is used as the reflector LLM to contrast successful and failed trajectories from the same category and synthesize up to three new experiences\. For pruning, we update the EMA utility score of each experience according to Eq\.[14](https://arxiv.org/html/2608.21946#A1.E14)and remove experiences whose utility falls below the pruning thresholdη\\eta\. For both ALFWorld and WebShop, we use the hyperparameters in Table[3](https://arxiv.org/html/2608.21946#A2.T3)\. All training experiments are conducted on8×808\\times 80GB GPUs\.

ConfigurationValueRL\-TrainingActor learning rate1​e−61\\text\{e\}^\{\-6\}Maximum prompt length4096Maximum response length512Training batch size16Rollout group size \(GG\)8Training mini\-batch size128Rollout temperature1\.0Training steps200EDGE configurationDistillation weightλ\\lambda0\.1Distillation top\-kktokens20Utility pruning thresholdη\\eta\-0\.1Retrieval pool size \(top\-mm\)6Expansion success thresholdξ\\xi0\.4EMA momentumμ\\mu0\.5Maximum new experiences per step3Table 3:RL HyperparametersTable 4:Experience\-bank co\-evolution examples\.Each case is extracted from saved failure trajectories, LLM reflection logs, retrieval logs, and utility traces\.

## Appendix CFurther Analysis

Figure 5:Sensitivity of validation success rate to distillation weightλ\\lambda\.### C\.1Sensitivity to Distillation Weightλ\\lambda\.

Figure[5](https://arxiv.org/html/2608.21946#A3.F5)sweeps the distillation coefficientλ\\lambdaon Qwen2\.5\-1\.5B\-Instruct, revealing a clear trade\-off between the RL and distillation objectives\. Atλ=1\\lambda=1, the distillation term dominates the gradient and effectively freezes the policy near its initial performance \(∼\\sim15%\), preventing autonomous exploration beyond the scaffold\-prescribed behavior\. Atλ=0\.01\\lambda=0\.01, the policy recovers RL\-driven improvement but internalizes scaffold\-induced patterns too slowly, converging roughly 5 points below the best setting\. The moderate valueλ=0\.1\\lambda=0\.1balances both pressures, sustaining steady improvement to approximately 80%—enough distillation to accelerate internalization without suppressing the RL objective’s exploratory signal\. We adopt this value for all remaining experiments\.

### C\.2Case Study: Experience Bank Evolution Dynamics

This section complements the aggregate gain\-tracking curves in Figure[4](https://arxiv.org/html/2608.21946#S4.F4)with entry\-level evidence from ALFWorld training logs \(Qwen2\.5\-7B\-Instruct\)\. We trace four representative cases \(Table[4](https://arxiv.org/html/2608.21946#A2.T4)\), each linking a rollout failure to the experience it produces, its later retrieval, and the resulting change in agent behavior\. Cases 1–2 \(Figure[6](https://arxiv.org/html/2608.21946#A3.F6)\) illustrate the expansion phase, while Cases 3–4 \(Figure[7](https://arxiv.org/html/2608.21946#A3.F7)\) illustrate late\-stage refinement and non\-stationary utility\.

Case 1: From destination\-first wandering to object\-first executionFailure \(step 150\)\.The task isput a clean bowl in shelf\. The failed rollout immediately navigates to the destination and loops around empty shelves:go to shelf 3→\\rightarrowgo to cabinet 1→\\rightarrowgo to shelf 3→\\rightarrowgo to cabinet 3→\\rightarrowexamine shelf 3\.Successful contrast\.The agent first searches likely sources, finds a bowl in the fridge, takes and cleans it, then navigates to a shelf:go to fridge 1→\\rightarrowopen fridge 1→\\rightarrowtake bowl 1 from fridge 1→\\rightarrowclean bowl 1 with sinkbasin 1→\\rightarrowgo to shelf 1\.Reflected experience\.Entrytask\_943:*“First identify where the target bowl is, retrieve it, clean it if the goal requires a clean bowl, and only then move it to a shelf\.”*Later retrieval\.At step 155, retrieved forput a clean bowl in diningtable\(similarity0\.8860\.886\)\. The rollout no longer starts by visiting the table; it locates, cleans, and places the bowl\. Utility:U=0\.125U\{=\}0\.125\.

Case 2: From hallucinated source to source verificationFailure \(step 58\)\.The task isput some pillow on sofa\. The agent goes to the sofa, observesbox,creditcard,keychain, andnewspaper—but no pillow\. It issuestake pillow from sofa 1, receivesNothing happens, and repeats similar invalid source assumptions\.Successful contrast\.The paired trajectory searches alternative sources, reachesarmchair 1, observespillow 1, takes it, and moves it to the sofa—only issuingtakewhen the observation lists the target\.Reflected experience\.Entrystep\_1067:*“Only usetake <object\> from <container\>when the object is actually listed at that location; if not observed, do not repeat the same invalid action\.”*Later retrieval\.At step 60, retrieved forput some keychain on ottoman, where the observation lacks the keychain—the same structural error\. Utility evolves from0\.0000\.000at creation to0\.5780\.578at step 200 after 78 retrievals\.

Figure 6:Experience bank evolution: early\-stage expansion\.Case 1 illustrates destination\-first wandering corrected by an object\-first scaffold; Case 2 illustrates hallucinated source actions corrected by an observation\-gated rule\.Case 3: From unstructured Pick2 search to systematic collectionFailure \(step 165\)\.Forfind two ladle and put them in drawer, the policy spends actions on destination drawers or repeated navigation before confirming where the two targets are\. Unlike early scaffolds, this failure involves managing object count, source search, and final placement simultaneously\.Reflected experience\.Entrytask\_995:*“Identify likely locations of the target item, inspect those locations methodically, pick up each required object, and then place them into the specified container\. Avoid random movement or opening unrelated containers before locating the targets\.”*Later retrieval\.At step 166, retrieved forfind two spatula and put them in drawer\(similarity0\.8540\.854\)\. The rollout locatesspatula 3on a countertop, picks it up, checks other locations, and deposits it in a drawer\. Utility reachesU=0\.375U\{=\}0\.375by steps 195–200\.

Case 4: Utility tracking separates useful retrieval from harmful repetitionCreation \(step 144\)\.Entrystep\_914,*“Avoid repeating a no\-op movement,”*is reflected from afind two tissuebox and put them in drawerfailure\. The rule: if the agent is already at the target location, it should inspect or act rather than re\-issuing the same navigation action\.Non\-stationary utility\.The entry is retrieved frequently but its utility is initially unstable:U=−0\.026U\{=\}\{\-\}0\.026at step 150 and−0\.085\{\-\}0\.085at step 165—hovering near the pruning thresholdη=−0\.1\\eta\{=\}\{\-\}0\.1but remaining above it\. Frequency\-based retention would treat this entry as important despite its near\-zero or negative marginal gain; conversely, a positive threshold would have discarded it prematurely\.Recovery\.As the policy reaches more states where repeated no\-op movements are the dominant error, the entry becomes useful: utility turns positive atU=0\.048U\{=\}0\.048by step 180 and rises toU=0\.452U\{=\}0\.452at step 200 after 142 retrievals\. This illustrates whyEDGEcombines smoothed EMA tracking \(Eq\. \([14](https://arxiv.org/html/2608.21946#A1.E14)\)\) with a mildly negative pruning threshold: experience utility can shift as the policy distribution changes, and premature pruning would discard entries whose value has not yet materialized\.

Figure 7:Experience bank evolution: late\-stage refinement and non\-stationary utility\.Case 3 illustrates specialized multi\-object coordination guidance; Case 4 illustrates an entry whose utility is initially near the pruning threshold but recovers as the policy’s failure distribution shifts\.The four cases reveal a natural curriculum driven by the policy’s evolving failure distribution\. Early failures are structural and broadly shared: destination\-first wandering \(Case 1\) and hallucinated object presence \(Case 2\) affect many task variants, producing generic scaffolds with sustained high utility—these entries drive the initial bank growth visible in Figure[4](https://arxiv.org/html/2608.21946#S4.F4)\. As the policy internalizes these broad strategies, the remaining failures narrow in scope\. By step 165, single\-object manipulation is reliable but coordinating two targets remains fragile; Case 3’s search\-then\-collect scaffold addresses precisely this gap, and its moderate final utility \(U=0\.375U\{=\}0\.375\) is consistent with Pick2 remaining the hardest subtask in Table[2](https://arxiv.org/html/2608.21946#S4.T2)\.

Case 4 provides the most direct evidence for non\-stationary utility\. Entrystep\_914\(“avoid repeating no\-op movements”\) hovers near the pruning boundary for roughly 20 steps \(U=−0\.026U\{=\}\{\-\}0\.026at step 150,−0\.085\{\-\}0\.085at step 165\) before recovering toU=0\.452U\{=\}0\.452by step 200\. The mechanism is interpretable: once destination\-first errors \(Case 1\) are resolved, no\-op navigation loops become the dominant Pick2 failure mode, re\-activating a previously marginal entry\. A positive pruning threshold would have discarded it; the adoptedη=−0\.1\\eta\{=\}\{\-\}0\.1retains such latent\-value entries until the policy’s distribution shifts in their favor\.

Together, these cases explain the aggregate bank trajectory in Figure[4](https://arxiv.org/html/2608.21946#S4.F4)—initial growth as generic scaffolds accumulate, followed by contraction as the policy absorbs broad patterns and the bank converges to a compact set of specialized, currently relevant entries\.

## Appendix DPseudocode

Algorithm[1](https://arxiv.org/html/2608.21946#alg1)outlines the full EDGE training loop, which alternates among three stages per iteration: experience\-guided exploration scaffolding \(§[3\.1](https://arxiv.org/html/2608.21946#S3.SS1)\), gain\-gated privileged distillation \(§[3\.2](https://arxiv.org/html/2608.21946#S3.SS2)\), and experience bank evolution \(§[3\.3](https://arxiv.org/html/2608.21946#S3.SS3)\)\.

Algorithm 1EDGE: Experience\-Distillation for Guided Exploration1:

2:Policy

πθ\\pi\_\{\\theta\}; experience bank

ℰ\\mathcal\{E\};

3:Rollout group size

GG; distillation weight

λ\\lambda;

4:EMA momentum

μ\\mu; pruning threshold

η\\eta;

5:Success\-rate threshold

ξ\\xi\.

6:foreach training iterationdo

7:foreach task

xxin batchdo

8:Stage 1: Experience\-Guided Exploration Scaffolding

9:Retrieve top experience

e∈ℰe\\in\\mathcal\{E\}by embedding similarity\.

10:Sample

GGtrajectories from

πθ\\pi\_\{\\theta\}:

11:

𝒯𝖳\\mathcal\{T\}^\{\\mathsf\{T\}\}\(

G/2G/2\): conditioned on

c𝖳=x⊕ec^\{\\mathsf\{T\}\}=x\\oplus e
12:

𝒯𝖲\\mathcal\{T\}^\{\\mathsf\{S\}\}\(

G/2G/2\): conditioned on

c𝖲=xc^\{\\mathsf\{S\}\}=x
13:Compute marginal gain

Δe\\Delta\_\{e\}via Eq\. \([5](https://arxiv.org/html/2608.21946#S3.E5)\)\.

14:Stage 2: Gain\-Gated Privileged Distillation

15:if

Δe\>0\\Delta\_\{e\}\>0then

16:Compute advantages over all

GGtrajectories\.

17:foreach

τ∈𝒯𝖲\\tau\\in\\mathcal\{T\}^\{\\mathsf\{S\}\}do

18:Forward

πsg⁡\(θ\)\\pi\_\{\\mathrm\{sg\}\(\\theta\)\}on

τ\\tauwith

c𝖳c^\{\\mathsf\{T\}\}\.

19:Compute

ℒdistill\\mathcal\{L\}\_\{\\text\{distill\}\}via Eq\. \([18](https://arxiv.org/html/2608.21946#A1.E18)\)\.

20:endfor

21:else

22:Compute advantages over

𝒯𝖲\\mathcal\{T\}^\{\\mathsf\{S\}\}only\.

23:

ℒdistill←0\\mathcal\{L\}\_\{\\text\{distill\}\}\\leftarrow 0\.

24:endif

25:

ℒactor←ℒRL\+λ​ℒdistill\\mathcal\{L\}\_\{\\text\{actor\}\}\\leftarrow\\mathcal\{L\}\_\{\\text\{RL\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\text\{distill\}\}\.

26:Stage 3: Experience Bank Evolution

27:Update EMA utility:

Ue\(t\)←\(−μ\)​Ue\(t−1\)\+μ​ΔeU\_\{e\}^\{\(t\)\}\\\!\\leftarrow\\\!\(1\\\!\-\\\!\\mu\)\\,U\_\{e\}^\{\(t\-1\)\}\+\\mu\\,\\Delta\_\{e\}\.

28:endfor

29:Update

πθ\\pi\_\{\\theta\}with

ℒactor\\mathcal\{L\}\_\{\\text\{actor\}\}\.

30:Prune experiences with

Ue\(t\)<ηU\_\{e\}^\{\(t\)\}<\\etafrom

ℰ\\mathcal\{E\}\.

31:Identify task categories with success rate

<ξ<\\xi\.

32:Generate new experiences via

freflect​\(τ\+,τ−\)f\_\{\\text\{reflect\}\}\(\\tau^\{\+\},\\tau^\{\-\}\); insert into

ℰ\\mathcal\{E\}\.

33:endfor

34:Trained policy

πθ\\pi\_\{\\theta\}\(deployed without

ℰ\\mathcal\{E\}\)\.

## Appendix EPrompts

This section provides the prompt templates used in our experiments\. We include both rollout prompts and experience update prompts for ALFWorld and WebShop\. The rollout prompts are used for action generation, while the update prompts are used to generate new state\-aware experiences from contrasted trajectories\.

### E\.1ALFWorld

Figure[8](https://arxiv.org/html/2608.21946#A5.F8)shows the ALFWorld rollout prompts\. The first prompt provides retrieved experiences as additional context, while the second prompt removes external experiences and asks the agent to act only based on the current task, history, observation, and admissible actions\.

Prompt: ALFWorld Agent Execution with ExperienceSystem Prompt:You are an expert agent operating in the ALFRED Embodied Environment\. Your task is to:\{task\_description\}\# Retrieved Relevant Experience\{retrieved\_experiences\}Warning: These experiences may be outdated\. Use them only if they align with your current observation\.\# Current ProgressPrior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}Your admissible actions of the current situation are: \[\{admissible\_actions\}\]\.Now it is your turn to take an action\. You should first reason step\-by\-step about the current situation\. This reasoning processMUSTbe enclosed within<think\></think\>tags\. Once you have finished your reasoning, you should choose an admissible action for the current step and present it within<action\></action\>tags\.

Prompt: ALFWorld Agent Execution without ExperienceSystem Prompt:You are an expert agent operating in the ALFRED Embodied Environment\. Your task is to:\{task\_description\}\# Retrieved Relevant ExperienceNo external general and task\-specific experiences are provided\. Use your learned strategy and current observation\.\# Current ProgressPrior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}Your admissible actions of the current situation are: \[\{admissible\_actions\}\]\.Now it is your turn to take an action\. You should first reason step\-by\-step about the current situation\. This reasoning processMUSTbe enclosed within<think\></think\>tags\. Once you have finished your reasoning, you should choose an admissible action for the current step and present it within<action\></action\>tags\.

Figure 8:ALFWorld rollout prompts with and without retrieved experience\.Figure[9](https://arxiv.org/html/2608.21946#A5.F9)shows the prompt used to update the ALFWorld experience bank\. Given a failed trajectory and a successful reference trajectory for the same task, the model generates compact state\-aware experiences in JSON format\.

Prompt: ALFWorld Experience UpdateSystem Prompt:You are an expert updating an ALFWorld state\-aware experience bank\. You are given one failed trajectory and one successful trajectory for the same task\. Analyze the contrasted rollout snippets below and proposeNEW or revised experiencesthat help the agent act better in the same environment state\.\# Task InformationTask:\{task\_text\}Task Type:\{task\_type\}The task was\{successfully/unsuccessfully\}completed\.\# Contrasted Rollout SnippetsFailed trajectory:\{failed\_text\}Successful trajectory \(reference\):\{success\_text\}\# Requirements•Each experience must bestate\-aware, not a generic tip\.•Focus on what the successful no\-experience rollout did, and what the retrieved experience may have caused the agent to do incorrectly\.•Prefer compact trigger language that can be matched at retrieval time\.•Use actionable wording tied to admissible actions and visible state cues\.•Return only JSON\.Generate 1–\{max\_new\_skills\_per\_update\}new experiences\.\# Output Format Example[⬇](data:text/plain;base64,ewogICJ0aXRsZSI6ICJQbGFuIG9iamVjdCBsb2NhdGlvbiBiZWZvcmUgYWN0aW5nIiwKICAicHJpbmNpcGxlIjogIkZvciB0aGlzIHRhc2sgdHlwZSwgZmlyc3QgaWRlbnRpZnkgd2hlcmUgdGhlIG9iamVjdCBpcywgdGhlbiBwbGFuIHRoZSBzZXF1ZW5jZSBvZiBhY3Rpb25zLiIsCiAgIndoZW5fdG9fYXBwbHkiOiAiV2hlbiB0aGUgdGFzayBpbnZvbHZlcyBmaW5kaW5nIG9yIG1vdmluZyBhIHNwZWNpZmljIG9iamVjdCIKfQ==)\{"title":"Planobjectlocationbeforeacting","principle":"Forthistasktype,firstidentifywheretheobjectis,thenplanthesequenceofactions\.","when\_to\_apply":"Whenthetaskinvolvesfindingormovingaspecificobject"\}Figure 9:Prompt for updating the ALFWorld state\-aware experience bank from contrasted failed and successful trajectories\. The model is instructed to generate compact, state\-aware experiences that can guide future retrieval and decision making\.
### E\.2WebShop

Figure[10](https://arxiv.org/html/2608.21946#A5.F10)shows the WebShop rollout prompts\. The experience\-conditioned prompt includes retrieved memories, while the experience\-free prompt relies only on the shopping instruction, interaction history, current observation, and admissible actions\.

Prompt: WebShop Agent Execution with ExperienceSystem Prompt:You are an expert autonomous agent operating in the WebShop e\-commerce environment\. Your task is to:\{task\_description\}\.\# Retrieved Relevant Experience\{retrieved\_memories\}Warning: These experiences may be outdated\. Use them only if they align with your current observation\.\# Current ProgressPrior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}\.Your admissible actions of the current situation are:\{available\_actions\}Now it is your turn to take one action for the current step\. You should first reason step\-by\-step about the current situation, then think carefully which admissible action best advances the shopping goal\. This reasoning processMUSTbe enclosed within<think\></think\>tags\. Once you have finished your reasoning, you should choose an admissible action for current step and present it within<action\></action\>tags\.

Prompt: WebShop Agent Execution without ExperienceSystem Prompt:You are an expert autonomous agent operating in the WebShop e\-commerce environment\. Your task is to:\{task\_description\}\.\# Retrieved Relevant ExperienceNo external general and task\-specific experiences are provided\. Use your learned strategy and current observation\.\# Current ProgressPrior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}\.Your admissible actions of the current situation are:\{available\_actions\}Now it is your turn to take one action for the current step\. You should first reason step\-by\-step about the current situation, then think carefully which admissible action best advances the shopping goal\. This reasoning processMUSTbe enclosed within<think\></think\>tags\. Once you have finished your reasoning, you should choose an admissible action for current step and present it within<action\></action\>tags\.

Figure 10:WebShop rollout prompts with and without retrieved experience\.Figure[11](https://arxiv.org/html/2608.21946#A5.F11)shows the prompt used to update the WebShop experience bank\. The model compares failed and successful shopping trajectories and produces new state\-aware experiences tied to visible state cues and available actions\.

Prompt: WebShop Experience UpdateSystem Prompt:You are an expert updating a WebShop shopping state\-aware experience bank\. You are given one failed trajectory and one successful trajectory for the same task\. Analyze the contrasted rollout snippets below and proposeNEW or revised experiencesthat help the agent act better in the same environment state\.\# Task InformationTask:\{task\_text\}Task Type:\{task\_type\}The task was\{successfully/unsuccessfully\}completed\.\# Contrasted Rollout SnippetsFailed trajectory:\{failed\_text\}Successful trajectory \(reference\):\{success\_text\}\# Requirements•Each experience must bestate\-aware, not a generic tip\.•Focus on what the successful no\-experience rollout did, and what the retrieved experience may have caused the agent to do incorrectly\.•Prefer compact trigger language that can be matched at retrieval time\.•Use actionable wording tied to available actions and visible state cues\.•Return only JSON\.Generate 1–\{max\_new\_skills\_per\_update\}new experiences\.\# Output Format Example[⬇](data:text/plain;base64,ewogICJ0aXRsZSI6ICJWZXJpZnkgRWFybHksIEFib3J0IEZhc3QiLAogICJwcmluY2lwbGUiOiAiT24gdGhlIHByb2R1Y3QgcGFnZSwgaW1tZWRpYXRlbHkgY2hlY2sgY2F0ZWdvcnksIGNvcmUgYXR0cmlidXRlcywgYW5kIHByaWNlOyBpZiBhIGtleSBjb25zdHJhaW50IGlzIHZpb2xhdGVkLCBsZWF2ZSB0aGUgcGFnZSBhdCBvbmNlLiIsCiAgIndoZW5fdG9fYXBwbHkiOiAiV2l0aGluIHRoZSBmaXJzdCBvYnNlcnZhdGlvbiBvbiBldmVyeSBwcm9kdWN0IGRldGFpbCBwYWdlLiIKfQ==)\{"title":"VerifyEarly,AbortFast","principle":"Ontheproductpage,immediatelycheckcategory,coreattributes,andprice;ifakeyconstraintisviolated,leavethepageatonce\.","when\_to\_apply":"Withinthefirstobservationoneveryproductdetailpage\."\}Figure 11:Prompt for updating the WebShop shopping state\-aware experience bank from contrasted failed and successful trajectories\. The model is instructed to generate compact, state\-aware experiences that can guide future retrieval and shopping decisions\.

## Appendix FDataset License

Our experiments are based on the publicly available ALFWorld\([26](https://arxiv.org/html/2608.21946#bib.bib6)\)and WebShop\([37](https://arxiv.org/html/2608.21946#bib.bib5)\)environments\. Training, evaluation, and experience\-bank construction are performed using trajectories generated from the task instances and interaction interfaces provided by these environments\. We strictly follow the licenses and usage terms of ALFWorld and WebShop, and use the resulting data only for academic research purposes\. No private, personally identifiable, or proprietary user data is used in any stage of training, evaluation, or experience construction\.

## Appendix GLLMs Usage Statement

We employed a Large Language Model \(LLM\) to assist exclusively in the editorial stage of manuscript preparation\. Its role was limited to refining phrasing, correcting grammar, and enhancing clarity and readability across different sections\. The LLM had no involvement in formulating research ideas, designing experiments, or conducting analyses\. All scientific contributions and findings are entirely the work of the authors\. The authors have ensured that the use of the LLM complies with ethical standards, avoiding plagiarism and scientific misconduct\.

Similar Articles

Learning Agentic Policy from Action Guidance

arXiv cs.CL

The paper proposes ActGuide-RL, a method for training agentic policies in LLMs by using human action data as guidance to overcome exploration barriers in reinforcement learning without extensive supervised fine-tuning.

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

arXiv cs.AI

This paper introduces ExpG, a mechanism for building and refining adaptive guidance that captures each tool's capability boundaries and best practices, enabling agents to use tools more robustly across diverse runtime conditions. Experiments show consistent improvements in tool selection, tool calling, and response generation, allowing smaller agents to outperform larger ones without ExpG.

Sample-Efficient Learning from Agent Experience

Hugging Face Daily Papers

Proposes Experience Distillation, a method that internalizes in-context learning gains from agent interaction histories into model weights without requiring additional environment interaction, achieving significant sample efficiency improvements on software engineering and text-adventure tasks.