@Xudong07452910: 一个 Agent 做了 20 步任务,最后只收到一个「成功 / 失败」。 训练时怎么知道,真正救了这次任务的到底是哪一步? 清华、浙大和美团团队这篇 AgentOPSD,研究的就是长时 Agent 的信用分配问题。 GRPO 通常把最终奖…
摘要
清华、浙大和美团团队提出AgentOPSD,一种递归自蒸馏的信用分配方法,将稀疏的最终奖励转化为逐步信用信号,提升长时Agent强化学习性能。在ALFWorld等任务上显著优于GRPO基线。
查看缓存全文
缓存时间: 2026/08/11 11:46
一个 Agent 做了 20 步任务,最后只收到一个「成功 / 失败」。
训练时怎么知道,真正救了这次任务的到底是哪一步?
清华、浙大和美团团队这篇 AgentOPSD,研究的就是长时 Agent 的信用分配问题。
GRPO 通常把最终奖励平均作用到整条轨迹。但真实任务里,有些步骤只是常规操作,少数关键决策才真正改变了结果;失败轨迹里也可能包含值得保留的正确判断。
AgentOPSD 会利用自蒸馏信号,判断每一轮行动让「最终成功」的可能性发生了多大变化,再根据这种变化重新分配训练信号。
更重要的是,它会结合前面的历史递归更新,而不是孤立评价每一步。
在 Qwen2.5-7B 上,ALFWorld 成功率从 GRPO 的 81.2% 提升到 89.1%。而且任务越长,优势越明显:GRPO 每增加一轮交互平均损失 2.91 个成功率点,AgentOPSD 只有 0.54。
长时 Agent 的训练,可能不能只知道「最后做成没有」,还需要逐渐学会判断:
这条轨迹里,究竟是哪几个决定真正改变了结果。
当奖励开始从轨迹级走向步骤级,Agent 才更有机会从自己的完整经历里学到东西。
arxiv: https://arxiv.org/abs/2608.05987
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Source: https://arxiv.org/html/2608.05987 Zi-Han Wang1,3,Zhengxi Lu2,Zhiyuan Yao2,Jinyang Wu1,Jie Wu1,Zhengzhou Cai3, Yueqing Sun3,Ziang Ye3,Linji Hao3,Qi Gu3,Xunliang Cai3,Yongliang Shen2,Yujiu Yang122footnotemark:2 1Tsinghua University2Zhejiang University3Meituan [email protected] [email protected]
Abstract
Reinforcement learning(RL) with verifiable rewards constructs trajectory-level advantage, yet often fails to credit the few pivotal decisions that drive outcomes in long-horizon multi-turn agentic RL. Some recent works introduce privileged self-distillation into credit assignment for RL, offering denser supervision, but it still remains unclear how such a local signal should expresssequentialcredit. We therefore proposeAgentOPSD, a critic-free recursive turn-level credit assignment for agentic reinforcement learning.AgentOPSDaggregates token-level teacher-student log-probability gaps into turn-level evidence, recursively updates a Bayesian belief state in log-odds space. This provides a principled reweighting scheme that transforms sparse outcome supervision into turn-level credit signals and identifies pivotal turns by the marginal revision between consecutive states while remaining fully compatible with standard policy optimization, requiring neither additional rollouts. We evaluateAgentOPSDon ALFWorld, WebShop, and Search-QA with two Qwen model scales (3B and 7B).AgentOPSDimproves over GRPO and strong self-distillation baselines, reaching89.1%89.1\%success on ALFWorld with Qwen2.5-7B, and ablations attribute the gains to turn-level aggregation and history-dependent recursive belief updates. Our code is available athttps://github.com/ZethWang/AgentOPSD.
Figure 1:Training dynamics and horizon-robustness ofAgentOPSDon Qwen2.5-7B-Instruct / ALFWorld.(a)Validation success rate over training.(b)Sensitivity to task horizon: success points lost per extra turn (OLS slope of per-sub-task success against the measured mean turns of successful episodes).(c)Policy entropy over training.## 1Introduction
Agentic post-training has become an important approach to improving the ability of large language models to solve complex tasks(Guoet al.,2025; Teamet al.,2025; Yanget al.,2025; Comaniciet al.,2025; Teamet al.,2026b). Unlike static, single-turn reasoning, agents must continuously interact with partially observable environments, at each turn acting on the current observation and receiving a new one as the environment transitions(Shenet al.,2023; Shiet al.,2025; Jimenezet al.,2023). However, many interactive environments provide verifiable rewards only after an entire trajectory terminates, forcing training algorithms to infer the contribution of each intermediate decision from a single sparse outcome. The decisions within a trajectory can nevertheless play substantially different roles: even successful trajectories may contain spurious, redundant, or misleading actions, whereas failed trajectories may still include useful reasoning.
Group-relative policy optimization methods such as GRPO(Shaoet al.,2024; Yuet al.,2025)and its agentic variants(Donget al.,2025; Fenget al.,2025)construct a trajectory-level advantage from outcome rewards and broadcast it uniformly across the trajectory. Such uniform credit cannot distinguish a few pivotal decisions from routine operations. This limitation becomes increasingly pronounced as the interaction horizon grows. Turn-level credit assignment is therefore essential for identifying the decisions that meaningfully influence the outcome and providing more precise supervision throughout long-horizon interactions.
A complementary line of work provides denser, token-level supervision. On-policy distillation(Yeet al.,2026a; Yanget al.,2026b; Teamet al.,2026a)trains a student on its own rollouts under a teacher, while itsself-distillation variants(Zhaoet al.,2026; Heet al.,2026)remove the need for a separate teacher by conditioning the same policy on privileged information available only during training(Luet al.,2026c). Recent studies further incorporate OPSD signals into reinforcement learning as an auxiliary source of supervision(Luet al.,2026b; Wanget al.,2026a).
However, applying OPSD to agentic reinforcement learning introduces two mismatches. First, OPSD’s token-level signals are not naturally aligned with agentic interaction(Luet al.,2026b), where multiple tokens jointly form an action and the environment responds only at turn boundaries. Second, even existing step-aware methods consider each turn in isolation(Zhanget al.,2026), without accounting for the evidence accumulated through preceding interactions. The central challenge is therefore to transform local OPSD signals into history-dependent turn-level credit.
Our key insight is thatthe credit of a turn should be determined not by its local signal in isolation, but by how much that signal changes the estimated probability of eventual success. To formalize this intuition, we interpret the per-turn self-distillation gap as new evidence that induces a Bayesian belief update(Åström,1965; Kaelblinget al.,1998). We define the corresponding belief state as the probability that the trajectory will ultimately succeed given the interaction history.
Based on this insight, we proposeAgentOPSD(Recursive Self-Distillation for Agentic Reinforcement Learning), a turn-level credit-assignment method for long-horizon agents.AgentOPSDaggregates token-level teacher–student log-probability gaps into turn-level evidence. Starting from the average group success rate, it then recursively updates Bayesian belief state at each turn in log-odds space without additional rollouts or a learned critic. The outcome verifier determines the global direction of optimization, while the bayesian belief updates redistribute the trajectory-level learning signal across turns. We evaluateAgentOPSDon three interactive environments—ALFWorld(Shridharet al.,2020), WebShop(Yaoet al.,2022), and Search-QA(Jinet al.,2025)—and across two model scales. As shown in Figure1,AgentOPSDconsistently outperforms GRPO and strong self-distillation baselines. Further ablations show that gains from aggregating token-level signals at turn boundaries aligned with environment transitions, and transforming an isolated local gap into a recursive revision of belief.
Our contributions are summarized as follows:
- •We formalize turn-level credit as the revision of a success belief induced by each turn. In log-odds space, this connects per-turn evidence to a recursive Bayesian update and reveals that an isolated self-distillation gap is not, by itself, sequential credit.
- •We introduceAgentOPSD, which aggregates token-level teacher–student log-probability gaps into environment-aligned turn-level evidence and recursively propagates this evidence through the trajectory-success belief, without additional rollouts or a learned critic.
- •Experiments and ablations across three interactive environments and two model scales demonstrate thatAgentOPSDconsistently outperforms GRPO and strong self-distillation baselines. Ablations further verify the complementary benefits of turn-boundary aggregation and recursive belief revision.
Figure 2:Overview ofAgentOPSD.Left:the agent loop, interacting with the environment over turns1,…,K1,\dots,K.Middle:AgentOPSDconverts GRPO’s single sequence-level advantage into turn-level reshaped advantages in three steps:(1)aggregate the token-level teacher–student gapsδk,t\delta_{k,t}within a turn into a turn-level gapeke_{k};(2)recursively update a belief stateBkB_{k}(initialized from the group success rate) and read off its marginal revisionΔBk=Bk−Bk−1\Delta B_{k}=B_{k}-B_{k-1};(3)reshape the sequence-level advantageAseq(i)A^{(i)}_{seq}per turn intoA~k(i)\tilde{A}^{(i)}_{k}.Right:vanilla GRPO instead broadcasts the sameAseq(i)A^{(i)}_{seq}to every token/turn. Each token in turnkkinheritsA~k\tilde{A}_{k}.
2Methodology
2.1Problem Setup
Given a taskxxand initial observationo0o_{0}, the agent starts froms1=(x,o0)s_{1}=(x,o_{0}). At turnkk, it samples whereπθ\pi_{\theta}is current policy,sks_{k}is its visible interaction history,yk,ty_{k,t}is thett-th token of actionaka_{k}, andLkL_{k}is the action length. After observingoko_{k}, the history becomessk+1=(sk,ak,ok)s_{k+1}=(s_{k},a_{k},o_{k}). AKK-turn episode forms𝝉=(s1,a1,o1,…,sK,aK,oK)\boldsymbol{\tau}=(s_{1},a_{1},o_{1},\ldots,s_{K},a_{K},o_{K})and receives a binary outcome rewardR(𝝉)R(\boldsymbol{\tau}).
ak=(yk,1,…,yk,Lk)∼πθ(⋅∣sk),a_{k}=(y_{k,1},\ldots,y_{k,L_{k}})\sim\pi_{\theta}(\cdot\mid s_{k}),(1)For each task, group-relative policy optimization samplesGGtrajectories and computes the sequence-level advantage. Hereiiindexes one of theGGsampled trajectories, whileR¯\bar{R}andσ^R\widehat{\sigma}_{R}are the group reward mean and standard deviation, andϵ0\epsilon_{0}is a small positive constant (reused throughout to avoid division by zero or infinite log-odds). GRPO assignsAseq(i)A_{\mathrm{seq}}^{(i)}to every token in trajectoryii, leaving turn-level credit unresolved.
Aseq(i)=R(i)−R¯σ^R+ϵ0,R¯=1G∑j=1GR(j).A_{\mathrm{seq}}^{(i)}=\frac{R^{(i)}-\bar{R}}{\widehat{\sigma}_{R}+\epsilon_{0}},\qquad\bar{R}=\frac{1}{G}\sum_{j=1}^{G}R^{(j)}.(2)
2.2From Outcome Contribution to Bayesian Turn Evidence
Directly measuring the counterfactual contribution of turnkkwould require marginalizing the outcome reward over all possible continuations followingaka_{k}, which is intractable in long-horizon interactions. We therefore adopt a hindsight-based evidential perspective. LetCCdenote the event that the trajectory eventually succeeds. Ifaka_{k}supports success, it should be more characteristic of successful behavior than of unsuccessful behavior. Bayes’ rule expresses the resulting change in the belief aboutCCas an action-side likelihood ratio(Åström,1965; Kaelblinget al.,1998):
logitp(C∣sk,ak)−logitp(C∣sk)=logp(ak∣sk,C)p(ak∣sk,¬C).\operatorname{logit}p(C\mid s_{k},a_{k})-\operatorname{logit}p(C\mid s_{k})=\log\frac{p(a_{k}\mid s_{k},C)}{p(a_{k}\mid s_{k},\neg C)}.(3)Herelogit(u)=logu1−u\operatorname{logit}(u)=\log\frac{u}{1-u}. The right-hand side is the ideal Bayes factor(Kass and Raftery,1995)between the success-conditional and failure-conditional likelihoods ofaka_{k}. Its sign indicates whether the action increases or decreases support for eventual success.
Because these outcome-conditional behavioral distributions are not directly available, we estimate this Bayesian evidence retrospectively using a computable per-turn self-distillation contrast between a privileged, success-associated self-teacher and the standard student policy. This contrast provides a tractable proxy for the otherwise inaccessible belief update; we characterize the approximation and its conditions below. The student and teacher share parametersθ\thetaand score the same student-generated action(Zhaoet al.,2026). Their token contexts are
hk,t=(sk,yk,<t),hk,t+=(sk,c+,yk,<t).h_{k,t}=(s_{k},y_{k,<t}),\qquad h_{k,t}^{+}=(s_{k},c^{+},y_{k,<t}).(4)whereyk,<ty_{k,<t}is the token prefix within turnkk, andc+c^{+}is a training-only retrieved skill describing useful subgoals and action patterns(Xiaet al.,2026). The skill-conditioned branch approximates success-associated behavior, while the unconditioned branch provides the background likelihood.
For tokenyk,ty_{k,t}, define the detached likelihood contrast
δk,t=logπθ(yk,t∣hk,t+)−logπθ(yk,t∣hk,t).\delta_{k,t}=\log\pi_{\theta}(y_{k,t}\mid h_{k,t}^{+})-\log\pi_{\theta}(y_{k,t}\mid h_{k,t}).(5)Positiveδk,t\delta_{k,t}means thatc+c^{+}increases the likelihood of the generated token. Summing over theLkL_{k}tokens gives the turn-level evidence
ek=∑t=1Lkδk,t=logπθ(ak∣sk,c+)πθ(ak∣sk).e_{k}=\sum_{t=1}^{L_{k}}\delta_{k,t}=\log\frac{\pi_{\theta}(a_{k}\mid s_{k},c^{+})}{\pi_{\theta}(a_{k}\mid s_{k})}.(6)Accordingly,eke_{k}provides a tractable hindsight approximation to the ideal Bayesian turn evidence in Eq. (3), under the conditions detailed in AppendixA.1. More generally, Bayes’ rule gives
logp(ak∣sk,C)p(ak∣sk)=logp(C∣sk,ak)p(C∣sk).\log\frac{p(a_{k}\mid s_{k},C)}{p(a_{k}\mid s_{k})}=\log\frac{p(C\mid s_{k},a_{k})}{p(C\mid s_{k})}.(7)Thus,eke_{k}can be interpreted as an evidential score whose sign indicates whetheraka_{k}raises or lowers support for eventual success. This sign-consistent Bayesian evidence is precisely the property thatAgentOPSDrelies on. We therefore treateke_{k}as a tractable Bayesian-inspired evidence proxy.
2.3Recursive Belief update
The local scoreeke_{k}does not indicate whether the same evidence is pivotal or redundant given earlier turns. We therefore maintain a decaying evidence accumulator and measure each turn by how much it revises the current support state:
B0\displaystyle B_{0}=clip(R¯,ϵ0,1−ϵ0),\displaystyle=\operatorname{clip}(\bar{R},\epsilon_{0},1-\epsilon_{0}),c0\displaystyle c_{0}=0,\displaystyle=0,(8)ck\displaystyle c_{k}=γck−1+ek,\displaystyle=\gamma\,c_{k-1}+e_{k},ℓk\displaystyle\ell_{k}=logit(B0)+ck=logit(B0)+∑j=1kγk−jej,\displaystyle=\operatorname{logit}(B_{0})+c_{k}=\operatorname{logit}(B_{0})+\sum_{j=1}^{k}\gamma^{\,k-j}e_{j},withBk=σ(ℓk)B_{k}=\sigma(\ell_{k})andσ(u)=(1+e−u)−1\sigma(u)=(1+e^{-u})^{-1}. HereR¯=S/G\bar{R}=S/Gis the fraction of successful trajectories in the group of sizeGG—the standard GRPO group mean (Prop.7)—andB0B_{0}clips it to[ϵ0,1−ϵ0][\epsilon_{0},1-\epsilon_{0}]withϵ0=10−4\epsilon_{0}{=}10^{-4}so its log-odds stay finite for all-correct or all-wrong groups.ckc_{k}is the accumulated evidence, andγ∈(0,1]\gamma\in(0,1]is a decay factor that down-weights older turns geometrically; only the evidenceckc_{k}decays, while the priorlogit(B0)\operatorname{logit}(B_{0})is retained at every step. Settingγ=1\gamma{=}1recovers the undiscounted accumulation of a log-likelihood ratio familiar from sequential testing(Wald,1945);γ<1\gamma{<}1makes the state recency-weighted, so that evidence from many turns ago no longer pins the support level. Sinceeke_{k}is estimated by the self-teacher (§2.2),BkB_{k}is treated as relative support rather than a calibrated success probability.
The importance of turnkkis its marginal support revision:
ΔBk\displaystyle\Delta B_{k}=Bk−Bk−1=σ(ℓk)−σ(ℓk−1),\displaystyle=B_{k}-B_{k-1}=\sigma(\ell_{k})-\sigma(\ell_{k-1}),(9)ΔBk\displaystyle\Delta B_{k}≈Bk−1(1−Bk−1)(ek−(1−γ)ck−1).\displaystyle\approx B_{k-1}(1-B_{k-1})\big(e_{k}-(1-\gamma)\,c_{k-1}\big).The incrementℓk−ℓk−1=ek−(1−γ)ck−1\ell_{k}-\ell_{k-1}=e_{k}-(1-\gamma)c_{k-1}is the new evidence net of the decayed carry-over, and it is weighted by the current state sensitivityBk−1(1−Bk−1)B_{k-1}(1-B_{k-1}): evidence has greatest effect under uncertainty and is suppressed once support saturates. Atγ=1\gamma{=}1this reduces toBk−1(1−Bk−1)ekB_{k-1}(1-B_{k-1})\,e_{k}. We update at turn boundaries; a token-level variant is used only as an ablation.
Outcome-aligned recursive credit.
We align the revision with the terminal update and read off its magnitude and direction:
qk=sign(Aseq)ΔBk.q_{k}=\operatorname{sign}(A_{\mathrm{seq}})\,\Delta B_{k}.(10)Themagnitude|ΔBk|=|qk||\Delta B_{k}|=|q_{k}|measures how much support the turn revises, while itssignsign(qk)\operatorname{sign}(q_{k})records whether that revision agrees with the verifier’s outcome signal. After the within-trajectory standardization below, turns with above-averageqkq_{k}are amplified and those below-average are attenuated; since the multiplier stays strictly positive, this never reverses the GRPO update direction.
2.4Bounded Advantage Reshaping
The raw creditqkq_{k}only modulates the magnitude of the verifier-derived advantage. For trajectoryii, we normalize itsKiK_{i}turn credits and apply a bounded multiplier:
μq(i)\displaystyle\mu_{q}^{(i)}=Ki−1∑j=1Kiqj(i),\displaystyle=K_{i}^{-1}\sum_{j=1}^{K_{i}}q_{j}^{(i)},σq(i)\displaystyle\sigma_{q}^{(i)}=Ki−1∑j=1Ki(qj(i)−μq(i))2,\displaystyle=\sqrt{K_{i}^{-1}\sum_{j=1}^{K_{i}}\big(q_{j}^{(i)}-\mu_{q}^{(i)}\big)^{2}},(11)zk(i)\displaystyle z_{k}^{(i)}=qk(i)−μq(i)σq(i)+ϵ0,\displaystyle=\frac{q_{k}^{(i)}-\mu_{q}^{(i)}}{\sigma_{q}^{(i)}+\epsilon_{0}},wk(i)\displaystyle w_{k}^{(i)}=clip(1+bzk(i),1−b,1+b),\displaystyle=\operatorname{clip}\!\left(1+bz_{k}^{(i)},\,1-b,\,1+b\right),A~k(i)\displaystyle\widetilde{A}_{k}^{(i)}=Aseq(i)[(1−λ)+λwk(i)],\displaystyle=A_{\mathrm{seq}}^{(i)}\big[(1-\lambda)+\lambda w_{k}^{(i)}\big],b\displaystyle b∈(0,1),λ∈[0,1].\displaystyle\in(0,1),\quad\lambda\in[0,1].Hereμq(i)\mu_{q}^{(i)}andσq(i)\sigma_{q}^{(i)}are the within-trajectory mean and standard deviation,zk(i)z_{k}^{(i)}is the normalized credit (ϵ0\epsilon_{0}stabilizes the normalization),bbsetswk(i)∈[1−b,1+b]w_{k}^{(i)}\in[1-b,1+b], andλ\lambdacontrols reshaping strength.
TokenttinheritsA~κi(t)(i)\widetilde{A}_{\kappa_{i}(t)}^{(i)}, yielding
ℒAgentOPSD(θ)\displaystyle\mathcal{L}_{\text{AgentOPSD}}(\theta)=−1G∑i=1G1∑tMi,t∑tMi,tmin(ri,tA~κi(t)(i),clip(ri,t,1−ε,1+ε)A~κi(t)(i))+βℒKL,\displaystyle=-\frac{1}{G}\sum_{i=1}^{G}\frac{1}{\sum_{t}M_{i,t}}\sum_{t}M_{i,t}\min\!\Big(r_{i,t}\widetilde{A}_{\kappa_{i}(t)}^{(i)},\operatorname{clip}(r_{i,t},1-\varepsilon,1+\varepsilon)\widetilde{A}_{\kappa_{i}(t)}^{(i)}\Big)+\beta\mathcal{L}_{\mathrm{KL}},(12)ri,t\displaystyle r_{i,t}=πθ(yi,t∣hi,t)πθold(yi,t∣hi,t).\displaystyle=\frac{\pi_{\theta}(y_{i,t}\mid h_{i,t})}{\pi_{\theta_{\mathrm{old}}}(y_{i,t}\mid h_{i,t})}.HereMi,t∈{0,1}M_{i,t}\in\{0,1\}masks valid response tokens,κi(t)\kappa_{i}(t)maps tokenttto its turn,ri,tr_{i,t}is the importance ratio against the rollout policyπθold\pi_{\theta_{\mathrm{old}}},ε\varepsilonis the clipping radius, andβ\betais its coefficient. No separate distillation loss is introduced; the detached self-teacher signal acts only throughA~\widetilde{A}.
Table 1:Performance on ALFWorld, Search-QA and WebShop.We report success rate (%) on ALFWorld, accuracy (%) on Search-QA, and Score/Acc (%) on WebShop. skills are training-only unless marked with∗*(validation with skills).AgentOPSDuses no skills at inference.Bestandsecond-bestare highlighted.
3Experiments
3.1Experimental Setup
Benchmarks.
We evaluate on three environments.ALFWorld(Shridharet al.,2020)is a text embodied benchmark over six household task categories—Pick and Place (Pick), Look at Object in Light (Look), Pick Clean then Place (Clean), Pick Heat then Place (Heat), Pick Cool then Place (Cool), and Pick Two and Place (Pick2).Search-QAfollows the Search-R1 setup(Jinet al.,2025)and covers single-hop QA (NQ(Kwiatkowskiet al.,2019), TriviaQA(Joshiet al.,2017), PopQA(Mallenet al.,2023)) and multi-hop QA (HotpotQA(Yanget al.,2018), 2Wiki(Hoet al.,2020), MuSiQue(Trivediet al.,2022), Bamboogle(Presset al.,2023)), with NQ and HotpotQA in-domain and the rest held out; retrieval uses E5(Wanget al.,2022).WebShop(Yaoet al.,2022)is an interactive online-shopping environment; we evaluate on the 128 fixed validation tasks ofFenget al.(2025).
Implementation.
We train Qwen2.5-3B/7B-Instruct on8×8\timesH800 GPUs. The privileged skills are retrieved from theSkillBankof SkillRL(Xiaet al.,2026)by keyword matching and are used only during training; inference uses no external skills. PriorB0B_{0}set to the fraction of successful trajectories in each GRPO group (the standard group meanR¯\bar{R}). All other optimization settings are shared with the SDAR baseline. Full training andAgentOPSDhyperparameters are listed in AppendixF(Table3).
Baselines.
We compare against three groups.(1) Training-free:Vanilla(the base model) andSkill-Prompt, which prepends retrieved skills at inference.(2) Group-relative RL:GRPO(Shaoet al.,2024)andSkill-GRPO, which injects skills into the training prompt (evaluated with,Skill-GRPO*, or without retrieved skills).(3) Self-distillation RL:OPSD(Zhaoet al.,2026),GRPO+OPSD,Skill-SD(Wanget al.,2026a),RLSD(Yanget al.,2026a),SDAR(Luet al.,2026b)andStepOPSD(Zhanget al.,2026), all of which use the teacher–student gap but inject it as a gate, magnitude, or auxiliary loss. All methods share the same backbone, data, and budget. Full algorithm details are in AppendixB.
3.2Main Results
The gain comes from credit construction, not privileged access.
Under our unified setup,AgentOPSDand the privileged baselines use the same retrieved skills; they differ primarily in how the skill-induced teacher–student discrepancy enters learning.AgentOPSDoutperforms GRPO+OPSD, Skill-SD, and RLSD on all eight aggregate comparisons across the two model scales, and exceeds SDAR on six of eight. This controlled-information comparison isolates the benefit ofAgentOPSD: a local teacher-student gap is not yet a reliable credit signal. Accumulating that gap into a belief state and assigning credit according to belief revision more effectively identifies the turns that change the predicted outcome.
The advantage grows with the interaction horizon.
AgentOPSDis designed for the regime where uniform credit is most harmful, so we also ask how performance degrades as tasks require more turns. Figure1(b) regresses per-sub-task success on the measured mean number of turns of successful episodes on ALFWorld (Qwen2.5-7B), reporting the success points lost per additional turn. The uniform-credit methods degrade fastest (−3.59-3.59for RLSD and−2.91-2.91for GRPO points per turn), whereasAgentOPSDis the flattest at−0.54-0.54. This is consistent with the motivation for turn-level credit: the longer the trajectory, the more decisions a single broadcast advantage has to cover, and the more a history-dependent revision helps.
3.3Mechanism Ablation
Table2evaluates each design choice on ALFWorld with Qwen2.5-7B by removing or replacing one component at a time. The full method achieves a success rate of89.189.1.
ComponentAblationALFWorldAgentOPSD(full)turn-level, bounded,λ=0.5\lambda{=}0.589.1Turn-level granularityper-token accumulation85.9Recursive state revision (8)raw local gapeke_{k}in place ofΔBk\Delta B_{k}82.8Signed direction (10)magnitude|ΔBk||\Delta B_{k}|only (drop outcome sign)80.5State priorB0B_{0}anchordrop empirical-rate initialization78.9Table 2:Component ablation ofAgentOPSDon ALFWorld with Qwen2.5-7B (success rate, %). Each row removes or replaces a single mechanism while holding all other settings fixed. The signed direction and the state prior anchor have the largest impact on performance, while the recursive state revision and turn-level granularity provide smaller but consistent improvements.#### Granularity and recursion.
Replacing turn-level belief tracking with per-token accumulation reduces the success rate to85.985.9: environment feedback is associated with a complete action rather than an individual token, so token-level accumulation fragments a single decision and weakens the alignment between the gap and outcomes. Replacing the recursive revisionΔBk\Delta B_{k}with the raw local gapeke_{k}further reduces performance to82.882.8. A raweke_{k}scores each turn in isolation, whereasΔBk\Delta B_{k}measures how that gap revises the belief state accumulated over the preceding history—so the same local gap that is decisive while the outcome is open becomes redundant once the accumulated state already points to an outcome. This controlled comparison isolates the value of the recursion and directly confirms our central principle that a local gap is not sequential credit.
Outcome-aligned signed direction.
Keeping only the magnitude|ΔBk||\Delta B_{k}|and dropping the sign (Eq.10), i.e. standardizing|ΔBk||\Delta B_{k}|instead of the signedqkq_{k}, lowers performance to80.580.5. The magnitude identifies where the belief state changes, but cannot determine whether that change agrees with the verifier outcome. For a successful trajectory, an upward belief revision is consistent with the outcome, whereas for a failed trajectory the same revision is inconsistent. The signed direction makes this distinction explicit, allowing outcome-consistent revisions to receive more credit and contradictory revisions to receive less.
State-prior anchoring.
Removing the empirical priorB0=clip(R¯,ϵ0,1−ϵ0)B_{0}=\operatorname{clip}(\bar{R},\epsilon_{0},1-\epsilon_{0})reduces the success rate to78.978.9. The group success rateR¯\bar{R}provides a verifier-grounded estimate of task difficulty before the trajectory-specific gap is accumulated. Moreover,B0B_{0}determines the initial log-odds and thus the operating region of theB(1−B)B(1-B)gate. Without this anchor, trajectories begin from an arbitrary uncertainty level, which can mis-scale early belief revisions and distort which early turns appear pivotal. The ablations therefore separate three roles: belief revision localizes credit, the signed direction aligns it with the final outcome, and prior anchoring stabilizes its reference point.
3.4Hyperparameter Sensitivity
Whereas the mechanism ablation asks whether each component is necessary, we now examine how sensitiveAgentOPSDis to its continuous hyperparameters. We sweep one knob at a time while holding the others at the full-AgentOPSDsetting (λ=0.5\lambda{=}0.5,γ=0.95\gamma{=}0.95,ϵhigh=0.24\epsilon_{\mathrm{high}}{=}0.24; AppendixF) across three configurations (Figure3): long-horizon ALFWorld with Qwen2.5-7B (89.189.1) and Qwen2.5-3B (84.484.4), and short-horizon Search-QA with Qwen2.5-3B (46.746.7).
Figure 3:Hyperparameter sensitivity ofAgentOPSD.Rows: ALFWorld (Qwen2.5-7B), Search-QA (Qwen2.5-3B), and ALFWorld (Qwen2.5-3B). Columns sweep one knob (λ\lambda,γ\gamma,ϵhigh\epsilon_{\mathrm{high}}) with the others held at our setting. Curves are rolling means; shaded bands show the local±1\pm 1standard deviation.#### Reshaping weightλ\lambda.
Sweepingλ∈{0.5,0.25,0.1,0.01}\lambda\in\{0.5,0.25,0.1,0.01\}interpolates between pure GRPO (λ=0\lambda{=}0) and full belief reshaping. This is the knob with the clearest effect:λ=0.5\lambda{=}0.5is best and any smaller value reduces performance (89.189.1atλ=0.5\lambda{=}0.5vs.84.4/85.9/83.684.4/85.9/83.6; Search46.746.7vs.45.1/40.2/45.445.1/40.2/45.4), consistent with a smallerλ\lambdadown-weighting the bounded multiplier and discarding turn-level credit. We useλ=0.5\lambda{=}0.5throughout.
Evidence decayγ\gamma.
Turn-level evidence is accumulated with a geometric decayck=γck−1+ekc_{k}=\gamma\,c_{k-1}+e_{k}(equivalentlyℓk=logit(B0)+∑j≤kγk−jej\ell_{k}=\operatorname{logit}(B_{0})+\sum_{j\leq k}\gamma^{\,k-j}e_{j}); sweepingγ∈{1.0,0.95,0.9,0.8}\gamma\in\{1.0,0.95,0.9,0.8\}moves the result within a few points (87.5/82.0/85.287.5/82.0/85.2; Search45.1/44.5/45.545.1/44.5/45.5) without a monotone trend, so the recursion is not particularly sensitive to how fast old evidence is discounted. We use the mild settingγ=0.95\gamma{=}0.95for all main results.
Policy clippingϵhigh\epsilon_{\mathrm{high}}.
Fixingϵlow=0.2\epsilon_{\mathrm{low}}{=}0.2and varyingϵhigh∈{0.2,0.24,0.28}\epsilon_{\mathrm{high}}\in\{0.2,0.24,0.28\}(clip-higher(Yuet al.,2025)),AgentOPSDis largely unaffected (88.388.3at both0.20.2and0.280.28; Search46.946.9and45.745.7), indicating that the reshaped objective inherits the trust-region robustness of GRPO. Overall, onlyλ\lambdaproduces a systematic effect, and the spread across all knobs shrinks sharply on the four-turn Search-QA task—the settings that matter on long-horizon ALFWorld are largely inert when little history accumulates, which is again consistent with the method acting where long-horizon credit assignment is needed.
4Related Work
4.1Agentic Post-Training with Verifiable Rewards
Reinforcement learning with verifiable rewards has advanced from single-turn reasoning(Shaoet al.,2024; Guoet al.,2025; Yuet al.,2025; Luet al.,2026a)to long-horizon agents in embodied text worlds, web shopping, and retrieval-augmented question answering(Shridharet al.,2020; Yaoet al.,2022; Jinet al.,2025; Luet al.,2025). In these interactive settings, sparse terminal rewards make turn-level credit assignment particularly challenging. Standard GRPO broadcasts a trajectory-level advantage uniformly across turns. GiGPO(Fenget al.,2025)improves reward-side credit assignment by combining episode-level advantages with step-level advantages estimated from repeated anchor states across trajectories. In contrast,AgentOPSDderives turn-level evidence from privileged teacher–student likelihood gaps and assigns credit through recursive belief revision. GiGPO andAgentOPSDtherefore operate on complementary signal sources—environment rewards and self-distillation evidence, respectively.
4.2On-Policy (Self-)Distillation
On-policy distillation trains a policy on its own rollouts under a teacher(Agarwalet al.,2024; Guet al.,2026; Wenet al.,2023). Its recent self-distillation variants remove the need for a separate teacher while the teacher branch is conditioned on privileged information available only during training(Zhaoet al.,2026; Heet al.,2026; Luet al.,2026c). Recent studies incorporate the resulting teacher–student log-probability gap into RLVR by using it to scale or reshape the advantage(Yanget al.,2026a), as a detached auxiliary objective(Luet al.,2026b; Wanget al.,2026a), or as reweighted, scheduled, or reward-densifying local supervision(Xuet al.,2026; Wanget al.,2026b; Yeet al.,2026a; Heet al.,2026). Existing methods predominantly treat the distillation gap as a local token-level or step-level signal. Token-level signals are not naturally aligned with action turns, and the contribution of a turn depends on the evidence accumulated through preceding interactions. StepOPSD(Zhanget al.,2026)aggregates the teacher–student signal over action-centered step spans but still scores each span by its local log-ratio. In contrast,AgentOPSDfirst aggregates token-level gaps within each turn and then recursively accumulates the resulting evidence into a running support state.
4.3Long-Horizon Credit Assignment
Assigning credit across a long horizon is a classical problem. PPO learns a value function and, via GAE, derives a per-step temporal-difference signals(Schulmanet al.,2017;2016). When rewards are sparse and delayed, return-decomposition methods such as RUDDER redistribute a terminal reward to the steps responsible for it(Arjona-Medinaet al.,2019), while process reward models and Monte-Carlo credit methods such as VinePPO estimate intermediate value by additional rollouts or a learned scorer(Cuiet al.,2025; Kazemnejadet al.,2024). These approaches recover per-step structure but reintroduce the cost GRPO removed: a trained critic, a reward model, or many extra rollouts.AgentOPSDrestores a per-turn value signal in the critic-free group-relative setting, at the cost of a single teacher forward pass. The belief state plays the role of GAE’s value baseline and its per-turn revision the role of the TD signal, but without a learned value network cost.
5Conclusion
We studied credit assignment for long-horizon language agents, where trajectory-level rewards provide limited supervision for distinguishing pivotal decisions from routine or redundant actions. Our key insight is that turn-level credit should depend not only on a local signal, but also on how that signal revises the accumulated belief in eventual trajectory success. Based on this insight, we proposedAgentOPSD, which aggregates token-level self-distillation gaps at environment-aligned turn boundaries and recursively updates a trajectory-success belief in log-odds space. These belief revisions redistribute the trajectory-level advantage across turns without requiring additional rollouts or a learned critic. Experiments across three interactive environments and two model scales show thatAgentOPSDconsistently outperforms GRPO and strong self-distillation baselines. Ablations further confirm the importance of both turn-level signal aggregation and history-dependent belief revision. Overall, our results suggest that recursive belief updating provides a simple and effective approach to restoring fine-grained temporal credit in critic-free agentic reinforcement learning.
References
- R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. Ramos, M. Geist, and O. Bachem (2024)On-policy distillation of language models: learning from self-generated mistakes.External Links:2306.13649,LinkCited by:§4.2.
- J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter (2019)RUDDER: return decomposition for delayed rewards.External Links:1806.07857,LinkCited by:§4.3.
- K. J. Åström (1965)Optimal control of markov processes with incomplete state information i.Journal of Mathematical Analysis and Applications10,pp. 174–205.External Links:DocumentCited by:§1,§2.2.
- Y. Chen, Z. Cai, X. Ji, W. Zhao, A. Zhang, X. Wang, and T. Chua (2026a)Understanding multilingualism in mixture-of-experts llms: routing mechanism, expert specialization, and layerwise steering.arXiv preprint arXiv:2601.14050.Cited by:Appendix C.
- Y. Chen, Y. Wang, Y. Zhang, Z. Ye, Z. Cai, Y. Shi, Q. Gu, H. Su, X. Cai, X. Wang,et al.(2026b)Learning to self-verify makes language models better reasoners.arXiv preprint arXiv:2602.07594.Cited by:Appendix C.
- G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen,et al.(2025)Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261.Cited by:§1.
- G. Cui, L. Yuan, Z. Wang, H. Wang, Y. Zhang, J. Chen, W. Li, B. He, Y. Fan, T. Yu, Q. Xu, W. Chen, J. Yuan, H. Chen, K. Zhang, X. Lv, S. Wang, Y. Yao, X. Han, H. Peng, Y. Cheng, Z. Liu, M. Sun, B. Zhou, and N. Ding (2025)Process reinforcement through implicit rewards.External Links:2502.01456,LinkCited by:§4.3.
- G. Dong, H. Mao, K. Ma, L. Bao, Y. Chen, Z. Wang, Z. Chen, J. Du, H. Wang, F. Zhang,et al.(2025)Agentic reinforced policy optimization.arXiv preprint arXiv:2507.19849.Cited by:§1.
- L. Feng, Z. Xue, T. Liu, and B. An (2025)Group-in-group policy optimization for llm agent training.arXiv preprint arXiv:2505.10978.Cited by:Appendix C,§1,§3.1,§4.1.
- Y. Gu, L. Dong, F. Wei, and M. Huang (2026)MiniLLM: on-policy distillation of large language models.External Links:2306.08543,LinkCited by:§4.2.
- D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi,et al.(2025)Deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948.Cited by:§1,§4.1.
- Y. He, S. Kaur, A. Bhaskar, Y. Yang, J. Liu, N. Ri, L. Fowl, A. Panigrahi, D. Chen, and S. Arora (2026)Self-distillation zero: self-revision turns binary rewards into dense supervision.External Links:2604.12002,LinkCited by:§1,§4.2.
- X. Ho, A. D. Nguyen, S. Sugawara, and A. Aizawa (2020)Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.InProceedings of the 28th International Conference on Computational Linguistics,pp. 6609–6625.Cited by:Appendix C,§3.1.
- X. Ji, Y. Chen, Z. Cai, X. Wang, A. Zhang, and T. Chua (2026)Tiny brains, giant impact: uncovering the keystone neurons of llm with just a few prompts.arXiv preprint arXiv:2605.24846.Cited by:Appendix C.
- C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan (2023)Swe-bench: can language models resolve real-world github issues?.arXiv preprint arXiv:2310.06770.Cited by:§1.
- B. Jin, H. Zeng, Z. Yue, J. Yoon, S. Arik, D. Wang, H. Zamani, and J. Han (2025)Search-r1: training llms to reason and leverage search engines with reinforcement learning.arXiv preprint arXiv:2503.09516.Cited by:Appendix C,§1,§3.1,§4.1.
- M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer (2017)Triviaqa: a large scale distantly supervised challenge dataset for reading comprehension.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),pp. 1601–1611.Cited by:Appendix C,§3.1.
- L. P. Kaelbling, M. L. Littman, and A. R. Cassandra (1998)Planning and acting in partially observable stochastic domains.Artificial Intelligence101(1–2),pp. 99–134.External Links:DocumentCited by:§1,§2.2.
- R. E. Kass and A. E. Raftery (1995)Bayes factors.Journal of the American Statistical Association90(430),pp. 773–795.External Links:DocumentCited by:§2.2.
- A. Kazemnejad, M. Aghajohari, E. Portelance, A. Sordoni, S. Reddy, A. Courville, and N. L. Roux (2024)VinePPO: refining credit assignment in rl training of llms.External Links:2410.01679,LinkCited by:§4.3.
- T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee,et al.(2019)Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics7,pp. 453–466.Cited by:Appendix C,§3.1.
- Z. Lu, Y. Chai, Y. Guo, X. Yin, L. Liu, H. Wang, H. Xiao, S. Ren, P. Zhao, G. Liu,et al.(2026a)Ui-r1: enhancing efficient action prediction of gui agents by reinforcement learning.InProceedings of the AAAI Conference on Artificial Intelligence,Vol.40,pp. 17608–17616.Cited by:§4.1.
- Z. Lu, Z. Yao, Z. Han, Z. Wang, J. Wu, Q. Gu, X. Cai, W. Lu, J. Xiao, Y. Zhuang,et al.(2026b)Self-distilled agentic reinforcement learning.arXiv preprint arXiv:2605.15155.Cited by:Appendix D,§1,§1,§3.1,§4.2.
- Z. Lu, Z. Yao, J. Wu, C. Han, Q. Gu, X. Cai, W. Lu, J. Xiao, Y. Zhuang, and Y. Shen (2026c)SKILL0: in-context agentic reinforcement learning for skill internalization.External Links:2604.02268,LinkCited by:§1,§4.2.
- Z. Lu, J. Ye, F. Tang, Y. Shen, H. Xu, Z. Zheng, W. Lu, M. Yan, F. Huang, J. Xiao,et al.(2025)Ui-s1: advancing gui automation via semi-online reinforcement learning.arXiv preprint arXiv:2509.11543.Cited by:§4.1.
- A. Mallen, A. Asai, V. Zhong, R. Das, D. Khashabi, and H. Hajishirzi (2023)When not to trust language models: investigating effectiveness of parametric and non-parametric memories.InProceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers),pp. 9802–9822.Cited by:Appendix C,§3.1.
- O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis (2023)Measuring and narrowing the compositionality gap in language models.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp. 5687–5711.Cited by:Appendix C,§3.1.
- J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel (2016)High-dimensional continuous control using generalized advantage estimation.External Links:1506.02438,LinkCited by:§4.3.
- J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017)Proximal policy optimization algorithms.External Links:1707.06347,LinkCited by:§4.3.
- Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu,et al.(2024)Deepseekmath: pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300.Cited by:Appendix D,§1,§3.1,§4.1.
- Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang (2023)Hugginggpt: solving ai tasks with chatgpt and its friends in hugging face.Advances in Neural Information Processing Systems36,pp. 38154–38180.Cited by:§1.
- Z. Shi, S. Gao, L. Yan, Y. Feng, X. Chen, Z. Chen, D. Yin, S. Verberne, and Z. Ren (2025)Tool learning in the wild: empowering language models as automatic tool agents.InProceedings of the ACM on Web Conference 2025,pp. 2222–2237.Cited by:§1.
- M. Shridhar, X. Yuan, M. Côté, Y. Bisk, A. Trischler, and M. Hausknecht (2020)Alfworld: aligning text and embodied environments for interactive learning.arXiv preprint arXiv:2010.03768.Cited by:Appendix C,§1,§3.1,§4.1.
- C. Team, B. Xiao, B. Xia, B. Yang, B. Gao, B. Shen, C. Zhang, C. He, C. Lou, F. Luo, G. Wang, G. Xie, H. Zhang, H. Lv, H. Li, H. Chen, H. Xu, H. Zhang, H. Liu, J. Duo, J. Wei, J. Xiao, J. Dong, J. Shi, J. Hu, K. Bao, K. Zhou, L. Li, L. Zhao, L. Zhang, P. Li, Q. Chen, S. Liu, S. Yu, S. Cao, S. Chen, S. Yu, S. Liu, T. Zhou, W. Su, W. Wang, W. Ma, X. Deng, B. Mao, B. Ye, C. Cai, C. Wang, C. Zhu, C. Ma, C. Chen, C. Li, D. Zhu, D. Xiao, D. Zhang, D. Zhang, F. Liu, F. Yang, F. Shi, G. Wang, H. Tian, H. Wu, H. Qu, H. Yi, H. An, H. Guan, X. Zhang, Y. Song, Y. Yan, Y. Zhao, Y. Lai, Y. Gao, Y. Cheng, Y. Tian, Y. Wang, Z. Tang, Z. Tang, Z. Wen, Z. Song, Z. Zheng, Z. Jiang, J. Wen, J. Sun, J. Li, J. Xue, J. Xia, K. Fang, M. Zhu, N. Chen, Q. Tu, Q. Zhang, Q. Wang, R. Li, R. Ma, S. Zhang, S. Wang, S. Li, S. Gu, S. Ren, S. Deng, T. Guo, T. Lu, W. Zhuang, W. Zhang, W. Xiong, W. Huang, W. Yang, X. Zhang, X. Yong, X. Wang, X. Xie, Y. Jiang, Y. Yang, Y. He, Y. Tu, Y. Dong, Y. Liu, Y. Ma, Y. Yu, Y. Xiang, Z. Huang, Z. Lin, Z. Xu, Z. Chen, Z. Deng, Z. Zhang, and Z. Yue (2026a)MiMo-v2-flash technical report.External Links:2601.02780,LinkCited by:§1.
- K. Team, Y. Bai, Y. Bao, Y. Charles, C. Chen, G. Chen, H. Chen, H. Chen, J. Chen, N. Chen,et al.(2025)Kimi k2: open agentic intelligence.arXiv preprint arXiv:2507.20534.Cited by:§1.
- M. L. Team, A. Gui, B. Li, B. Tao, B. Zhou, B. Chen, C. Zhang, C. Gao, C. Zhang, C. Han,et al.(2026b)Longcat-flash-thinking-2601 technical report.arXiv preprint arXiv:2601.16725.Cited by:§1.
- H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal (2022)MuSiQue: multi-hop questions via single-hop question composition.Transactions of the Association for Computational Linguistics10,pp. 539–554.Cited by:Appendix C,§3.1.
- A. Wald (1945)Sequential tests of statistical hypotheses.The Annals of Mathematical Statistics16(2),pp. 117–186.External Links:DocumentCited by:§2.3.
- H. Wang, G. Wang, H. Xiao, Y. Zhou, Y. Pan, J. Wang, K. Xu, Y. Wen, X. Ruan, X. Chen, and H. Qi (2026a)Skill-sd: skill-conditioned self-distillation for multi-turn llm agents.External Links:2604.10674,LinkCited by:Appendix D,§1,§3.1,§4.2.
- J. Wang, W. Zhang, W. Shi, Y. Li, and J. Cheng (2026b)TCOD: exploring temporal curriculum in on-policy distillation for multi-turn autonomous agents.External Links:2604.24005,LinkCited by:§4.2.
- L. Wang, N. Yang, X. Huang, B. Jiao, L. Yang, D. Jiang, R. Majumder, and F. Wei (2022)Text embeddings by weakly-supervised contrastive pre-training.arXiv preprint arXiv:2212.03533.Cited by:Appendix C,§3.1.
- Y. Wen, Z. Li, W. Du, and L. Mou (2023)F-divergence minimization for sequence-level knowledge distillation.External Links:2307.15190,LinkCited by:§4.2.
- P. Xia, J. Chen, H. Wang, J. Liu, K. Zeng, Y. Wang, S. Han, Y. Zhou, X. Zhao, H. Chen, Z. Zheng, C. Xie, and H. Yao (2026)SkillRL: evolving agents via recursive skill-augmented reinforcement learning.External Links:2602.08234,LinkCited by:§2.2,§3.1.
- H. Xu, Z. Wang, Z. Zhu, L. Pan, X. Chen, S. Fan, L. Chen, and K. Yu (2025)Alignment for efficient tool calling of large language models.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp. 17787–17803.Cited by:Appendix C.
- H. Xu, Z. Zhu, L. Pan, Z. Wang, S. Zhu, D. Ma, R. Cao, L. Chen, and K. Yu (2024)Reducing tool hallucination via reliability alignment.arXiv preprint arXiv:2412.04141.Cited by:Appendix C.
- Y. Xu, H. Sang, Z. Zhou, R. He, Z. Wang, and A. Geramifard (2026)TIP: token importance in on-policy distillation.External Links:2604.14084,LinkCited by:§4.2.
- A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv,et al.(2025)Qwen3 technical report.arXiv preprint arXiv:2505.09388.Cited by:§1.
- C. Yang, C. Qin, Q. Si, M. Chen, N. Gu, D. Yao, Z. Lin, W. Wang, J. Wang, and N. Duan (2026a)Self-distilled rlvr.External Links:2604.03128,LinkCited by:Appendix D,§3.1,§4.2.
- W. Yang, W. Liu, R. Xie, K. Yang, S. Yang, and Y. Lin (2026b)Learning beyond teacher: generalized on-policy distillation with reward extrapolation.External Links:2602.12125,LinkCited by:§1.
- Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning (2018)HotpotQA: a dataset for diverse, explainable multi-hop question answering.InProceedings of the 2018 conference on empirical methods in natural language processing,pp. 2369–2380.Cited by:Appendix C,§3.1.
- S. Yao, H. Chen, J. Yang, and K. Narasimhan (2022)Webshop: towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems35,pp. 20744–20757.Cited by:Appendix C,§1,§3.1,§4.1.
- T. Ye, L. Dong, X. Wu, S. Huang, and F. Wei (2026a)On-policy context distillation for language models.External Links:2602.12275,LinkCited by:§1,§4.2.
- Z. Ye, W. Shi, Y. Liu, Y. Wang, Z. Cai, Y. Shi, Q. Gu, X. Cai, and F. Feng (2026b)Look before you leap: autonomous exploration for llm agents.arXiv preprint arXiv:2605.16143.Cited by:Appendix C.
- Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, W. Dai, T. Fan, G. Liu, L. Liu,et al.(2025)Dapo: an open-source llm reinforcement learning system at scale.arXiv preprint arXiv:2503.14476.Cited by:§1,§3.4,§4.1.
- Y. Zhang, X. Lin, and C. Wu (2026)StepOPSD: step-aware online preference distillation for agent reinforcement learning.External Links:2605.27140,LinkCited by:Appendix D,§1,§3.1,§4.2.
- S. Zhao, Z. Xie, M. Liu, J. Huang, G. Pang, F. Chen, and A. Grover (2026)Self-distilled reasoner: on-policy self-distillation for large language models.External Links:2601.18734,LinkCited by:Appendix D,§1,§2.2,§3.1,§4.2.
- H. Zhou, Y. Chen, S. Guo, X. Yan, K. H. Lee, Z. Wang, K. Y. Lee, G. Zhang, K. Shao, L. Yang,et al.(2025)Memento: fine-tuning llm agents without fine-tuning llms.arXiv preprint arXiv:2508.16153.Cited by:Appendix C.
Table of Contents
- 1Introduction
- 2Methodology1. 2.1Problem Setup 2. 2.2From Outcome Contribution to Bayesian Turn Evidence 3. 2.3Recursive Belief update 4. 2.4Bounded Advantage Reshaping
- 3Experiments1. 3.1Experimental Setup 2. 3.2Main Results 3. 3.3Mechanism Ablation 4. 3.4Hyperparameter Sensitivity
- 4Related Work1. 4.1Agentic Post-Training with Verifiable Rewards 2. 4.2On-Policy (Self-)Distillation 3. 4.3Long-Horizon Credit Assignment
- 5Conclusion
- References
- ATheoretical Analysis1. A.1From the Bayes factor to the self-teacher contrast 2. A.2Properties of the reshaping
- BAlgorithm
- CDatasets
- DBaseline Details
- EEvaluation Metrics
- FHyperparameters
- GTraining Dynamics
- HPrompt
Appendix ATheoretical Analysis
A.1From the Bayes factor to the self-teacher contrast
AgentOPSDapproximates the ideal per-turn Bayes factor
ℬk=logp(ak∣sk,C)p(ak∣sk,¬C)=logitp(C∣sk,ak)−logitp(C∣sk)\mathcal{B}_{k}\;=\;\log\frac{p(a_{k}\mid s_{k},C)}{p(a_{k}\mid s_{k},\neg C)}\;=\;\operatorname{logit}p(C\mid s_{k},a_{k})-\operatorname{logit}p(C\mid s_{k})(13)by the self-teacher contrast
ek=logπθ(ak∣sk,c+)πθ(ak∣sk),e_{k}\;=\;\log\frac{\pi_{\theta}(a_{k}\mid s_{k},c^{+})}{\pi_{\theta}(a_{k}\mid s_{k})},(14)whereCCdenotes eventual success andρk=p(C∣sk)\rho_{k}=p(C\mid s_{k}). We use two assumptions:(A1)the skill-conditioned branch is success-conditional,πθ(ak∣sk,c+)≈p(ak∣sk,C)\pi_{\theta}(a_{k}\mid s_{k},c^{+})\approx p(a_{k}\mid s_{k},C);(A2)when success is rare (ρk\rho_{k}small) the marginal is failure-dominated,πθ(ak∣sk)≈p(ak∣sk,¬C)\pi_{\theta}(a_{k}\mid s_{k})\approx p(a_{k}\mid s_{k},\neg C).
The marginal action distribution is the success/failure mixture
πθ(ak∣sk)=ρkp(ak∣sk,C)+(1−ρk)p(ak∣sk,¬C).\pi_{\theta}(a_{k}\mid s_{k})=\rho_{k}\,p(a_{k}\mid s_{k},C)+(1-\rho_{k})\,p(a_{k}\mid s_{k},\neg C).(15)Substituting (15) into (14) under (A1),
ek≈ℬk−log(1−ρk+ρkeℬk)→ρk→0ℬk,e_{k}\;\approx\;\mathcal{B}_{k}-\log\!\big(1-\rho_{k}+\rho_{k}\,e^{\mathcal{B}_{k}}\big)\;\xrightarrow[\;\rho_{k}\to 0\;]{}\;\mathcal{B}_{k},(16)so (A2) is theρk→0\rho_{k}\to 0limit in which the contrast recovers the Bayes factor. Under (A1) alone,eke_{k}is the pointwise mutual information
ek≈logp(ak∣sk,C)p(ak∣sk)=logp(C∣sk,ak)p(C∣sk),e_{k}\;\approx\;\log\frac{p(a_{k}\mid s_{k},C)}{p(a_{k}\mid s_{k})}\;=\;\log\frac{p(C\mid s_{k},a_{k})}{p(C\mid s_{k})},(17)positive iffaka_{k}raises the posterior success probability. The correction in (16) is monotone inℬk\mathcal{B}_{k}, hencesign(ek)=sign(ℬk)\operatorname{sign}(e_{k})=\operatorname{sign}(\mathcal{B}_{k})andeke_{k}preserves the ranking of turns by evidential strength;AgentOPSDuseseke_{k}only through this sign and ranking.
A.2Properties of the reshaping
LetA(i)A^{(i)}be the group-relative advantage,ΔBk=Bk−Bk−1\Delta B_{k}=B_{k}-B_{k-1}the per-turn belief revision,zkz_{k}its within-trajectory standardization,mk=clip(1+bsign(A(i))zk,1−b,1+b)m_{k}=\mathrm{clip}\!\big(1+b\,\operatorname{sign}(A^{(i)})z_{k},\,1-b,\,1+b\big)withb∈(0,1)b\in(0,1), andA~k=A(i)((1−λ)+λmk)\tilde{A}_{k}=A^{(i)}\big((1-\lambda)+\lambda m_{k}\big)withλ∈[0,1]\lambda\in[0,1].
Proposition 1(Boundedness).
|A~k−A(i)|≤λb|A(i)|\big|\tilde{A}_{k}-A^{(i)}\big|\leq\lambda b\,|A^{(i)}|, hence(1−λb)|A(i)|≤|A~k|≤(1+λb)|A(i)|(1-\lambda b)|A^{(i)}|\leq|\tilde{A}_{k}|\leq(1+\lambda b)|A^{(i)}|.
Proof.
mk∈[1−b,1+b]m_{k}\in[1-b,1+b]gives|mk−1|≤b|m_{k}-1|\leq b, andA~k−A(i)=A(i)λ(mk−1)\tilde{A}_{k}-A^{(i)}=A^{(i)}\lambda(m_{k}-1). ∎
Proposition 2(Sign Preservation).
sign(A~k)=sign(A(i))\operatorname{sign}(\tilde{A}_{k})=\operatorname{sign}(A^{(i)})for every turnkk.
Proof.
(1−λ)+λmk≥1−λb>0(1-\lambda)+\lambda m_{k}\geq 1-\lambda b>0sinceλ≤1,b<1\lambda\leq 1,\,b<1; a strictly positive factor preserves sign. ∎
Proposition 3(Recovery of GRPO).
Atλ=0\lambda=0,A~k=A(i)\tilde{A}_{k}=A^{(i)}for every token and theAgentOPSDgradient equals the GRPO gradient.
Proof.
λ=0\lambda=0gives(1−λ)+λmk=1(1-\lambda)+\lambda m_{k}=1, soA~k=A(i)\tilde{A}_{k}=A^{(i)}identically, independent of the belief signal. ∎
Proposition 4(First-Order Decomposition of the Belief Revision).
Forck=γck−1+ekc_{k}=\gamma c_{k-1}+e_{k},ℓk=logit(B0)+ck\ell_{k}=\operatorname{logit}(B_{0})+c_{k}, andΔℓk=ek−(1−γ)ck−1\Delta\ell_{k}=e_{k}-(1-\gamma)c_{k-1},
ΔBk=Bk−1(1−Bk−1)Δℓk+O((Δℓk)2).\Delta B_{k}=B_{k-1}(1-B_{k-1})\,\Delta\ell_{k}+O\!\big((\Delta\ell_{k})^{2}\big).(18)
Proof.
Bk=σ(ℓk)B_{k}=\sigma(\ell_{k})withσ′=σ(1−σ)\sigma^{\prime}=\sigma(1-\sigma); a first-order expansion aroundℓk−1\ell_{k-1}givesBk=Bk−1+Bk−1(1−Bk−1)Δℓk+O((Δℓk)2)B_{k}=B_{k-1}+B_{k-1}(1-B_{k-1})\Delta\ell_{k}+O((\Delta\ell_{k})^{2}). The gateB(1−B)B(1-B)is maximal atB=12B=\tfrac{1}{2}and vanishes asB→{0,1}B\to\{0,1\}. ∎
Proposition 5(Exact Budget of the Idealized Recursion).
The idealized increments telescope to the endpoint change,∑kΔBk=BK−B0\sum_{k}\Delta B_{k}=B_{K}-B_{0}.
Proof.
∑k=1K(Bk−Bk−1)=BK−B0\sum_{k=1}^{K}(B_{k}-B_{k-1})=B_{K}-B_{0}. ∎
Proposition 6(Non-Identifiability of Per-Turn Contribution).
There exist two trajectories with identical outcome reward—hence identical broadcast advantage—whose per-turn contributions differ; per-turn credit is not identifiable from the trajectory return alone.
Proof (by construction).
Takeτ1,τ2\tau_{1},\tau_{2}in one group withR(τ1)=R(τ2)R(\tau_{1})=R(\tau_{2}), soA(τ1)=A(τ2)A(\tau_{1})=A(\tau_{2})and GRPO assigns the same scalar to every turn. Letτ1\tau_{1}succeed through a single decisive turn (ΔB\Delta Bconcentrated) andτ2\tau_{2}through evenly spread progress. The returns coincide but the per-turn contributions differ, so an additional per-turn signal is required to recover them. ∎
Proposition 7(B0B_{0}as the Group Success-Rate Estimate).
For a task with success probabilityθx\theta_{x}and a group ofGGtrajectories yieldingSSsuccesses under a binary reward, the maximum-likelihood estimate ofθx\theta_{x}is the group success fractionR¯=S/G\bar{R}=S/G, which is the standard GRPO group mean;AgentOPSDsetsB0=clip(R¯,ϵ0,1−ϵ0)B_{0}=\operatorname{clip}(\bar{R},\epsilon_{0},1-\epsilon_{0}).
Proof.
Under a Binomial(G,θx)(G,\theta_{x})likelihood the MLE isS/GS/G. The clip (ϵ0=10−4\epsilon_{0}{=}10^{-4}) only keepslogit(B0)\operatorname{logit}(B_{0})finite for all-correct or all-wrong groups, which have zero group-relative advantage and hence do not contribute to the update. ∎
Appendix BAlgorithm
We give pseudocode for oneAgentOPSDtraining iteration at turn-level granularity in Algorithm1. The only addition over GRPO is a single teacher forward pass per turn and the per-turn belief reshaping block; everything else is the standard group-relative update.
Algorithm 1AgentOPSD: Recursive State Updates for Turn-Level Credit1:policy
πθ\pi_{\theta}, verifier
RR, group size
GG, skill retriever; mixing
λ\lambda, bound
bb, evidence decay
γ\gamma 2:foreach training iterationdo
3:sample a batch of tasks
{x}\{x\} 4:foreach task
xxwith retrieved skill
c+c^{+}do
5:sample
GGtrajectories
{y(1),…,y(G)}∼πθ(⋅∣x)\{y^{(1)},\dots,y^{(G)}\}\sim\pi_{\theta}(\cdot\mid x); trajectory
iihas
KiK_{i}turns⊳\trianglerighton-policy rollout
6:for
i=1,…,Gi=1,\dots,Gdo
7:obtain reward
R(i)=R(x,y(i))∈{0,1}R^{(i)}=R(x,y^{(i)})\in\{0,1\}from the verifier
8:endfor
9:
Aseq(i)←(R(i)−μG)/σGA_{\mathrm{seq}}^{(i)}\leftarrow(R^{(i)}-\mu_{G})/\sigma_{G}⊳\trianglerightgroup-relative advantage
10:for
i=1,…,Gi=1,\dots,Gdo
11:
B0←clip(R¯,ϵB,1−ϵB)B_{0}\leftarrow\mathrm{clip}(\bar{R},\epsilon_{B},1-\epsilon_{B});
ℓ0←logit(B0)\ell_{0}\leftarrow\mathrm{logit}(B_{0})⊳\trianglerightstandard GRPO group success rateR¯=S/G\bar{R}{=}S/G(Prop.7)
12:
c0←0c_{0}\leftarrow 0 13:for
k=1,…,Kik=1,\dots,K_{i}do⊳\trianglerightper-turn belief state
14:
ek←∑tsg[logπθ(yk,t∣sk+)−logπθ(yk,t∣sk)]e_{k}\leftarrow\sum_{t}\mathrm{sg}[\log\pi_{\theta}(y_{k,t}\mid s_{k}^{+})-\log\pi_{\theta}(y_{k,t}\mid s_{k})]⊳\trianglerightone extra teacher forward
15:
ck←γck−1+ekc_{k}\leftarrow\gamma\,c_{k-1}+e_{k};
ℓk←ℓ0+ck\ell_{k}\leftarrow\ell_{0}+c_{k};
Bk←σ(ℓk)B_{k}\leftarrow\sigma(\ell_{k});
ΔBk←Bk−Bk−1\Delta B_{k}\leftarrow B_{k}-B_{k-1} 16:endfor
17:
qk←sign(Aseq(i))ΔBkq_{k}\leftarrow\mathrm{sign}(A_{\mathrm{seq}}^{(i)})\,\Delta B_{k}for all
kk⊳\trianglerightoutcome-aligned credit
18:
zk←(qk−mean(q))/(std(q)+ϵ)z_{k}\leftarrow(q_{k}-\mathrm{mean}(q))/(\mathrm{std}(q)+\epsilon)⊳\trianglerightwithin-trajectory standardization
19:
wk←clip(1+bzk,1−b,1+b)w_{k}\leftarrow\mathrm{clip}(1+b\,z_{k},\,1-b,\,1+b)⊳\trianglerightbounded multiplier
20:
A~k(i)←Aseq(i)[(1−λ)+λwk]\widetilde{A}_{k}^{(i)}\leftarrow A_{\mathrm{seq}}^{(i)}\big[(1-\lambda)+\lambda w_{k}\big]; each token inherits
A~\widetilde{A}of its turn
21:endfor
22:endfor
23:update
θ\thetaby maximizing the clipped GRPO objective
ℒAgentOPSD(θ)\mathcal{L}_{\text{AgentOPSD}}(\theta)with
{A~}\{\widetilde{A}\}⊳\trianglerightpolicy update
24:endfor
Token-level variant.
Replace the per-turn recursion with the same recursion over the flattened token sequence under a response mask: accumulateδt\delta_{t}directly, standardizeΔBt\Delta B_{t}over the trajectory’s tokens, and assignA~t\widetilde{A}_{t}per token. This is the granularity ablation reported in Table2.
Cost.
The overhead over GRPO is one teacher forward pass per trajectory; the belief reshaping block is elementwise and adds no rollouts and no learned parameters.
Appendix CDatasets
Our experiments span three multi-turn agentic environments covering embodied household reasoning, web navigation, and search-augmented question answering.
ALFWorld
(Shridharet al.,2020)is a text-based embodied environment with six task categories—Pick and Place, Look at Object in Light, Pick Clean then Place, Pick Heat then Place, Pick Cool then Place, and Pick Two and Place. Given a language goal and textual observations, the agent selects admissible actions until the goal is satisfied.
WebShop
(Yaoet al.,2022)is an interactive online-shopping environment. For each user request the agent searches the product catalog, inspects candidate items, selects the required attributes, and attempts a purchase satisfying the specified constraints. We evaluate on the128128fixed validation tasks ofFenget al.(2025).
Search-QA
follows the Search-R1 setup(Jinet al.,2025)over seven datasets: single-hop NQ(Kwiatkowskiet al.,2019), TriviaQA(Joshiet al.,2017), PopQA(Mallenet al.,2023)and multi-hop HotpotQA(Yanget al.,2018), 2Wiki(Hoet al.,2020), MuSiQue(Trivediet al.,2022), Bamboogle(Presset al.,2023), with NQ and HotpotQA in-domain and the rest held out. The agent issues search queries, inspects retrieved documents (retrieval via E5(Wanget al.,2022)), and synthesizes the collected evidence before returning its final answer(Yeet al.,2026b; Chenet al.,2026b;a; Jiet al.,2026; Zhouet al.,2025; Xuet al.,2024;2025).
Appendix DBaseline Details
We compare against three groups of baselines. Unless a method is marked with∗*, evaluation uses only the standard task prompt and the interaction history returned by the environment;∗*indicates that a retrieved skill is additionally supplied during validation and testing.
Vanilla.
The instruction-tuned backbone evaluated without any post-training.
Skill-Prompt∗.
The same frozen parameters as Vanilla, but a retrieved task-relevant skill is prepended to the context at validation/test time, measuring the inference-time value of skills without any parameter update.
GRPO
(Shaoet al.,2024). A critic-free group-relative RL algorithm: it samples a group of trajectories per task, normalizes their terminal rewards into relative advantages, and optimizes a clipped surrogate objective; every token inherits its trajectory’s sequence-level advantage.
Skill-GRPO / Skill-GRPO∗.
GRPO with a retrieved skill injected into the training prompt. The skill is removed at inference for Skill-GRPO (testing whether the guidance has been internalized), and kept at inference for Skill-GRPO∗.
OPSD
(Zhaoet al.,2026). On-policy self-distillation: a teacher branch conditioned on training-only privileged context re-scores the student’s sampled tokens and produces dense token-level targets through distribution matching; the teacher outputs are detached and the privileged context is not used at inference.
GRPO+OPSD.
Jointly optimizes the trajectory-level GRPO loss and the token-level OPSD objective, a straightforward combination of outcome-based RL and generic self-distillation.
Skill-SD
(Wanget al.,2026a). Supplies the retrieved skill only to the teacher branch and trains the student to absorb the skill-conditioned guidance via an importance-weighted distillation loss, without requiring skills at evaluation.
RLSD
(Yanget al.,2026a). Converts the teacher–student log-probability gap into a bounded coefficient that scales the magnitude of each token’s GRPO update; the sign of the update remains determined by the outcome-derived advantage.
SDAR
(Luet al.,2026b). Adds a separately gated auxiliary self-distillation loss on top of GRPO, leaving the original GRPO advantage unchanged and using a bounded gate to modulate each teacher signal.
StepOPSD
(Zhanget al.,2026). Applies the teacher–student self-distillation signal at the turn (step) level rather than per token, but uses each step’s local signal in isolation.
All post-training baselines share the same backbone models, environment interfaces, data, and training budget asAgentOPSD; they differ primarily in their optimization objective and in whether skills are available during training or evaluation.
Appendix EEvaluation Metrics
ALFWorld.
We report the overall success rate over the evaluation tasks (the fraction of episodes that reach the specified goal), which is the sample-weighted average of the six per-category success rates.
Search-QA.
We report the overall accuracy over all evaluation questions aggregated across the seven datasets.
WebShop.
We report a normalized completion Score (averaged over partial-constraint satisfaction and scaled by100100) and an exact-completion success rate Succ. (the percentage of episodes that satisfy all specified requirements).
Appendix FHyperparameters
Table3summarizes the hyperparameters used byAgentOPSDacross all our experiments. We deliberately use asinglesetting for every environment and model scale rather than tuning per task:AgentOPSDruns at turn-level granularity with reshaping weightλ=0.5\lambda{=}0.5, multiplier bandb=0.2b{=}0.2, gap accumulation with decayγ=0.95\gamma{=}0.95, policy clippingϵlow=0.2\epsilon_{\mathrm{low}}{=}0.2/ϵhigh=0.24\epsilon_{\mathrm{high}}{=}0.24, and the empirical group success rateR¯=S/G\bar{R}=S/G(clipped) as the state priorB0B_{0}. The sensitivity study in §3.4sweepsλ\lambda,γ\gammaandϵhigh\epsilon_{\mathrm{high}}around this setting and finds no swept value that improves on it by a meaningful margin, which is why one shared configuration is used throughout rather than per-environment tuning.
Table 3:Hyperparameters.η\eta: learning rate;GG: group size;ϵlow/ϵhigh\epsilon_{\mathrm{low}}/\epsilon_{\mathrm{high}}: PPO clip range;αKL\alpha_{\mathrm{KL}}: KL penalty coefficient toward the reference policy; SRS: skill retrieval strategy (KM = keyword matching).AgentOPSD-specific reshaping knobs:λ\lambda(mult_lambda, reshaping weight),bb(mult_band, multiplier band), andγ\gamma(gap_decay_gamma, gap-accumulation decay). A single shared setting is used across all environments and model scales.We use a single shared optimization recipe across environments (learning rate, group size, PPO clip range, and KL coefficient in Table3; dual-clip constantc=3.0c{=}3.0, gradient clipping1.01.0, entropy coefficient0.0010.001, one PPO epoch per update, and FSDP on a single node). The remaining settings are environment-specific and summarized in Table4.
Table 4:Per-environment training configuration.Optimization settings shared across all environments are listed in Table3; the environment-specific settings are below.ALFWorldWebShopSearch-QATraining steps150150150Train batch size1616128Rollout group sizeGG888Max prompt length204840964096Max response length512512512Max interaction turns50154Rollout temperature (train / val)1.0 / 0.41.0 / 0.41.0 / 0.4GPUs (tensor-parallel size)8 (2)2 (2)4 (1)
Appendix GTraining Dynamics
We present the full training dynamics ofAgentOPSDacross all model scales and environments in Figures4–5, tracking the teacher–student gap and the reward throughout training.
Figure 4:Teacher–Student Gapδ¯\bar{\delta}when training Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop and Search-QA.
Figure 5:Reward Curvewhen training Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop and Search-QA.
Appendix HPrompt
Figures6–8present the full prompt templates used byAgentOPSDfor the three evaluation environments, where{skill_context}is populated with the retrieved skill during training and left empty at inference time.
Prompt ofAgentOPSDon ALFWorldYou are an expert agent operating in the ALFRED Embodied Environment. Your task is to: {task_description}.{skill_context}Prior to this step, you have already taken {step_count} step(s). Below are the most recent {history_length} observations and the corresponding actions you took: {action_history}You are now at step {current_step} and your current observation is: {current_observation}Your admissible actions of the current situation are: [{admissible_actions}].Now it’s your turn to take an action. You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within<think> </think>tags. Once you’ve finished your reasoning, you should choose an admissible action for current step and present it within<action> </action>tags.Figure 6:Prompt template used byAgentOPSDfor the ALFWorld task environment.Prompt ofAgentOPSDon Search-based QAYou are an expert agent tasked with answering the given question step-by-step.{skill_context}Your question: {task_description}.Prior to this step, you have already taken {step_count} step(s). Below is the interaction history where<search> </search>wrapped your past search queries and<information> </information>wrapped the corresponding search results returned by the external search engine. History:{memory_context}Now it’s your turn to respond for the current step. You should first conduct a reasoning process. This process MUST be enclosed within<think> </think>tags. After completing your reasoning, choose only one of the following actions (do not perform both):1.If you find you lack some knowledge, youMUSTcall a search engine to get more external information using format:<search> your query </search>.2.If you have enough knowledge to answer the question confidently, provide your final answer within<answer> </answer>tags, without detailed illustrations. For example,<answer>Beijing</answer>.Figure 7:Prompt template used byAgentOPSDfor the Search-based QA task environment.Prompt ofAgentOPSDon WebShopYou are an expert autonomous agent operating in the WebShop e-commerce environment.{skill_context}Your task is to: {task_description}.Prior to this step, you have already taken {step_count} step(s). Below are the most recent {history_length} observations and the corresponding actions you took: {action_history}You are now at step {current_step} and your current observation is: {current_observation}.Your admissible actions of the current situation are: [ {available_actions} ].Now it’s your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within<think> </think>tags. Once you’ve finished your reasoning, you should choose an admissible action for current step and present it within<action> </action>tags.Figure 8:Prompt template used byAgentOPSDfor the WebShop task environment.
相似文章
@Xudong07452910: RL 训练 LLM Agent 有个经典难题: 一次长任务失败后,模型到底该从哪里学起? 最终奖励通常只能告诉 Agent「成功或失败」,却很难指出中间哪些判断值得保留,哪些动作把整条轨迹带偏了。 这篇论文提出 SEED,用「自进化在线蒸…
这篇论文提出SEED方法,通过自进化在线蒸馏将轨迹中的事后技能内化到模型参数中,解决长任务RL训练中奖励稀疏的问题,在ALFWorld等基准上取得了显著提升。
@Xudong07452910: Agent 记忆最危险的时候,往往是它太相信过去。 很多 Memory Agent 找到相似经验后,会直接把它塞进上下文。但任务相似,不代表当前状态相同,旧经验有时反而会把决策带偏。 这篇论文提出 MemHarness,把 Agent 使…
MemHarness 提出将 Agent 记忆从简单回放改为基于当前状态的重构,通过 GRPO 端到端训练,在 ALFWorld 和 WebShop 上显著提升成功率。
@xiaohu: 昨天看很多人转发Apodex 1.1 一个专门面向深度研究而打造的 Agent 专门解决那种"没有现成答案、需要大量调研才能搞定"的硬问题 好奇测试了下,跑了俩任务,一下午都没跑完 执行时间是真长 这玩意能你只要给它个目标,它就能能长时间…
Apodex 1.1 是一个专为深度研究设计的AI代理,能处理需要大量调研的复杂任务,通过主代理分解问题并异步派发多个子代理执行,支持长时间运行和自动恢复。
@wsl8297: 用 AI Agent 跑复杂任务,最难受的往往不是模型不够强,而是对话一变长,上下文就开始爆仓。 你还得一遍遍补背景、重讲流程,再加上工具调用吐出来的冗余日志,Token 像开了口子一样往外流。 最近看到腾讯开源的 TencentDB A…
腾讯开源了 TencentDB Agent Memory,通过分层记忆管理(符号化短期记忆+分层长期记忆)解决AI Agent长对话上下文爆仓问题,实测Token消耗最高降低61%,任务通过率提升超50%。
@Xudong07452910: 最近看 Agent 自进化相关的工作,我一直比较关注一个问题: Agent 留下来的 trajectory,到底什么时候才值得被下一轮训练继续学习? 这次我拿 SEED 的一篇 self-evolving Agent 论文,丢进 Apod…
这篇文章探讨了AI代理自进化中轨迹学习的重要性,并介绍了Apodex 1.1系统和开源工具FrontierAgent,用于执行和评估长时间的研究任务。