@Xudong07452910: An Agent completed a 20-step task and only received a "success/failure" at the end. During training, how do you know which step actually saved the task? This AgentOPSD paper by Tsinghua, Zhejiang University, and Meituan team studies the credit assignment problem for long-horizon agents. GRPO usually assigns the final reward…

X AI KOLs Timeline Papers

Summary

Tsinghua, Zhejiang University, and Meituan team propose AgentOPSD, a recursive self-distillation credit assignment method that converts sparse final rewards into step-wise credit signals, improving reinforcement learning performance for long-horizon agents. It significantly outperforms the GRPO baseline on tasks like ALFWorld.

An Agent completed a 20-step task and received only a single "success/failure" at the end. During training, how do you know which step actually saved the task? This AgentOPSD paper by Tsinghua, Zhejiang University, and Meituan team studies the credit assignment problem for long-horizon agents. GRPO typically averages the final reward across the entire trajectory. But in real tasks, some steps are just routine operations, and only a few key decisions truly change the outcome; failed trajectories may also contain correct judgments worth preserving. AgentOPSD uses self-distillation signals to determine how much each action changes the likelihood of "final success", and then reassigns training signals based on that change. More importantly, it recursively updates by incorporating previous history, rather than evaluating each step in isolation. On Qwen2.5-7B, the ALFWorld success rate improves from 81.2% with GRPO to 89.1%. And the longer the task, the more obvious the advantage: GRPO loses an average of 2.91 success-rate points per additional interaction round, while AgentOPSD loses only 0.54. Training long-horizon agents may require more than just knowing "whether it succeeded in the end" — they also need to gradually learn to judge: which decisions in this trajectory truly changed the outcome. When rewards start moving from trajectory-level to step-level, agents have a better chance of learning from their complete experience. arxiv: https://arxiv.org/abs/2608.05987
Original Article
View Cached Full Text

Cached at: 08/11/26, 11:46 AM

An agent performs a 20-step task and only receives a single “success/failure” at the end. During training, how do we know which step actually saved the task? This AgentOPSD work by Tsinghua, Zhejiang University, and Meituan investigates the credit assignment problem for long-horizon agents. GRPO usually broadcasts the final reward uniformly over the entire trajectory. But in real tasks, some steps are only routine operations, and a few key decisions truly change the outcome; failed trajectories may also contain correct judgments worth preserving. AgentOPSD uses self-distillation signals to estimate how much each turn’s action changed the probability of eventual success, then reassigns training signals accordingly. More importantly, it updates recursively conditioned on the preceding history rather than evaluating each step in isolation. On Qwen2.5-7B, ALFWorld success rate improves from GRPO’s 81.2% to 89.1%. And the advantage grows as tasks get longer: GRPO loses an average of 2.91 success-rate points per additional interaction turn, while AgentOPSD loses only 0.54. Training long-horizon agents may require more than just knowing whether the task was finally completed; the model also needs to gradually learn to judge: within this trajectory, which decisions truly changed the outcome. When rewards begin to move from trajectory-level to step-level, agents have a better chance of learning from their full experience. arXiv: https://arxiv.org/abs/2608.05987 — # AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Source: https://arxiv.org/html/2608.05987 Zi-Han Wang1,3,Zhengxi Lu2,Zhiyuan Yao2,Jinyang Wu1,Jie Wu1,Zhengzhou Cai3, Yueqing Sun3,Ziang Ye3,Linji Hao3,Qi Gu3,Xunliang Cai3,Yongliang Shen2,Yujiu Yang122footnotemark:2 1Tsinghua University2Zhejiang University3Meituan [email protected] [email protected] ###### Abstract Reinforcement learning(RL) with verifiable rewards constructs trajectory-level advantage, yet often fails to credit the few pivotal decisions that drive outcomes in long-horizon multi-turn agentic RL. Some recent works introduce privileged self-distillation into credit assignment for RL, offering denser supervision, but it still remains unclear how such a local signal should expresssequentialcredit. We therefore proposeAgentOPSD, a critic-free recursive turn-level credit assignment for agentic reinforcement learning.AgentOPSDaggregates token-level teacher-student log-probability gaps into turn-level evidence, recursively updates a Bayesian belief state in log-odds space. This provides a principled reweighting scheme that transforms sparse outcome supervision into turn-level credit signals and identifies pivotal turns by the marginal revision between consecutive states while remaining fully compatible with standard policy optimization, requiring neither additional rollouts. We evaluateAgentOPSDon ALFWorld, WebShop, and Search-QA with two Qwen model scales (3B and 7B).AgentOPSDimproves over GRPO and strong self-distillation baselines, reaching89.1%89.1\%success on ALFWorld with Qwen2.5-7B, and ablations attribute the gains to turn-level aggregation and history-dependent recursive belief updates. Our code is available athttps://github.com/ZethWang/AgentOPSD. Refer to captionFigure 1:Training dynamics and horizon-robustness ofAgentOPSDon Qwen2.5-7B-Instruct / ALFWorld.(a)Validation success rate over training.(b)Sensitivity to task horizon: success points lost per extra turn (OLS slope of per-sub-task success against the measured mean turns of successful episodes).(c)Policy entropy over training.## 1Introduction Agentic post-training has become an important approach to improving the ability of large language models to solve complex tasks(Guoet al.,2025 (https://arxiv.org/html/2608.05987#bib.bib22); Teamet al.,2025 (https://arxiv.org/html/2608.05987#bib.bib27); Yanget al.,2025 (https://arxiv.org/html/2608.05987#bib.bib26); Comaniciet al.,2025 (https://arxiv.org/html/2608.05987#bib.bib30); Teamet al.,2026b (https://arxiv.org/html/2608.05987#bib.bib31)). Unlike static, single-turn reasoning, agents must continuously interact with partially observable environments, at each turn acting on the current observation and receiving a new one as the environment transitions(Shenet al.,2023 (https://arxiv.org/html/2608.05987#bib.bib33); Shiet al.,2025 (https://arxiv.org/html/2608.05987#bib.bib29); Jimenezet al.,2023 (https://arxiv.org/html/2608.05987#bib.bib34)). However, many interactive environments provide verifiable rewards only after an entire trajectory terminates, forcing training algorithms to infer the contribution of each intermediate decision from a single sparse outcome. The decisions within a trajectory can nevertheless play substantially different roles: even successful trajectories may contain spurious, redundant, or misleading actions, whereas failed trajectories may still include useful reasoning. Group-relative policy optimization methods such as GRPO(Shaoet al.,2024 (https://arxiv.org/html/2608.05987#bib.bib23); Yuet al.,2025 (https://arxiv.org/html/2608.05987#bib.bib51))and its agentic variants(Donget al.,2025 (https://arxiv.org/html/2608.05987#bib.bib35); Fenget al.,2025 (https://arxiv.org/html/2608.05987#bib.bib20))construct a trajectory-level advantage from outcome rewards and broadcast it uniformly across the trajectory. Such uniform credit cannot distinguish a few pivotal decisions from routine operations. This limitation becomes increasingly pronounced as the interaction horizon grows. Turn-level credit assignment is therefore essential for identifying the decisions that meaningfully influence the outcome and providing more precise supervision throughout long-horizon interactions. A complementary line of work provides denser, token-level supervision. On-policy distillation(Yeet al.,2026a (https://arxiv.org/html/2608.05987#bib.bib36); Yanget al.,2026b (https://arxiv.org/html/2608.05987#bib.bib37); Teamet al.,2026a (https://arxiv.org/html/2608.05987#bib.bib39))trains a student on its own rollouts under a teacher, while itsself-distillation variants(Zhaoet al.,2026 (https://arxiv.org/html/2608.05987#bib.bib24); Heet al.,2026 (https://arxiv.org/html/2608.05987#bib.bib38))remove the need for a separate teacher by conditioning the same policy on privileged information available only during training(Luet al.,2026c (https://arxiv.org/html/2608.05987#bib.bib6)). Recent studies further incorporate OPSD signals into reinforcement learning as an auxiliary source of supervision(Luet al.,2026b (https://arxiv.org/html/2608.05987#bib.bib67); Wanget al.,2026a (https://arxiv.org/html/2608.05987#bib.bib5)). However, applying OPSD to agentic reinforcement learning introduces two mismatches. First, OPSD’s token-level signals are not naturally aligned with agentic interaction(Luet al.,2026b (https://arxiv.org/html/2608.05987#bib.bib67)), where multiple tokens jointly form an action and the environment responds only at turn boundaries. Second, even existing step-aware methods consider each turn in isolation(Zhanget al.,2026 (https://arxiv.org/html/2608.05987#bib.bib55)), without accounting for the evidence accumulated through preceding interactions. The central challenge is therefore to transform local OPSD signals into history-dependent turn-level credit. Our key insight is thatthe credit of a turn should be determined not by its local signal in isolation, but by how much that signal changes the estimated probability of eventual success. To formalize this intuition, we interpret the per-turn self-distillation gap as new evidence that induces a Bayesian belief update(Åström,1965 (https://arxiv.org/html/2608.05987#bib.bib69); Kaelblinget al.,1998 (https://arxiv.org/html/2608.05987#bib.bib68)). We define the corresponding belief state as the probability that the trajectory will ultimately succeed given the interaction history. Based on this insight, we proposeAgentOPSD(Recursive Self-Distillation for Agentic Reinforcement Learning), a turn-level credit-assignment method for long-horizon agents.AgentOPSDaggregates token-level teacher–student log-probability gaps into turn-level evidence. Starting from the average group success rate, it then recursively updates Bayesian belief state at each turn in log-odds space without additional rollouts or a learned critic. The outcome verifier determines the global direction of optimization, while the bayesian belief updates redistribute the trajectory-level learning signal across turns. We evaluateAgentOPSDon three interactive environments—ALFWorld(Shridharet al.,2020 (https://arxiv.org/html/2608.05987#bib.bib11)), WebShop(Yaoet al.,2022 (https://arxiv.org/html/2608.05987#bib.bib9)), and Search-QA(Jinet al.,2025 (https://arxiv.org/html/2608.05987#bib.bib12))—and across two model scales. As shown in Figure1 (https://arxiv.org/html/2608.05987#S0.F1),AgentOPSDconsistently outperforms GRPO and strong self-distillation baselines. Further ablations show that gains from aggregating token-level signals at turn boundaries aligned with environment transitions, and transforming an isolated local gap into a recursive revision of belief. Our contributions are summarized as follows: - •We formalize turn-level credit as the revision of a success belief induced by each turn. In log-odds space, this connects per-turn evidence to a recursive Bayesian update and reveals that an isolated self-distillation gap is not, by itself, sequential credit. - •We introduceAgentOPSD, which aggregates token-level teacher–student log-probability gaps into environment-aligned turn-level evidence and recursively propagates this evidence through the trajectory-success belief, without additional rollouts or a learned critic. - •Experiments and ablations across three interactive environments and two model scales demonstrate thatAgentOPSDconsistently outperforms GRPO and strong self-distillation baselines. Ablations further verify the complementary benefits of turn-boundary aggregation and recursive belief revision. Refer to captionFigure 2:Overview ofAgentOPSD.Left:the agent loop, interacting with the environment over turns1,…,K1,\dots,K.Middle:AgentOPSDconverts GRPO’s single sequence-level advantage into turn-level reshaped advantages in three steps:(1)aggregate the token-level teacher–student gapsδk,t\delta_{k,t}within a turn into a turn-level gapeke_{k};(2)recursively update a belief stateBkB_{k}(initialized from the group success rate) and read off its marginal revisionΔBk=Bk−Bk−1\Delta B_{k}=B_{k}-B_{k-1};(3)reshape the sequence-level advantageAseq(i)A^{(i)}_{seq}per turn intoA~k(i)\tilde{A}^{(i)}_{k}.Right:vanilla GRPO instead broadcasts the sameAseq(i)A^{(i)}_{seq}to every token/turn. Each token in turnkkinheritsA~k\tilde{A}_{k}. ## 2Methodology ### 2.1Problem Setup Given a taskxxand initial observationo0o_{0}, the agent starts froms1=(x,o0)s_{1}=(x,o_{0}). At turnkk, it samples whereπθ\pi_{\theta}is current policy,sks_{k}is its visible interaction history,yk,ty_{k,t}is thett-th token of actionaka_{k}, andLkL_{k}is the action length. After observingoko_{k}, the history becomessk+1=(sk,ak,ok)s_{k+1}=(s_{k},a_{k},o_{k}). AKK-turn episode formsτ=(s1,a1,o1,…,sK,aK,oK)\boldsymbol{\tau}=(s_{1},a_{1},o_{1},\ldots,s_{K},a_{K},o_{K})and receives a binary outcome rewardR(τ)R(\boldsymbol{\tau}). ak=(yk,1,…,yk,Lk)∼πθ(⋅∣sk),a_{k}=(y_{k,1},\ldots,y_{k,L_{k}})\sim\pi_{\theta}(\cdot\mid s_{k})),(1)For each task, group-relative policy optimization samplesGGtrajectories and computes the sequence-level advantage. Hereiiindexes one of theGGsampled trajectories, whileR ̄\bar{R}andσ^R\widehat{\sigma}_{R}are the group reward mean and standard deviation, andε0\epsilon_{0}is a small positive constant (reused throughout to avoid division by zero or infinite log-odds). GRPO assignsAseq(i)A_{\mathrm{seq}}^{(i)}to every token in trajectoryii, leaving turn-level credit unresolved. Aseq(i)=R(i)−R ̄σ^R+ε0,R ̄=1G∑j=1GR(j).A_{\mathrm{seq}}^{(i)}=\frac{R^{(i)}-\bar{R}}{\widehat{\sigma}_{R}+\epsilon_{0}},\qquad\bar{R}=\frac{1}{G}\sum_{j=1}^{G}R^{(j)}.(2) ### 2.2From Outcome Contribution to Bayesian Turn Evidence Directly measuring the counterfactual contribution of turnkkwould require marginalizing the outcome reward over all possible continuations followingaka_{k}, which is intractable in long-horizon interactions. We therefore adopt a hindsight-based evidential perspective. LetCCdenote the event that the trajectory eventually succeeds. Ifaka_{k}supports success, it should be more characteristic of successful behavior than of unsuccessful behavior. Bayes’ rule expresses the resulting change in the belief aboutCCas an action-side likelihood ratio(Åström,1965 (https://arxiv.org/html/2608.05987#bib.bib69); Kaelblinget al.,1998 (https://arxiv.org/html/2608.05987#bib.bib68)): logit⁡p(C∣sk,ak)−logit⁡p(C∣sk)=log⁡p(ak∣sk,C)p(ak∣sk,¬C).\operatorname{logit}p(C\mid s_{k},a_{k})-\operatorname{logit}p(C\mid s_{k})=\log\frac{p(a_{k}\mid s_{k},C)}{p(a_{k}\mid s_{k},\neg C)}.(3)Herelogit⁡(u)=log⁡u1−u\operatorname{logit}(u)=\log\frac{u}{1-u}. The right-hand side is the ideal Bayes factor(Kass and Raftery,1995 (https://arxiv.org/html/2608.05987#bib.bib62))between the success-conditional and failure-conditional likelihoods ofaka_{k}. Its sign indicates whether the action increases or decreases support for eventual success. Because these outcome-conditional behavioral distributions are not directly available, we estimate this Bayesian evidence retrospectively using a computable per-turn self-distillation contrast between a privileged, success-associated self-teacher and the standard student policy. This contrast provides a tractable proxy for the otherwise inaccessible belief update; we characterize the approximation and its conditions below. The student and teacher share parametersθ\thetaand score the same student-generated action(Zhaoet al.,2026 (https://arxiv.org/html/2608.05987#bib.bib24)). Their token contexts are hk,t=(sk,yk,0(1-\lambda)+\lambda m_{k}\geq 1-\lambda b>0sinceλ≤1,b<1\lambda\leq 1,\,b<1; a strictly positive factor preserves sign. ∎ ###### Proposition 3(Recovery of GRPO). Atλ=0\lambda=0,A~k=A(i)\tilde{A}_{k}=A^{(i)}for every token and theAgentOPSDgradient equals the GRPO gradient. ###### Proof. λ=0\lambda=0gives(1−λ)+λmk=1(1-\lambda)+\lambda m_{k}=1, soA~k=A(i)\tilde{A}_{k}=A^{(i)}identically, independent of the belief signal. ∎ ###### Proposition 4(First-Order Decomposition of the Belief Revision). Forck=γck−1+ekc_{k}=\gamma c_{k-1}+e_{k},lk=logit⁡(B0)+ck\ell_{k}=\operatorname{logit}(B_{0})+c_{k}, andΔlk=ek−(1−γ)ck−1\Delta\ell_{k}=e_{k}-(1-\gamma)c_{k-1}, ΔBk=Bk−1(1−Bk−1)Δlk+O((Δlk)2).\Delta B_{k}=B_{k-1}(1-B_{k-1})\,\Delta\ell_{k}+O\!\big((\Delta\ell_{k})^{2}\big).(18) ###### Proof. Bk=σ(lk)B_{k}=\sigma(\ell_{k})withσ′=σ(1−σ)\sigma^{\prime}=\sigma(1-\sigma); a first-order expansion aroundlk−1\ell_{k-1}givesBk=Bk−1+Bk−1(1−Bk−1)Δlk+O((Δlk)2)B_{k}=B_{k-1}+B_{k-1}(1-B_{k-1})\Delta\ell_{k}+O((\Delta\ell_{k})^{2}). The gateB(1−B)B(1-B)is maximal atB=12B=\tfrac{1}{2}and vanishes asB→{0,1}B\to\{0,1\}. ∎ ###### Proposition 5(Exact Budget of the Idealized Recursion). The idealized increments telescope to the endpoint change,∑kΔBk=BK−B0\sum_{k}\Delta B_{k}=B_{K}-B_{0}. ###### Proof. ∑k=1K(Bk−Bk−1)=BK−B0\sum_{k=1}^{K}(B_{k}-B_{k-1})=B_{K}-B_{0}. ∎ ###### Proposition 6(Non-Identifiability of Per-Turn Contribution). There exist two trajectories with identical outcome reward—hence identical broadcast advantage—whose per-turn contributions differ; per-turn credit is not identifiable from the trajectory return alone. ###### Proof (by construction). Takeτ1,τ2\tau_{1},\tau_{2}in one group withR(τ1)=R(τ2)R(\tau_{1})=R(\tau_{2}), soA(τ1)=A(τ2)A(\tau_{1})=A(\tau_{2})and GRPO assigns the same scalar to every turn. Letτ1\tau_{1}succeed through a single decisive turn (ΔB\Delta Bconcentrated) andτ2\tau_{2}through evenly spread progress. The returns coincide but the per-turn contributions d

Similar Articles

@Xudong07452910: A classic challenge in RL training of LLM agents: after a long task fails, where should the model start learning? The final reward can usually only tell the agent 'success' or 'failure', but it's hard to pinpoint which intermediate judgments are worth keeping and which actions led the entire trajectory astray. This paper proposes SEED, using 'self-evolving online distillation...'

X AI KOLs Timeline

This paper proposes SEED, a method that internalizes post-hoc skills from trajectories into model parameters through self-evolving online distillation, solving the reward sparsity problem in long-horizon RL training, achieving significant improvements on benchmarks such as ALFWorld.

@Xudong07452910: Agent memory is most dangerous when it trusts the past too much. Many Memory Agents stuff similar experiences directly into context after retrieval. But similar tasks do not mean the current state is the same; old experiences can sometimes steer decisions astray. This paper proposes MemHarness, turning Agent...

X AI KOLs Timeline

MemHarness proposes changing Agent memory from simple replay to reconstruction based on the current state, trained end-to-end with GRPO, significantly improving success rates on ALFWorld and WebShop.

@xiaohu: Yesterday, I saw many people sharing Apodex 1.1, an AI agent specifically built for deep research to solve those hard problems that 'have no ready-made answers and require extensive investigation'. Curious, I tested it with two tasks, and they ran all afternoon without finishing. The execution time is indeed long. This agent can, as long as you give it a goal, run for an extended period…

X AI KOLs Timeline

Apodex 1.1 is an AI agent designed for deep research, capable of handling complex tasks that require extensive investigation. It uses a main agent to decompose problems and asynchronously dispatches multiple sub-agents for execution, supporting long-running operations and automatic recovery.

@wsl8297: When running complex tasks with AI agents, the most painful thing is often not that the model isn't strong enough, but that as the conversation gets longer, the context starts to overflow. You have to keep filling in background details, re-explaining the process, plus the redundant logs from tool calls — tokens just gush out like a broken pipe. Recently, I saw TencentDB Agent Memory open-sourced by Tencent...

X AI KOLs Timeline

Tencent has open-sourced TencentDB Agent Memory, which solves the AI agent long-context overflow problem through hierarchical memory management (symbolic short-term memory + hierarchical long-term memory). Benchmarks show token consumption reduced by up to 61% and task success rate improved by over 50%.

@Xudong07452910: Recently, while looking into agent self-evolution research, I've been focusing on a question: When are the trajectories left by agents worth continued learning for the next round of training? This time, I took a self-evolving agent paper from SEED and ran it entirely in Apodex…

X AI KOLs Timeline

This article discusses the importance of trajectory learning in AI agent self-evolution and introduces the Apodex 1.1 system and open-source tool FrontierAgent for executing and evaluating long-duration research tasks.