T2SPO:面向智能体强化学习的轨迹到步骤策略优化
摘要
T2SPO 使用历史交互轨迹提供步骤级奖励信号来优化 LLM 智能体的强化学习策略:借助冻结的 TabPFN 回归器估计每个状态到成功的剩余距离,并将其变化作为辅助信用分配来增强 GRPO,实验在 ALFWorld 和 WebShop 上提升了任务成功率。
查看缓存全文
缓存时间: 2026/10/03 09:53
# T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning
Source: [https://arxiv.org/html/2610.00388](https://arxiv.org/html/2610.00388)
Bo\-Wen Zhang1,2,3,\*†\\daggerJunwei He3,\*Maoqi Liu3Feiran Li3Song\-Lin Lv1,2Wentao Ma3Rongyi Lin3,‡\\ddaggerShuhan Zhong3Lan\-Zhe Guo1,2,§1State Key Laboratory of Novel Software Technology, Nanjing University2School of Intelligence Science and Technology, Nanjing University3ByteDanceguolz@nju\.edu\.cn
###### Abstract
Reinforcement learning enables large language model \(LLM\) agents to learn multi\-step behaviors through interaction with their environments\. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task\. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions\. We introduceTrajectory\-to\-Step Policy Optimization \(T2SPO\), a method that uses past interaction trajectories to provide step\-level feedback for policy learning\. T2SPO derives remaining\-distance targets from successful trajectories and pairs them with representations of the states visited along the way\. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout\. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task\-level supervision\. As training proceeds, newly completed trajectories refresh the estimator’s context, incorporating new experience without updating its parameters\. Experiments with 1\.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO\.
00footnotetext:\*Equal contribution\.†\\daggerWork done during internship at ByteDance\.‡\\ddaggerProject lead\.§Corresponding author\.Figure 1:T2SPO converts historical trajectories into step\-level credit\.\(a\) Context construction\.Successful historical trajectories pair each visited state with its own observed number of turns to completion \(Section[4\.1](https://arxiv.org/html/2610.00388#S4.SS1)\)\. State representations and these targets form a labeled context\.\(b\) Step credit for policy optimization\.A frozen TabPFN predicts the remaining distance at adjacent states of a current rollout using the same historical context\. Their difference is normalized, gated, and scaled into auxiliary step credit, which augments the outcome advantage in GRPO\. Current trajectories refresh the context only after completion and scoring\.## 1 Introduction
Reinforcement learning enables LLM agents to improve through environment interaction\([Zhou et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib25);[Wang et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib19)\)\. In multi\-step tasks, the final outcome evaluates the whole trajectory but does not reveal how individual decisions contribute to success\.Estimating task progress and its changes across stepshelps connect individual decisions to task outcomes, providing fine\-grained signals for credit assignment during policy optimization\([Zhou et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib25);[Li et al\., 2026](https://arxiv.org/html/2610.00388#bib.bib8)\)\.
PPO uses a learned critic to estimate expected returns from intermediate states\([Schulman et al\., 2017](https://arxiv.org/html/2610.00388#bib.bib15)\)\. Combined with generalized advantage estimation, these predictions provide state\-dependent training signals for individual actions\([Schulman et al\., 2016](https://arxiv.org/html/2610.00388#bib.bib14)\)\. This requires training a value model alongside the policy\. GRPO avoids training a critic by estimating its baseline from the rewards of trajectories sampled for the same task\([Shao et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib16)\)\. Under outcome\-only supervision, all steps within a trajectory share the same outcome advantage, even when some decisions advance the task and others are redundant\([Feng et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib4)\)\. If all trajectories in a group receive the same outcome reward, these advantages are zero\. This motivates our central question:Can historical experience supply fine\-grained step credit without training an additional value model?
Successful trajectories collected during policy training provide supervision beyond their final outcomes: each visited state can be paired with the number of turns remaining along its own successful continuation\. These observed lengths serve as a proxy for task progress, providing supervision for estimating remaining distance at states encountered in later rollouts\. Changes in the predicted distance across consecutive states can then supply auxiliary credit for individual agent steps\.
In this paper, we introduceTrajectory\-to\-Step Policy Optimization \(T2SPO\), which implements this approach with a frozen TabPFN regressor\([Müller et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib10);[Hollmann et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib6)\)\. For each rollout batch, the regressor predicts remaining distance using a shared context of historical state representations and completion\-length labels\. Adjacent\-state predictions define a potential difference that provides auxiliary step credit\([Ng et al\., 1999](https://arxiv.org/html/2610.00388#bib.bib11)\)\. Newly completed trajectories update the context after scoring, allowing subsequent estimates to incorporate new experience while the encoder and regressor remain frozen\. The credit augments outcome\-based advantages in our main implementation and can also enter an actor–critic optimizer as an auxiliary reward \(Figure[1](https://arxiv.org/html/2610.00388#S0.F1)\)\.
Our contributions are:
- •We proposeT2SPO, a framework that reuses historical successful trajectories to supply auxiliary credit for individual agent steps while retaining task\-level supervision\.
- •We develop a progress estimator with a frozen TabPFN regressor and online context updates, yielding auxiliary step credit without training an additional scoring model\.
- •Experiments on ALFWorld and WebShop show higher overall success than the reported GRPO baselines at both scales\([Shridhar et al\., 2021](https://arxiv.org/html/2610.00388#bib.bib17);[Yao et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib20)\)\. Offline evaluations assess distance prediction across context sizes and SVD dimensions\.
## 2 Related Work
Reinforcement learning for LLM agents\.Reinforcement learning is widely used in language\-model post\-training, with feedback drawn from human preferences or verifiable task outcomes\([Ouyang et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib12);[DeepSeek\-AI et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib1);[Yu et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib22)\)\. For LLM agents, RL optimizes policies over multiple rounds of interaction, where earlier actions influence later decisions and eventual task outcomes\([Zhou et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib25);[Wang et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib19);[Jin et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib7)\)\. Sparse or delayed feedback makes it difficult to determine which actions contribute to success, motivating methods that connect outcomes to individual steps\([Feng et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib4)\)\.
Step\-level feedback and credit assignment\.Step\-level feedback can be obtained by training a model to evaluate intermediate states or actions\. For mathematical reasoning, process reward models learn from human step annotations or automatically generated supervision: Let’s Verify Step by Step uses human judgments, while Math\-Shepherd derives labels from sampled continuations\([Lightman et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib9);[Wang et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib18)\)\. For LLM agents, ArCHer and Turn\-PPO learn critics that estimate expected returns at the turn level and support policy updates\([Zhou et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib25);[Li et al\., 2026](https://arxiv.org/html/2610.00388#bib.bib8)\)\. Another approach compares returns following different actions from a shared state\. GiGPO identifies repeated environment states in collected trajectories and groups the corresponding actions to compute relative advantages\([Feng et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib4)\)\. ARPO instead uses entropy to guide branching during rollout collection and assigns advantages according to the shared and branched portions of the resulting trajectories\([Dong et al\., 2026a](https://arxiv.org/html/2610.00388#bib.bib2)\)\. Neither method trains an auxiliary scoring model\.
Tabular in\-context learning for value estimation\.Prior\-data fitted networks learn to make supervised predictions from labeled context examples by pretraining over synthetic tasks\([Müller et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib10)\)\. Tabular foundation models extend this approach to classification and regression, with TabPFN emphasizing prediction from small datasets and TabICL scaling in\-context classification to larger tables\([Hollmann et al\., 2023](https://arxiv.org/html/2610.00388#bib.bib5);[Hollmann et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib6);[Qu et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib13)\)\. Closely related applications areV0V\_\{0\}andV0\.5V\_\{0\.5\}, which combine semantic representations with TabPFN\-based prediction conditioned on prior instruction–performance pairs\([Zhang et al\., 2026b](https://arxiv.org/html/2610.00388#bib.bib24);[Zhang et al\., 2026a](https://arxiv.org/html/2610.00388#bib.bib23)\)\.V0V\_\{0\}predicts success probability at the initial prompt to allocate rollout budgets and route instructions, whileV0\.5V\_\{0\.5\}combines such predictions with sparse on\-policy outcomes for baseline estimation and adaptive sampling\. T2SPO instead conditions a frozen TabPFN regressor on representations of intermediate states paired with the remaining lengths of their successful source trajectories\. Differences in predicted remaining distance between adjacent states supply auxiliary step credit for policy optimization\.
## 3 Preliminaries
We first define the interaction and optimization setting, then introduce the regression interface used to transfer trajectory\-derived supervision to current states\.
### 3\.1 Multi\-Turn Reinforcement Learning
LLM agents interact with their environments by selecting actions conditioned on observations and prior context\([Yao et al\., 2023](https://arxiv.org/html/2610.00388#bib.bib21)\)\. We use*step*, or*turn*, to denote one complete agent–environment interaction\. Given an instructionqq, letsts\_\{t\}denote the context available to the policy before turntt, typically consisting of the instruction, the current observationoto\_\{t\}, and processed interaction history\([Dong et al\., 2026b](https://arxiv.org/html/2610.00388#bib.bib3)\)\. The policyπθ\(at∣st\)\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)generates a textual actionata\_\{t\}autoregressively, after which the environment returns rewardrtr\_\{t\}and observationot\+1o\_\{t\+1\}\. A trajectoryτ\\taucomprisesTTsuch transitions, with sparse task feedback provided by the outcome at the end of the episode\.
Advantage estimation\.Outcome\-supervised GRPO samples a group ofG\>1G\>1trajectories for the same task and normalizes their final task scoresRiR\_\{i\}\([Shao et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib16)\):
A^iout=Ri−R¯σR\+ε,\\hat\{A\}\_\{i\}^\{\\mathrm\{out\}\}=\\frac\{R\_\{i\}\-\\overline\{R\}\}\{\\sigma\_\{R\}\+\\varepsilon\},\(1\)whereR¯\\overline\{R\}andσR\\sigma\_\{R\}denote the group mean and standard deviation, andε\>0\\varepsilon\>0ensures numerical stability\. This estimator assigns the same advantage to all generated decisions in trajectoryiiwithout a learned critic\. For example, an above\-average outcome reinforces both a redundant search and the successful purchase that follows it\. When all group outcomes are equal, the outcome advantages are zero even if the trajectories differ in intermediate progress or redundant behavior\.
Learned critics\.PPO can obtain advantages that vary across turns through a learned value function and generalized advantage estimation \(GAE\)\([Schulman et al\., 2017](https://arxiv.org/html/2610.00388#bib.bib15);[Schulman et al\., 2016](https://arxiv.org/html/2610.00388#bib.bib14)\)\. At a nonterminal turn with zero environment reward, its temporal\-difference residual isδt=γVψ\(st\+1\)−Vψ\(st\)\\delta\_\{t\}=\\gamma V\_\{\\psi\}\(s\_\{t\+1\}\)\-V\_\{\\psi\}\(s\_\{t\}\), whereγ∈\(0,1\]\\gamma\\in\(0,1\]is the discount factor\. These values let task outcomes inform updates to earlier decisions\.
Policy updates\.We use the clipped surrogate\([Schulman et al\., 2017](https://arxiv.org/html/2610.00388#bib.bib15);[Shao et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib16)\)
ℓclip\(ρ,A^\)=min\{ρA^,clip\[1−ϵ,1\+ϵ\]\(ρ\)A^\},\\ell\_\{\\mathrm\{clip\}\}\(\\rho,\\hat\{A\}\)=\\min\\bigl\\\{\\rho\\hat\{A\},\\operatorname\{clip\}\_\{\[1\-\\epsilon,\\,1\+\\epsilon\]\}\(\\rho\)\\hat\{A\}\\bigr\\\},\(2\)whereclip\[l,u\]\\operatorname\{clip\}\_\{\[l,u\]\}clips its argument to\[l,u\]\[l,u\],ϵ\>0\\epsilon\>0is the clipping threshold, andρ\\rhois the likelihood ratio between the current policy and the fixed rollout policyπold\\pi\_\{\\mathrm\{old\}\}\. For GRPO, ratios are computed over generated tokens; environment observations only condition the policy\. Section[4\.4](https://arxiv.org/html/2610.00388#S4.SS4)describes the integration of auxiliary step credit\. Appendix[A\.5](https://arxiv.org/html/2610.00388#A1.SS5)details PPO critic targets, GAE, and complete\-turn likelihood ratios\.
### 3\.2 Prior\-Data Fitted Networks
A prior\-data fitted network \(PFN\) predicts the target for a queryxxfrom a labeled context𝒞=\{\(xj,yj\)\}j=1m\\mathcal\{C\}=\\\{\(x\_\{j\},y\_\{j\}\)\\\}\_\{j=1\}^\{m\}\([Müller et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib10)\)\. For regression,xj,x∈ℝdx\_\{j\},x\\in\\mathbb\{R\}^\{d\}andyj∈ℝy\_\{j\}\\in\\mathbb\{R\}\. PFNs are pretrained on synthetic supervised learning tasks to approximate posterior predictive inference under a task prior\. At deployment, a pretrained regressor evaluates
y^=fϕ\(x,𝒞\),\\hat\{y\}=f\_\{\\phi\}\(x;\\mathcal\{C\}\),\(3\)with fixed parametersϕ\\phi; adapting the context changes the prediction without optimizing those parameters\. TabPFN implements this approach for tabular classification and regression\([Hollmann et al\., 2023](https://arxiv.org/html/2610.00388#bib.bib5);[Hollmann et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib6)\)\. T2SPO uses this prediction interface with interaction states as queries and historical states paired with trajectory\-derived labels as context\. Section[4](https://arxiv.org/html/2610.00388#S4)specifies how to construct those labels and convert the resulting state predictions into auxiliary signals for policy optimization\.
## 4 Trajectory\-to\-Step Policy Optimization
T2SPO pairs historical states with labels derived from successful trajectories, uses these examples to predict remaining distance in subsequent rollouts, and converts adjacent\-state differences into auxiliary step credit \(Figure[1](https://arxiv.org/html/2610.00388#S0.F1)\)\. We describe target construction, context\-based prediction, and integration into policy updates\.
### 4\.1 Supervision from Historical Trajectories
A completed successful trajectory provides more than a terminal score: each visited state has an observed number of turns before success\. For each successful trajectoryτj\\tau\_\{j\}of lengthTjT\_\{j\}, we pair its decision statesj,ts\_\{j,t\}with the number of interactions remaining along that trajectory:
yj,t=Tj−t,t=0,…,Tj−1\.y\_\{j,t\}=T\_\{j\}\-t,\\qquad t=0,\\ldots,T\_\{j\}\-1\.\(4\)Each label records the remaining length of its own successful trajectory\. These targets provide a success\-conditioned distance proxy reflecting continuations observed under the collection policies\.
Unsuccessful trajectories do not reveal a completion distance\. The success\-only variant excludes their states from regression\. A failure\-augmented variant assigns them the rollout horizonHHas a heuristic target and caps their share of the context\. This target is not an observed time to success; the main experiments use success\-only context \(Appendix[A\.1](https://arxiv.org/html/2610.00388#A1.SS1)\)\.
### 4\.2 Estimating Current States from Past Trajectories
State representation\.To use historical targets at newly visited states, we express both reference and query states in a shared feature space\. Lethth\_\{t\}denote the instruction and interaction record available before turntt, including the current observation\. A frozen text encoder embeds the instruction, current observation, and the selected history fromhth\_\{t\}\. The estimator’s history window is configured separately from the policy contextsts\_\{t\}\(Appendix[A\.2](https://arxiv.org/html/2610.00388#A1.SS2)\)\. We concatenate numerical step and collection\-round features, then apply standardization and truncated singular value decomposition to obtainxt∈ℝdx\_\{t\}\\in\\mathbb\{R\}^\{d\}\. Both transformations are fitted on the selected historical examples and shared by reference and query states\. The current actionata\_\{t\}, future observations, and the episode outcome are excluded fromxtx\_\{t\}\. This representation allows the estimator to compare states through their features without requiring exact state matches across trajectories\.
Shared trajectory context\.At training roundnn, a bounded historical memoryℳn\\mathcal\{M\}\_\{n\}supplies the labeled context𝒞n=\{\(xj,t,yj,t\)\}\(j,t\)∈ℐn\\mathcal\{C\}\_\{n\}=\\\{\(x\_\{j,t\},y\_\{j,t\}\)\\\}\_\{\(j,t\)\\in\\mathcal\{I\}\_\{n\}\}\. Hereℐn\\mathcal\{I\}\_\{n\}indexes the selected historical trajectory–turn pairs, with\|ℐn\|=Mn\|\\mathcal\{I\}\_\{n\}\|=M\_\{n\}\. Successful state examples are sampled from completed batches and retained in a finite pool, with older entries evicted at capacity\. The failure\-augmented variant additionally samples unsuccessful states subject to its quota\. A pretrained TabPFN regressorfϕf\_\{\\phi\}conditions on this context to predict
d^n,t=clip\[0,Dmax\]\(fϕ\(xt,𝒞n\)\),\\hat\{d\}\_\{n,t\}=\\operatorname\{clip\}\_\{\[0,D\_\{\\max\}\]\}\\bigl\(f\_\{\\phi\}\(x\_\{t\};\\mathcal\{C\}\_\{n\}\)\\bigr\),\(5\)whereDmaxD\_\{\\max\}bounds the numerical output\. All batch queries share𝒞n\\mathcal\{C\}\_\{n\}and its fitted feature transformation, including adjacent states used for step credit\. Context updates refresh the labeled examples and their feature transformation, changing predictions without updatingϕ\\phi\([Müller et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib10);[Hollmann et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib6)\)\.
Prediction before insertion\.We score all current rollouts using𝒞n\\mathcal\{C\}\_\{n\}before their outcomes can supply labels for𝒞n\+1\\mathcal\{C\}\_\{n\+1\}\. This ordering lets a trajectory inform future updates while excluding its own outcome from the context used to evaluate it\. The main experiments initialize the context from a separate trajectory archive \(Appendix[A\.2](https://arxiv.org/html/2610.00388#A1.SS2)\)\. With empty memory, training uses the base signal until sufficient historical support is available\.
### 4\.3 From State Estimates to Step Credit
A remaining\-distance estimate describes a state, whereas policy optimization requires a signal for the decision that changes it\. T2SPO compares estimates before and after each transition under the same historical context\. DefineΦn\(ht\)=−d^n,t\\Phi\_\{n\}\(h\_\{t\}\)=\-\\hat\{d\}\_\{n,t\}and omitnnbelow when unambiguous\. Using a potential difference\([Ng et al\., 1999](https://arxiv.org/html/2610.00388#bib.bib11)\), we obtain
Ft=γΦn\(ht\+1\)−Φn\(ht\)=d^t−γd^t\+1\.F\_\{t\}=\\gamma\\Phi\_\{n\}\(h\_\{t\+1\}\)\-\\Phi\_\{n\}\(h\_\{t\}\)=\\hat\{d\}\_\{t\}\-\\gamma\\hat\{d\}\_\{t\+1\}\.\(6\)Hereγ∈\(0,1\]\\gamma\\in\(0,1\]is the discount factor\. Whenγ=1\\gamma=1, a positive difference indicates a reduction in predicted remaining distance\. Successful termination sets the successor potential to zero\. Unsuccessful termination receives zero final process signal, avoiding credit caused solely by ending an unsuccessful episode\. At truncation, PPO evaluates the actual successor state; GRPO sets the final process signal to zero\.
To stabilize the auxiliary scale, letzt=arsinh\(Ft\)z\_\{t\}=\\operatorname\{arsinh\}\(F\_\{t\}\)and letℬn\\mathcal\{B\}\_\{n\}contain the batch turns selected by the normalization mask\. We compute the root\-mean\-square scale and normalized signal as
νn2=1\|ℬn\|∑u∈ℬnzu2,F~t=clip\[−c,c\]\(ztνn\+ε\)\.\\nu\_\{n\}^\{2\}=\\frac\{1\}\{\|\\mathcal\{B\}\_\{n\}\|\}\\sum\_\{u\\in\\mathcal\{B\}\_\{n\}\}z\_\{u\}^\{2\},\\qquad\\tilde\{F\}\_\{t\}=\\operatorname\{clip\}\_\{\[\-c,c\]\}\\left\(\\frac\{z\_\{t\}\}\{\\nu\_\{n\}\+\\varepsilon\}\\right\)\.\(7\)Herec\>0c\>0limits individual contributions andε\>0\\varepsilon\>0ensures numerical stability\. The step credit supplied to the optimizer is
pi,t=wngi,tF~i,t,p\_\{i,t\}=w\_\{n\}g\_\{i,t\}\\tilde\{F\}\_\{i,t\},\(8\)wherewnw\_\{n\}ramps up after burn\-in andgi,t∈\[0,1\]g\_\{i,t\}\\in\[0,1\]controls whether a prediction contributes to the update\. With sufficient support, valid predictions have gate one; insufficient support or rejected predictions disable the contribution\. The ALFWorld configuration also suppresses unchanged observations and excludes them from normalization\. Turns outside the mask receive zero credit; all turns do so whenℬn\\mathcal\{B\}\_\{n\}is empty\.
Normalization, gating, and terminal exceptions modify the potential difference, so the standard policy\-invariance guarantee does not extend to the full T2SPO procedure\. Algorithm[1](https://arxiv.org/html/2610.00388#alg1)summarizes step\-credit computation for the main success\-only configuration\.
Algorithm 1Computing Auxiliary Step Credit1:Memory
ℳn\\mathcal\{M\}\_\{n\}, rollouts
𝒯n\\mathcal\{T\}\_\{n\}, frozen encoder and regressor
fϕf\_\{\\phi\}, scale
wnw\_\{n\}
2:Step credits
\{pi,t\}\\\{p\_\{i,t\}\\\}and updated memory
ℳn\+1\\mathcal\{M\}\_\{n\+1\}
3:Initialize
pi,t←0p\_\{i,t\}\\leftarrow 0for all turns in
𝒯n\\mathcal\{T\}\_\{n\}
4:Select successful historical support states from
ℳn\\mathcal\{M\}\_\{n\}
5:Assign each support state its source trajectory’s remaining length \(Equation[4](https://arxiv.org/html/2610.00388#S4.E4)\)
6:ifsufficient support is availablethen
7:Encode support and query states; append step and collection\-round features
8:Fit the feature transformation on support; apply it to both sets
9:Form the labeled context
𝒞n\\mathcal\{C\}\_\{n\}from transformed support states
10:Predict current and successor distances with the same
𝒞n\\mathcal\{C\}\_\{n\}\(Equation[5](https://arxiv.org/html/2610.00388#S4.E5)\)
11:Set successor distance to zero at successful termination
12:Compute
Fi,t←d^i,t−γd^i,t\+1F\_\{i,t\}\\leftarrow\\hat\{d\}\_\{i,t\}\-\\gamma\\hat\{d\}\_\{i,t\+1\}
13:Apply terminal and truncation rules from Section[4\.3](https://arxiv.org/html/2610.00388#S4.SS3)
14:Determine gates
gi,tg\_\{i,t\}and normalization set
ℬn\\mathcal\{B\}\_\{n\}\(Section[4\.3](https://arxiv.org/html/2610.00388#S4.SS3)\)
15:if
ℬn≠∅\\mathcal\{B\}\_\{n\}\\neq\\varnothingthen
16:
zi,t←arsinh\(Fi,t\)z\_\{i,t\}\\leftarrow\\operatorname\{arsinh\}\(F\_\{i,t\}\)for
\(i,t\)∈ℬn\(i,t\)\\in\\mathcal\{B\}\_\{n\}
17:
νn←\|ℬn\|−1∑\(i,t\)∈ℬnzi,t2\\nu\_\{n\}\\leftarrow\\sqrt\{\|\\mathcal\{B\}\_\{n\}\|^\{\-1\}\\sum\_\{\(i,t\)\\in\\mathcal\{B\}\_\{n\}\}z\_\{i,t\}^\{2\}\}
18:
pi,t←wngi,tclip\[−c,c\]\(zi,t/\(νn\+ε\)\)p\_\{i,t\}\\leftarrow w\_\{n\}g\_\{i,t\}\\operatorname\{clip\}\_\{\[\-c,c\]\}\\\!\\left\(z\_\{i,t\}/\(\\nu\_\{n\}\+\\varepsilon\)\\right\)for
\(i,t\)∈ℬn\(i,t\)\\in\\mathcal\{B\}\_\{n\}
19:endif
20:endif
21:Update memory with sampled successful states from
𝒯n\\mathcal\{T\}\_\{n\}, obtaining
ℳn\+1\\mathcal\{M\}\_\{n\+1\}
22:return
\{pi,t\},ℳn\+1\\\{p\_\{i,t\}\\\},\\mathcal\{M\}\_\{n\+1\}
### 4\.4 Policy Optimization with Step Credit
The auxiliary term adds local information to the task\-level learning signal\. We retain the base optimizer and specify how the same step credit enters its training targets\.
Group\-relative optimization\.LetA^i,tbase\\hat\{A\}^\{\\mathrm\{base\}\}\_\{i,t\}be the outcome\-based advantage supplied by the GRPO implementation\([Shao et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib16)\)\. T2SPO augments it at each turn:
A^i,tT2SPO=A^i,tbase\+pi,t\.\\hat\{A\}^\{\\mathrm\{T2SPO\}\}\_\{i,t\}=\\hat\{A\}^\{\\mathrm\{base\}\}\_\{i,t\}\+p\_\{i,t\}\.\(9\)The resulting scalar is assigned to the generated tokens of that turn and used in the token\-level clipped surrogate \(Equation[2](https://arxiv.org/html/2610.00388#S3.E2)\)\. The outcome component preserves task\-level supervision, whilepi,tp\_\{i,t\}introduces variation among decisions within each trajectory\. We retain the base implementation’s group normalization and KL regularization and introduce no learned critic\. Appendix[A\.4](https://arxiv.org/html/2610.00388#A1.SS4)details the multi\-turn aggregation convention\.
Actor–critic extension\.The same step credit can instead augment rewards before PPO advantage estimation:
ri,t\+=ri,tenv\+pi,t\.r^\{\+\}\_\{i,t\}=r^\{\\mathrm\{env\}\}\_\{i,t\}\+p\_\{i,t\}\.\(10\)The environment reward includes task outcomes and any configured invalid\-action penalty\. PPO computes advantages from these augmented rewards using GAE and trains its critic on the corresponding return targets\([Schulman et al\., 2016](https://arxiv.org/html/2610.00388#bib.bib14);[Li et al\., 2026](https://arxiv.org/html/2610.00388#bib.bib8)\)\. This integration uses complete\-turn likelihood ratios; Appendix[A\.5](https://arxiv.org/html/2610.00388#A1.SS5)specifies the advantages, bootstrap handling, and objective\.
The two integrations use the same source of step credit but apply it at different points: GRPO augments advantages directly, whereas PPO propagates augmented rewards through GAE\. Both keep the current batch’s targets fixed during optimization and incorporate completed trajectories into subsequent contexts\. The encoder and TabPFN receive no gradient updates\.
## 5 Experiments
### 5\.1 Experimental Setup
Benchmarks and metrics\.We evaluate on ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2610.00388#bib.bib17)\)and WebShop\([Yao et al\., 2022](https://arxiv.org/html/2610.00388#bib.bib20)\), which require agents to complete tasks through multiple rounds of environment interaction\. For ALFWorld, we report success rates for six task categories—Pick, Look, Clean, Heat, Cool, and Pick2—and overall success\. For WebShop, we report the average task score on a 0–100 scale and the success rate\. All success rates are expressed as percentages\.
Models and baselines\.We use Qwen2\.5\-1\.5B\-Instruct and Qwen2\.5\-7B\-Instruct as policy models\. Our primary comparison is with GRPO\([Shao et al\., 2024](https://arxiv.org/html/2610.00388#bib.bib16)\), whose outcome\-based training signal is augmented by T2SPO\. We also include the prompting\-only base model, ReAct\([Yao et al\., 2023](https://arxiv.org/html/2610.00388#bib.bib21)\), PPO, RLOO, and GiGPO with standard\-deviation normalization\([Feng et al\., 2025](https://arxiv.org/html/2610.00388#bib.bib4)\)\. All baseline results are taken from[Feng et al\. \(2025\)](https://arxiv.org/html/2610.00388#bib.bib4)\.
Training details\.The main T2SPO\-GRPO configuration uses a context budget of 256 state examples drawn from historical successful trajectories\. We evaluate checkpoints after 150 training updates\. Training, estimator, and context settings are detailed in Appendix[A](https://arxiv.org/html/2610.00388#A1)\.
### 5\.2 Main Results
Table 1:Main results on ALFWorld and WebShop\.ALFWorld reports success rates \(%\); WebShop reports score \(0–100\) and success rate \(%\)\. RL results are means over three seeds; subscripts denote standard deviations\. Shaded rows denote T2SPO\-GRPO using only successful context examples; bold numbers indicate the highest reported mean within each model size\.Δ\\Deltareports the difference between T2SPO\-GRPO and the GRPO reference within each model size: percentage points for success rates and points for WebShop score\.TypeMethodALFWorldWebShopPickLookCleanHeatCoolPick2AllScoreSucc\.Qwen2\.5\-1\.5B\-InstructPromptingQwen2\.55\.95\.95\.55\.53\.33\.39\.79\.74\.24\.20\.00\.04\.14\.123\.123\.15\.25\.2PromptingReAct17\.417\.420\.520\.515\.715\.76\.26\.27\.77\.72\.02\.012\.812\.840\.140\.111\.311\.3RL TrainingPPO \(with critic\)64\.8±3\.564\.8\_\{\\pm 3\.5\}40\.5±6\.940\.5\_\{\\pm 6\.9\}57\.1±4\.957\.1\_\{\\pm 4\.9\}60\.6±6\.660\.6\_\{\\pm 6\.6\}46\.4±4\.046\.4\_\{\\pm 4\.0\}47\.4±1\.947\.4\_\{\\pm 1\.9\}54\.4±3\.154\.4\_\{\\pm 3\.1\}73\.8±3\.073\.8\_\{\\pm 3\.0\}51\.5±2\.951\.5\_\{\\pm 2\.9\}RL TrainingRLOO88\.3±3\.088\.3\_\{\\pm 3\.0\}52\.8±8\.652\.8\_\{\\pm 8\.6\}71\.0±5\.971\.0\_\{\\pm 5\.9\}62\.8±8\.762\.8\_\{\\pm 8\.7\}66\.4±5\.566\.4\_\{\\pm 5\.5\}56\.9±4\.756\.9\_\{\\pm 4\.7\}69\.7±2\.569\.7\_\{\\pm 2\.5\}73\.9±5\.673\.9\_\{\\pm 5\.6\}52\.1±6\.752\.1\_\{\\pm 6\.7\}RL TrainingGRPO85\.3±1\.585\.3\_\{\\pm 1\.5\}53\.7±8\.053\.7\_\{\\pm 8\.0\}84\.5±6\.884\.5\_\{\\pm 6\.8\}78\.2±7\.978\.2\_\{\\pm 7\.9\}59\.7±5\.059\.7\_\{\\pm 5\.0\}53\.5±5\.653\.5\_\{\\pm 5\.6\}72\.8±3\.672\.8\_\{\\pm 3\.6\}75\.8±3\.575\.8\_\{\\pm 3\.5\}56\.8±3\.856\.8\_\{\\pm 3\.8\}RL TrainingGiGPO94\.4±5\.9\\mathbf\{94\.4\}\_\{\\pm 5\.9\}67\.5±4\.6\\mathbf\{67\.5\}\_\{\\pm 4\.6\}94\.8±3\.8\\mathbf\{94\.8\}\_\{\\pm 3\.8\}94\.4±7\.8\\mathbf\{94\.4\}\_\{\\pm 7\.8\}79\.8±4\.779\.8\_\{\\pm 4\.7\}76\.4±5\.4\\mathbf\{76\.4\}\_\{\\pm 5\.4\}86\.7±1\.7\\mathbf\{86\.7\}\_\{\\pm 1\.7\}83\.1±1\.683\.1\_\{\\pm 1\.6\}65\.0±3\.265\.0\_\{\\pm 3\.2\}RL TrainingT2SPO\-GRPO87\.4±7\.487\.4\_\{\\pm 7\.4\}56\.8±15\.956\.8\_\{\\pm 15\.9\}90\.3±10\.190\.3\_\{\\pm 10\.1\}65\.9±18\.265\.9\_\{\\pm 18\.2\}86\.5±3\.0\\mathbf\{86\.5\}\_\{\\pm 3\.0\}43\.7±23\.943\.7\_\{\\pm 23\.9\}74\.7±10\.074\.7\_\{\\pm 10\.0\}88\.5±0\.3\\mathbf\{88\.5\}\_\{\\pm 0\.3\}73\.7±3\.3\\mathbf\{73\.7\}\_\{\\pm 3\.3\}Δ\\Deltavs\. GRPO\+2\.1\+2\.1\+3\.1\+3\.1\+5\.8\+5\.8−12\.3\-12\.3\+26\.8\+26\.8−9\.8\-9\.8\+1\.9\+1\.9\+12\.7\+12\.7\+16\.9\+16\.9Qwen2\.5\-7B\-InstructPromptingQwen2\.533\.433\.421\.621\.619\.319\.36\.96\.92\.82\.83\.23\.214\.814\.826\.426\.47\.87\.8PromptingReAct48\.548\.535\.435\.434\.334\.313\.213\.218\.218\.217\.617\.631\.231\.246\.246\.219\.519\.5RL TrainingPPO \(with critic\)92\.3±4\.092\.3\_\{\\pm 4\.0\}64\.0±8\.464\.0\_\{\\pm 8\.4\}92\.5±2\.492\.5\_\{\\pm 2\.4\}89\.5±7\.0\\mathbf\{89\.5\}\_\{\\pm 7\.0\}80\.3±2\.080\.3\_\{\\pm 2\.0\}68\.8±8\.368\.8\_\{\\pm 8\.3\}80\.4±2\.780\.4\_\{\\pm 2\.7\}81\.4±3\.181\.4\_\{\\pm 3\.1\}68\.7±5\.168\.7\_\{\\pm 5\.1\}RL TrainingRLOO87\.6±4\.387\.6\_\{\\pm 4\.3\}78\.2±8\.378\.2\_\{\\pm 8\.3\}87\.3±5\.887\.3\_\{\\pm 5\.8\}81\.3±7\.681\.3\_\{\\pm 7\.6\}71\.9±5\.271\.9\_\{\\pm 5\.2\}48\.9±8\.448\.9\_\{\\pm 8\.4\}75\.5±4\.675\.5\_\{\\pm 4\.6\}80\.3±3\.280\.3\_\{\\pm 3\.2\}65\.7±4\.065\.7\_\{\\pm 4\.0\}RL TrainingGRPO90\.8±5\.190\.8\_\{\\pm 5\.1\}66\.1±6\.766\.1\_\{\\pm 6\.7\}89\.3±5\.489\.3\_\{\\pm 5\.4\}74\.7±6\.974\.7\_\{\\pm 6\.9\}72\.5±5\.472\.5\_\{\\pm 5\.4\}64\.7±7\.364\.7\_\{\\pm 7\.3\}77\.6±5\.277\.6\_\{\\pm 5\.2\}79\.3±2\.879\.3\_\{\\pm 2\.8\}66\.1±3\.766\.1\_\{\\pm 3\.7\}RL TrainingGiGPO97\.7±1\.6\\mathbf\{97\.7\}\_\{\\pm 1\.6\}82\.7±7\.9\\mathbf\{82\.7\}\_\{\\pm 7\.9\}98\.8±1\.6\\mathbf\{98\.8\}\_\{\\pm 1\.6\}83\.7±7\.283\.7\_\{\\pm 7\.2\}89\.3±8\.2\\mathbf\{89\.3\}\_\{\\pm 8\.2\}79\.2±6\.6\\mathbf\{79\.2\}\_\{\\pm 6\.6\}90\.8±1\.3\\mathbf\{90\.8\}\_\{\\pm 1\.3\}84\.4±2\.984\.4\_\{\\pm 2\.9\}72\.8±3\.272\.8\_\{\\pm 3\.2\}RL TrainingT2SPO\-GRPO91\.3±5\.791\.3\_\{\\pm 5\.7\}73\.8±2\.173\.8\_\{\\pm 2\.1\}97\.3±2\.497\.3\_\{\\pm 2\.4\}84\.5±16\.884\.5\_\{\\pm 16\.8\}78\.4±7\.778\.4\_\{\\pm 7\.7\}63\.5±10\.563\.5\_\{\\pm 10\.5\}83\.6±6\.483\.6\_\{\\pm 6\.4\}85\.4±4\.1\\mathbf\{85\.4\}\_\{\\pm 4\.1\}75\.8±3\.6\\mathbf\{75\.8\}\_\{\\pm 3\.6\}Δ\\Deltavs\. GRPO\+0\.5\+0\.5\+7\.7\+7\.7\+8\.0\+8\.0\+9\.8\+9\.8\+5\.9\+5\.9−1\.2\-1\.2\+6\.0\+6\.0\+6\.1\+6\.1\+9\.7\+9\.7
Qwen2\.5, ReAct, PPO, RLOO, GRPO, and GiGPO results are from[Feng et al\. \(2025\)](https://arxiv.org/html/2610.00388#bib.bib4)\.
Comparison with baselines\.Table[1](https://arxiv.org/html/2610.00388#S5.T1)compares T2SPO\-GRPO with published baseline results\. Overall ALFWorld success is 74\.7% for the 1\.5B model and 83\.6% for the 7B model, exceeding the reported GRPO means by 1\.9 and 6\.0 percentage points, respectively\. On WebShop, the corresponding differences are 16\.9 and 9\.7 percentage points in success rate and 12\.7 and 6\.1 points in task score\. T2SPO\-GRPO also exceeds GiGPO in WebShop score and success at both sizes, while GiGPO retains higher overall ALFWorld success\.
Failure\-augmented context\.Table[2](https://arxiv.org/html/2610.00388#S5.T2)compares success\-only context with a variant that admits up to 25% failure examples, labeled by the rollout horizon\. Both configurations use a context budget of 256 and a process weight of 0\.10\. On ALFWorld, failure augmentation raises mean success from 83\.6% to 89\.3% at 7B, while the 1\.5B mean remains similar at 74\.0% versus 74\.7%\. On WebShop, mean success is lower at both model sizes; task score decreases at 1\.5B but increases slightly at 7B\. The effect of failure examples therefore varies across benchmarks and model sizes\. Appendix[A\.3](https://arxiv.org/html/2610.00388#A1.SS3)reports the ALFWorld task\-category breakdown\.
Table 2:Failure\-augmented context on ALFWorld and WebShop\.Means over three seeds; subscripts report standard deviations\. Bold numbers indicate the higher mean within each model\-size pair\.
### 5\.3 Offline Distance Prediction
We evaluate remaining\-distance prediction on 1,951 states from successful WebShop trajectories generated by Qwen2\.5\-1\.5B\-Instruct\. Context and query trajectories have disjoint task identities, and both use the per\-trajectory targets in Equation[4](https://arxiv.org/html/2610.00388#S4.E4)\. Mean absolute error \(MAE\) is averaged over query states and measured in turns\. Inputs contain the task description, current observation, history, step index, and pre\-rollout metadata, excluding the current action and future information\.
Using the same queries, we test all nine combinations of context budgets 128, 256, and 512 and SVD dimensions 32, 64, and 128\. Larger contexts retain all examples from smaller contexts\.
Figure 2:Offline distance prediction on WebShop\.\(a\) Remaining\-distance MAE across nine context/SVD settings; parameter axes use logarithmic spacing, and facets connect evaluated points\. The square marks the main RL configuration; the star marks the lowest observed MAE\. \(b\) TabPFN and the lowest kNN MAE across five tested neighbor counts per context budget, with matched context examples, SVD32 features, and queries\.Increasing the context budget from the main RL setting of 256 to 512 reduces offline MAE at every tested SVD dimension \(Figure[2](https://arxiv.org/html/2610.00388#S5.F2)a\)\. At SVD32, MAE decreases from 0\.8215 to 0\.6500 turns, a 20\.9% reduction\. The lowest observed MAE is 0\.5948 with 512 context examples and SVD64, while increasing feature dimension does not improve prediction uniformly \(Appendix[B](https://arxiv.org/html/2610.00388#A2)\)\.
### 5\.4 Comparison with Nearest Neighbors
We compare TabPFN with kNN regression at SVD32 across context budgetsM∈\{128,256,512\}M\\in\\\{128,256,512\\\}\. At each budget, both estimators use the same labeled context, fitted feature transformation, and query states\. The kNN baseline averages labels of the nearest context states under cosine similarity\. We testk∈\{8,16,M/4,M/2,M\}k\\in\\\{8,16,M/4,M/2,M\\\}and plot the lowest observed kNN MAE at each budget \(Figure[2](https://arxiv.org/html/2610.00388#S5.F2)b\)\. TabPFN’s MAE decreases from 0\.8526 to 0\.6500, while the best tested kNN MAEs are 0\.9856, 0\.9978, and 1\.0022\. The corresponding MAE reductions are 13\.5%, 17\.7%, and 35\.1%, respectively\. Appendix[B](https://arxiv.org/html/2610.00388#A2)reports the full grid, tested neighbor counts, and offline evaluation protocol\.
### 5\.5 Computation Cost
We profile three WebShop Qwen2\.5\-1\.5B\-Instruct runs with the main context\-256/SVD32 configuration\. Across 360 optimizer steps excluding evaluation and checkpointing, mean step time is 380\.21 s\. In Figure[3](https://arxiv.org/html/2610.00388#S5.F3), rollout generation, actor updates, and advantage computation account for 45\.17%, 22\.31%, and 21\.22% of summed step time, respectively\. Within advantage computation, state estimation accounts for 21\.06% of step time, while preparing the next context takes 0\.075%\. State estimation includes embedding, context construction, SVD, and TabPFN fitting and prediction\. Appendix[C](https://arxiv.org/html/2610.00388#A3)details the protocol and stage timings\.
Figure 3:Online computation cost on WebShop\.\(a\) Mutually exclusive training stages across three seeds and 360 ordinary optimizer steps\. \(b\) Substages within advantage computation; times are means in seconds\. Both panels report shares of summed full\-step time\. Other advantage work includes GRPO advantage computation, shaping, and context commit\.
## 6 Conclusion
T2SPO converts historical completion lengths into auxiliary step credit through a frozen distance estimator whose context evolves with training\. Adjacent\-state differences augment GRPO advantages or PPO rewards, with PPO retaining its learned critic\. T2SPO\-GRPO achieves higher overall ALFWorld success and higher WebShop scores and success rates than the reported GRPO baselines at both model sizes\. In the offline WebShop evaluation, larger contexts reduce prediction error against observed remaining lengths\. At SVD32, TabPFN achieves lower MAE than the best tested kNN settings at all three context budgets\.
Limitations and future work\.The mixed effects of a fixed proportion of failure examples motivate supervision that adapts to the task and the agent’s experience\. For tasks with irreversible dead ends, such as Sokoban, the estimator should distinguish recoverable detours from states where success is no longer possible\. Historical context could evolve through soft updates based on weighted combinations of feature vectors and remaining\-distance targets from semantically compatible states\. When successful trajectories are scarce, memory updates and the strength of auxiliary credit could adapt to the coverage and reliability of available experience\. For more complex tasks, training or fine\-tuning this compact predictor could improve task\-specific progress estimation while keeping it substantially smaller than a separate LLM\-based critic\.
### AI Use Statement
AI tools assisted with manuscript editing, literature search, figure preparation, and the implementation of parts of the code\. They also supported discussions of experimental design and result interpretation\. The authors developed the core research idea, specified the algorithmic procedure, and take responsibility for the final manuscript, references, code, and reported results\.
## References
- DeepSeek\-AI et al\. \(2025\)DeepSeek\-AI et al\.DeepSeek\-R1: Incentivizing reasoning capability in LLMs via reinforcement learning\.arXiv preprint arXiv:2501\.12948, 2025\.URL[https://arxiv\.org/abs/2501\.12948v1](https://arxiv.org/abs/2501.12948v1)\.
- Dong et al\. \(2026a\)Guanting Dong, Hangyu Mao, Kai Ma, Licheng Bao, Yifei Chen, Zhongyuan Wang, Zhongxia Chen, Jiazhen Du, Huiyang Wang, Fuzheng Zhang, Guorui Zhou, Yutao Zhu, Ji\-Rong Wen, and Zhicheng Dou\.Agentic reinforced policy optimization\.In*International Conference on Learning Representations*, 2026a\.
- Dong et al\. \(2026b\)Guanting Dong, Xiaoshuai Song, Yuyang Hu, Jiajie Jin, Chenghao Zhang, Yifei Chen, Xiaoxi Li, Huaying Yuan, Xinyu Yang, Tongyu Wen, Jiejun Tan, Hongjin Qian, Shijue Huang, Junting Lu, Zhenyu Li, Wanjun Zhong, Yutao Zhu, Tat\-Seng Chua, Zhicheng Dou, and Ji\-Rong Wen\.Towards long\-horizon agents: A survey\.*Preprints*, July 2026b\.doi:10\.20944/preprints202607\.1328\.v1\.
- Feng et al\. \(2025\)Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An\.Group\-in\-group policy optimization for LLM agent training\.In*Advances in Neural Information Processing Systems*, volume 38, pp\. 46375–46408\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-1544\.
- Hollmann et al\. \(2023\)Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter\.TabPFN: A transformer that solves small tabular classification problems in a second\.In*International Conference on Learning Representations*, 2023\.
- Hollmann et al\. \(2025\)Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter\.Accurate predictions on small data with a tabular foundation model\.*Nature*, 637\(8045\):319–326, 2025\.doi:10\.1038/s41586\-024\-08328\-6\.
- Jin et al\. \(2025\)Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan O\. Arık, Dong Wang, Hamed Zamani, and Jiawei Han\.Search\-R1: Training LLMs to reason and leverage search engines with reinforcement learning\.In*Second Conference on Language Modeling*, 2025\.
- Li et al\. \(2026\)Junbo Li, Peng Zhou, Rui Meng, Meet P\. Vadera, Lihong Li, and Yang Li\.Turn\-PPO: Turn\-level advantage estimation with PPO for improved multi\-turn RL in agentic LLMs\.In*Findings of the Association for Computational Linguistics: EACL 2026*, pp\. 6227–6243\. Association for Computational Linguistics, 2026\.doi:10\.18653/v1/2026\.findings\-eacl\.328\.
- Lightman et al\. \(2024\)Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In*International Conference on Learning Representations*, 2024\.
- Müller et al\. \(2022\)Samuel Müller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter\.Transformers can do Bayesian inference\.In*International Conference on Learning Representations*, 2022\.
- Ng et al\. \(1999\)Andrew Y\. Ng, Daishi Harada, and Stuart J\. Russell\.Policy invariance under reward transformations: Theory and application to reward shaping\.In*Proceedings of the 16th International Conference on Machine Learning*, pp\. 278–287, 1999\.
- Ouyang et al\. \(2022\)Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe\.Training language models to follow instructions with human feedback\.In*Advances in Neural Information Processing Systems*, volume 35, pp\. 27730–27744\. Curran Associates, Inc\., 2022\.doi:10\.52202/068431\-2011\.
- Qu et al\. \(2025\)Jingang Qu, David Holzmüller, Gaël Varoquaux, and Marine Le Morvan\.TabICL: A tabular foundation model for in\-context learning on large data\.In*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, pp\. 50817–50847\. PMLR, 2025\.
- Schulman et al\. \(2016\)John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel\.High\-dimensional continuous control using generalized advantage estimation\.In*International Conference on Learning Representations*, 2016\.
- Schulman et al\. \(2017\)John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov\.Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347, 2017\.URL[https://arxiv\.org/abs/1707\.06347](https://arxiv.org/abs/1707.06347)\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y\. K\. Li, Y\. Wu, and Daya Guo\.DeepSeekMath: Pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300, 2024\.URL[https://arxiv\.org/abs/2402\.03300](https://arxiv.org/abs/2402.03300)\.
- Shridhar et al\. \(2021\)Mohit Shridhar, Xingdi Yuan, Marc\-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht\.ALFWorld: Aligning text and embodied environments for interactive learning\.In*International Conference on Learning Representations*, 2021\.
- Wang et al\. \(2024\)Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui\.Math\-Shepherd: Verify and reinforce LLMs step\-by\-step without human annotations\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 9426–9439\. Association for Computational Linguistics, 2024\.doi:10\.18653/v1/2024\.acl\-long\.510\.
- Wang et al\. \(2025\)Zihan Wang, Kangrui Wang, Qineng Wang, Pingyue Zhang, Linjie Li, Zhengyuan Yang, Xing Jin, Kefan Yu, Minh Nhat Nguyen, Licheng Liu, Eli Gottlieb, Yiping Lu, Kyunghyun Cho, Jiajun Wu, Fei\-Fei Li, Lijuan Wang, Yejin Choi, and Manling Li\.RAGEN: Understanding self\-evolution in LLM agents via multi\-turn reinforcement learning\.arXiv preprint arXiv:2504\.20073, 2025\.URL[https://arxiv\.org/abs/2504\.20073](https://arxiv.org/abs/2504.20073)\.
- Yao et al\. \(2022\)Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan\.WebShop: Towards scalable real\-world web interaction with grounded language agents\.In*Advances in Neural Information Processing Systems*, volume 35, pp\. 20744–20757\. Curran Associates, Inc\., 2022\.doi:10\.52202/068431\-1508\.
- Yao et al\. \(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao\.ReAct: Synergizing reasoning and acting in language models\.In*International Conference on Learning Representations*, 2023\.
- Yu et al\. \(2025\)Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Ru Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, Wei\-Ying Ma, Ya\-Qin Zhang, Lin Yan, Yonghui Wu, and Mingxuan Wang\.DAPO: An open\-source LLM reinforcement learning system at scale\.In*Advances in Neural Information Processing Systems*, volume 38, pp\. 113222–113244\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-3775\.
- Zhang et al\. \(2026a\)Yi\-Kai Zhang, Yueqing Sun, Hongyan Hao, Qi Gu, Xunliang Cai, Long Chen, De\-Chuan Zhan, and Han\-Jia Ye\.V0\.5V\_\{0\.5\}: Generalist value model as a prior for sparse RL rollouts\.In*Third Conference on Language Modeling*, 2026a\.
- Zhang et al\. \(2026b\)Yi\-Kai Zhang, Zhiyuan Yao, Hongyan Hao, Yueqing Sun, Qi Gu, Hui Su, Xunliang Cai, De\-Chuan Zhan, and Han\-Jia Ye\.V0V\_\{0\}: A generalist value model for any policy at state zero\.In*Proceedings of the 43rd International Conference on Machine Learning*, 2026b\.
- Zhou et al\. \(2024\)Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, and Aviral Kumar\.ArCHer: Training language model agents via hierarchical multi\-turn RL\.In*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 62178–62209\. PMLR, 2024\.
## Appendix AT2SPO Implementation and Evaluation Details
This appendix details policy training, evaluation, context construction, and optimizer integration\.
### A\.1 Training and Evaluation Protocol
Table[3](https://arxiv.org/html/2610.00388#A1.T3)summarizes the main T2SPO\-GRPO training settings for the Qwen2\.5\-1\.5B\-Instruct and Qwen2\.5\-7B\-Instruct models\. Each update collects eight trajectories for each of 16 task instances, giving 128 rollouts\. The policy receives the instruction, current observation, available actions, and up to two preceding observation–action pairs\. The estimator’s history window is configured separately, as detailed in Appendix[A\.2](https://arxiv.org/html/2610.00388#A1.SS2)\.
Table 3:Policy training and evaluation settings shared by both model sizes\.SettingALFWorldWebShopTask instances per update1616Rollouts per task instance88Actor learning rate10−610^\{\-6\}10−610^\{\-6\}KL loss coefficient0\.010\.01Invalid\-action penalty0\.100\.10Policy history window2 turns2 turnsMaximum prompt length \(tokens\)2,0484,096Maximum response length \(tokens\)512512Rollout horizon \(turns\)5015Training updates150150Evaluation interval \(updates\)55Episodes per evaluation128128Evaluation temperature0\.40\.4ALFWorld uses the environment’s training and in\-distribution evaluation splits\. WebShop reserves the first 500 instruction indices for evaluation and draws training tasks from the remaining indices\. Each WebShop evaluation samples 128 instructions without replacement from the reserved set\. Evaluation uses stochastic decoding at temperature 0\.4, and reported results use the checkpoint after update 150\. For both the success\-only and failure\-augmented T2SPO\-GRPO results, we report the mean and sample standard deviation over seeds 0, 1, and 2\. Baseline results in Table[1](https://arxiv.org/html/2610.00388#S5.T1)are taken from[Feng et al\. \(2025\)](https://arxiv.org/html/2610.00388#bib.bib4)\.
### A\.2 Estimator and Context Configurations
Table[4](https://arxiv.org/html/2610.00388#A1.T4)specifies the success\-only T2SPO\-GRPO configurations in Table[1](https://arxiv.org/html/2610.00388#S5.T1)\. All use a frozen Qwen3\-Embedding\-0\.6B encoder and the official TabPFN V2 regressor with one ensemble member and a random seed matched to the training run\. Both ALFWorld models use a structured state summary without an action suffix; WebShop uses raw observations\. The estimator includes the full available interaction history for both benchmarks\. Step and collection\-round features are added before context\-fitted standardization and SVD\.
Table 4:Estimator settings for the main T2SPO\-GRPO runs at both model sizes\. History refers to the estimator input; context and successful\-state pool budgets count state rows\.Shared settingsValueEstimator history windowFull prefixTabPFN package version8\.3\.0Shared context budget256Successful\-state pool capacity256New successful states per round≤64\\leq 64Recent trajectory window256Minimum support rows8Process clipcc5Final process weight0\.10Burn\-in / linear ramp \(rounds\)1 / 4Discountγ\\gamma1\.0Benchmark\-specific settingsALFWorldWebShopState textStructured summaryRawSVD dimension6432Suppress unchanged observationsYesNoContext initialization and refresh\.Each benchmark starts with a fixed archive of 128 training trajectories collected from a separate policy checkpoint at update 100\. ALFWorld uses a separate archive for each model size, while WebShop shares one frozen archive across both sizes\. Each archive is shared across training seeds\. Successful states from this archive fill the context while the online pool is small\. Each completed training batch contributes at most 64 sampled successful states to a pool of capacity 256, with older entries evicted at capacity\. Context selection takes online states first and fills any remaining positions from the archive, so archived examples are gradually replaced as online experience accumulates\. Both sources use each state’s own successful continuation length from Equation[4](https://arxiv.org/html/2610.00388#S4.E4)\. The next batch shares this context and its fitted feature transformation\.
Failure\-augmented context\.The variants in Table[2](https://arxiv.org/html/2610.00388#S5.T2)keep the context budget at 256 and the process weight at 0\.10 on both benchmarks\. Unsuccessful states receive the horizon labelH=50H=50on ALFWorld andH=15H=15on WebShop, and occupy at most 25% of the context; successful states fill the remaining positions\. Each variant retains the corresponding success\-only run’s policy training and state representation settings\.
### A\.3 Failure\-Augmented Results by Task Category
Table[5](https://arxiv.org/html/2610.00388#A1.T5)gives the complete ALFWorld category results for the success\-only and failure\-augmented configurations\. At 7B, all six category means increase with failure examples; at 1\.5B, the changes vary across categories\. The overall ALFWorld and WebShop comparisons are reported in Table[2](https://arxiv.org/html/2610.00388#S5.T2)\.
Table 5:Failure\-augmented context across ALFWorld task categories\.Success rates \(%\), averaged over three seeds; subscripts denote sample standard deviations\. T2SPO\-GRPO uses success\-only context, while T2SPO\-GRPO \+ Fail admits up to 25% failure examples\. Bold marks the higher mean at each size\.
### A\.4 Algorithm and Reproducibility Details
Update order\.For each training round, the policy first collects a rollout batch\. The estimator predicts remaining distances for the batch using only the existing context, after which the step credits and policy\-training targets are computed\. The completed trajectories can then be inserted into memory and the next predictor context prepared\. This insertion does not change the predictions or targets already computed for the current batch\. The actor is optimized using these fixed targets; in the PPO extension the critic is updated as well\. There are no gradient updates to the encoder or TabPFN\.
GRPO aggregation\.The base multi\-turn implementation computes within\-task outcome statistics over expanded turn rows\. Consequently, a trajectory with more turns contributes more entries to the group mean and standard deviation\. This convention differs from the standard estimator in Equation[1](https://arxiv.org/html/2610.00388#S3.E1), which contributes one score per trajectory\. T2SPO preserves the base convention and adds its process term only after the base advantages have been computed\.
Validity checks\.The direct predictor requires sufficient support, finite input features, a valid projection, and finite predictions\. Insufficient support disables the process contribution; rejected predictions receive zero gate in numerical fallback paths\. Strict error handling can instead stop an update on an estimator error\. For the ALFWorld configuration, unchanged observations set both the process gate and normalization mask to zero\. If no turns remain active, the process signal is zero\.
### A\.5 Actor–Critic Extension
We detail the PPO integration from Section[4\.4](https://arxiv.org/html/2610.00388#S4.SS4); Appendix[D](https://arxiv.org/html/2610.00388#A4)reports its ALFWorld results\.
Turn probabilities\.A complete responseat=\(at,1,…,at,Lt\)a\_\{t\}=\(a\_\{t,1\},\\ldots,a\_\{t,L\_\{t\}\}\)has autoregressive probability
πθ\(at∣st\)=∏k=1Ltπθ\(at,k∣st,at,<k\)\.\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)=\\prod\_\{k=1\}^\{L\_\{t\}\}\\pi\_\{\\theta\}\(a\_\{t,k\}\\mid s\_\{t\},a\_\{t,<k\}\)\.\(11\)Environment observations condition the policy but are excluded from its action likelihood and loss\.
Advantage estimation\.PPO uses a learned criticVψ\(st\)V\_\{\\psi\}\(s\_\{t\}\)to approximate the expected discounted return fromsts\_\{t\}\. Generalized advantage estimation \(GAE\) forms discounted sums of temporal\-difference residuals\([Schulman et al\., 2016](https://arxiv.org/html/2610.00388#bib.bib14)\); for a rollout segment ofKKturns\([Li et al\., 2026](https://arxiv.org/html/2610.00388#bib.bib8)\),
δt\\displaystyle\\delta\_\{t\}=rt\+γVψ\(st\+1\)−Vψ\(st\),\\displaystyle=r\_\{t\}\+\\gamma V\_\{\\psi\}\(s\_\{t\+1\}\)\-V\_\{\\psi\}\(s\_\{t\}\),\(12\)A^tGAE\\displaystyle\\hat\{A\}\_\{t\}^\{\\mathrm\{GAE\}\}=∑ℓ=0K−t−1\(γλ\)ℓδt\+ℓ,\\displaystyle=\\sum\_\{\\ell=0\}^\{K\-t\-1\}\(\\gamma\\lambda\)^\{\\ell\}\\delta\_\{t\+\\ell\},whereγ∈\(0,1\]\\gamma\\in\(0,1\]is the discount factor andλ∈\[0,1\]\\lambda\\in\[0,1\]controls the bias–variance tradeoff\. At a true terminal state, the successor value is zero; for a rollout truncated before termination, the final residual instead uses a bootstrapped successor value\. For T2SPO, the augmented rewardsr\+r^\{\+\}from Equation[10](https://arxiv.org/html/2610.00388#S4.E10)replacertr\_\{t\}in these residuals to obtainA^\+\\hat\{A\}^\{\+\}\. The critic targets areV^i,ttarget=A^i,t\+\+Vψ\(si,t\)\\hat\{V\}^\{\\mathrm\{target\}\}\_\{i,t\}=\\hat\{A\}^\{\+\}\_\{i,t\}\+V\_\{\\psi\}\(s\_\{i,t\}\), formed before advantage whitening\. At truncation, the PFN also evaluates the actual successor state when constructing the auxiliary reward\.
Policy updates\.Following turn\-level optimization\([Li et al\., 2026](https://arxiv.org/html/2610.00388#bib.bib8)\), define the likelihood ratio of the complete generated response and its clipped objective as
ρi,t\(θ\)\\displaystyle\\rho\_\{i,t\}\(\\theta\)=πθ\(ai,t∣si,t\)πold\(ai,t∣si,t\),\\displaystyle=\\frac\{\\pi\_\{\\theta\}\(a\_\{i,t\}\\mid s\_\{i,t\}\)\}\{\\pi\_\{\\mathrm\{old\}\}\(a\_\{i,t\}\\mid s\_\{i,t\}\)\},\(13\)𝒥turn\(θ\)\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{turn\}\}\(\\theta\)=1Ntok∑i,tℓclip\(ρi,t\(θ\),A¯i,t\+\)\.\\displaystyle=\\frac\{1\}\{N\_\{\\mathrm\{tok\}\}\}\\sum\_\{i,t\}\\ell\_\{\\mathrm\{clip\}\}\\bigl\(\\rho\_\{i,t\}\(\\theta\),\\bar\{A\}^\{\+\}\_\{i,t\}\\bigr\)\.HereA¯\+\\bar\{A\}^\{\+\}is whitened over valid turns, andNtokN\_\{\\mathrm\{tok\}\}is the batch’s total valid generated\-token count\. Each turn contributes one clipped surrogate\. KL regularization is retained from the base optimizer\. The complete\-turn log\-ratio sums generated\-token log\-ratios, clipped to\[−20,20\]\[\-20,20\]before exponentiation for numerical stability\.
## Appendix BOffline Distance Prediction
Table[6](https://arxiv.org/html/2610.00388#A2.T6)reports the WebShop context/SVD grid from Section[5\.3](https://arxiv.org/html/2610.00388#S5.SS3)\. The query set contains 514 trajectories and 3,073 states, with regression metrics computed on its 1,951 successful\-trajectory states\. Context and query targets use each successful trajectory’s own remaining length\. Tasks are split between context and query sets with seed 20260812\. A query\-independent ordering of 2,115 successful context states \(seed 20260920\) defines nested pools of 128, 256, and 512\.
All configurations encode raw observations with recent two\-turn history using Qwen3\-Embedding\-0\.6B and append the step and collection\-round features\. Standardization and SVD are fitted on the selected context and then applied to the query states\. The regressor uses the official TabPFN V2 checkpoint, one ensemble member, and random seed 0\. The same cached 1,024\-dimensional embeddings and held\-out query states are reused for every context budget and SVD dimension\.
Table 6:Offline remaining\-distance prediction on WebShop\.All rows use TabPFN and evaluate the same 1,951 successful query states on held\-out tasks\. MAE and RMSE are measured in turns;ρ\\rhodenotes Spearman correlation\. Bold numbers mark the best value for each metric\.### B\.1 Nearest\-Neighbor Comparison
Table[7](https://arxiv.org/html/2610.00388#A2.T7)reports the full kNN comparison underlying Figure[2](https://arxiv.org/html/2610.00388#S5.F2)b\. At each context budget, both methods use the same context examples, SVD32 feature transformation, and 1,951 successful query states\. The kNN estimates are unweighted means of labels selected by cosine similarity in the transformed feature space\. The figure uses the lowest observed kNN MAE among five tested neighbor counts at each budget:k=32k=32,6464, and88for context sizes 128, 256, and 512, respectively\.
Table 7:TabPFN and kNN on WebShop\.All results use SVD32; columns give context budgetsMM\. MAE is in turns, with the lowest value in bold\. Atk=Mk=M, kNN returns the context mean\.
## Appendix CComputation Cost
We measure online computation in the three success\-only WebShop Qwen2\.5\-1\.5B\-Instruct runs, each using two NVIDIA H20 GPUs, a context budget of 256, SVD32, and a 15\-turn horizon\. Each optimizer step collects 128 trajectories and queries 640–1,760 states, with a median of 928\. We retain log records from the final training process of each run and exclude steps containing evaluation or checkpoint saving, leaving 120 steps per seed and 360 steps in total\. Table[8](https://arxiv.org/html/2610.00388#A3.T8)reports descriptive timing statistics over these steps\. Each percentage is the sum of the component’s times divided by the sum of full\-step times\.
Table 8:Online training\-step timing on WebShop\.Means and standard deviations are over 360 optimizer steps from three seeds; times are in seconds\. Shares use summed full\-step time\. The lower block expands the advantage row in the upper block\.Timing hierarchy\.The six top\-level stage timers are mutually exclusive and together cover 99\.966% of recorded step time; the residual consists of between\-stage work and timer overhead\. State estimation and online\-context preparation are nested within advantage computation\. The state\-estimation timer covers feature and embedding handling, prediction\-context construction, SVD, TabPFN fitting, and prediction\. Preparing context updates for subsequent batches is timed separately\. Together, state estimation and context preparation account for 21\.138% of step time\. The remaining advantage time includes base GRPO computation, shaping, validation, and context commit\. This runtime has no separate weight\-synchronization timer, and GRPO uses no critic\. These timings describe component shares within the measured runs, rather than a matched\-control slowdown\.
## Appendix DAdditional Turn\-PPO Results
We evaluate the actor–critic integration in Appendix[A\.5](https://arxiv.org/html/2610.00388#A1.SS5)on ALFWorld with both model sizes\. Our Turn\-PPO control and T2SPO\-Turn\-PPO retain the same learned critic, GAE, and turn\-level policy objective; T2SPO adds auxiliary rewards with weight 0\.10\. Each update collects one rollout for each of 128 task instances, with a horizon of 50 turns and 150 training updates\. Both use actor and critic learning rates of10−610^\{\-6\}and10−510^\{\-5\}, respectively, discountγ=0\.99\\gamma=0\.99, GAE parameterλ=0\.9\\lambda=0\.9, one PPO epoch, and clipping ratio 0\.2\. The estimator uses raw state text with recent two\-turn history and full admissible\-action descriptions, SVD64, a context of 256 successful state examples with trajectory\-localT−tT\-ttargets, and the official TabPFN V2 regressor from package version 2\.2\.1\.
Table 9:Turn\-PPO with and without T2SPO on ALFWorld\.Success rates \(%\) are means with sample standard deviations over the same two training seeds, 0 and 1 \(n=2n=2\)\. Each final checkpoint is evaluated on 128 episodes after 150 updates;Δ\\Deltais the change in mean success in percentage points\.Table[9](https://arxiv.org/html/2610.00388#A4.T9)shows unchanged mean success at 1\.5B and a 1\.2\-point increase at 7B\. At each scale, one seed improves and the other declines\.相似文章
StepPO:面向智能体强化学习的步骤对齐策略优化
StepPO 引入了一种面向智能体强化学习的步骤中心范式,该范式将策略优化与智能体决策粒度对齐,在多轮交互任务中优于以令牌为中心的方法。
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Introduces SSPO, a step-level self-distilled policy optimization method for training deep search agents, which uses evidence anchors and advantage weights to improve credit assignment beyond sparse outcome rewards. SSPO outperforms GRPO on benchmarks like BrowseComp and GAIA with only ~5% overhead per step.
SAPO:用于智能体强化学习的单次展开自回归策略优化
SAPO 是一种用于智能体强化学习的低内存、高计算效率框架,它在单个自回归骨干中共享策略和价值函数,在 ALFWorld 和 WebShop 的实验中表现优于 PPO 和 GRPO。
PGPO:面向多轮智能体任务的势引导策略优化
PGPO 提出势引导策略优化用于多轮智能体任务,使得在 LLM 后训练中实现更细粒度的信用分配,并在 ALFWorld 和 WebShop 基准测试上展示出强劲结果。
PlanPO:面向多轮代理式大语言模型的群体规划感知策略优化
PlanPO是一种强化学习方法,它为多轮代理式大语言模型引入了从粗到细的优势信号,在ALFWorld、WebShop和SciWorld等基准测试中,性能比GRPO提高了27.2%。