面向智能体强化学习信用分配的关键决策定位
摘要
提出 ProVer 框架,通过 agentic judge 筛选可能的关键决策片段,并利用前向与后续采样的终端成功率来验证优势值,从而实现 agentic RL 中的细粒度信用分配。在 ALFWorld、WebShop 和 SearchQA 上,较 GRPO 最高提升 9.91%。
arXiv:2609.36178v1 Announce Type: new
Abstract: Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
查看缓存全文
缓存时间: 2026/09/30 09:50
# Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
Source: [https://arxiv.org/html/2609.36178](https://arxiv.org/html/2609.36178)
Dongwon Jung Hemanth Neelgund Ramesh Yifan Wang Xiaomin Li††thanks:Work done during an internship at Microsoft\.Andrzej Banburski\-Fahey Jaron LanierAffiliation:University of California, Davis Microsoft University of Washington Purdue University
###### Abstract
Group Relative Policy Optimization \(GRPO\) has become a promising approach for training large language model agents\. However, its uniform assignment of trajectory\-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success\. We introduceProVer, a framework that targets potentially pivotal decisions for fine\-grained credit assignment in agentic reinforcement learning\. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes\. Rather than directly trusting the judge’s assessment,ProVerverifies the proposed segment by estimating its advantage from the difference in terminal success rates between current\-policy continuations sampled before and after the segment\. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment\. By using model judgment only to select where to verify,ProVergrounds local credit in observed outcomes without exhaustively evaluating every intermediate state\. Across ALFWorld, WebShop, and SearchQA,ProVerachieves the strongest average performance at both model scales, with relative improvements over GRPO of 9\.91% and 7\.12% for Qwen3\.5\-2B and Qwen3\.5\-4B, respectively\. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier\-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine\-grained credit assignment in agentic reinforcement learning\.
## 1Introduction
Large language model \(LLM\) agents have become increasingly competent at solving complex tasks through multi\-turn interaction with external tools and environments, including web navigation\([Zhou et al\., 2024a](https://arxiv.org/html/2609.36178#bib.bib4);[Wei et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib7)\), information retrieval\([Chen et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib10);[Jin et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib27)\), and tool use\([Huang et al\., 2024](https://arxiv.org/html/2609.36178#bib.bib6);[Qian et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib11);[Feng et al\., 2026a](https://arxiv.org/html/2609.36178#bib.bib5)\)\. Recent work has therefore trained LLM agents directly through environment interaction using reinforcement learning with verifiable outcomes\([Wei et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib7);[Jin et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib27);[Feng et al\., 2026a](https://arxiv.org/html/2609.36178#bib.bib5)\)\. In particular, Group Relative Policy Optimization \(GRPO\) has become a widely used foundation for such training because it estimates relative advantages from multiple trajectories sampled for the same task without requiring a separately trained value model\([Shao et al\., 2024](https://arxiv.org/html/2609.36178#bib.bib2);[Guo et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib1)\)\. However, when applied to long\-horizon agents, its trajectory\-level supervision creates a fundamental challenge for credit assignment: terminal feedback distinguishes trajectories within a group but provides little information about which intermediate decisions were consequential to the outcome\. Thus, the same trajectory\-level advantage is assigned to all policy tokens within a trajectory, regardless of their contributions, potentially reinforcing mistakes in successful trajectories and penalizing useful progress in failed ones\.
This limitation has motivated approaches that provide finer\-grained, intermediate supervision to distinguish the contributions of intermediate decisions rather than uniformly propagating the trajectory\-level advantages\. Prior work obtains intermediate supervision through learned value functions or critic models that estimate expected returns\([Bahdanau et al\., 2017](https://arxiv.org/html/2609.36178#bib.bib41);[Zhou et al\., 2024b](https://arxiv.org/html/2609.36178#bib.bib8)\), or through process reward models and LLM evaluators that directly assess intermediate behavior\([Lightman et al\., 2024](https://arxiv.org/html/2609.36178#bib.bib13);[Zhang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib14);[Wang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib15)\)\. Although these approaches provide finer\-grained supervision, their model\-based evaluations may be inaccurate, unreliable as the policy distribution shifts, or susceptible to exploitation during optimization\([Liu et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib18);[Cui et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib20);[Hou et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib16)\)\. Monte Carlo policy evaluation instead estimates intermediate\-state values by averaging terminal returns from fresh continuations sampled from the current policy, grounding supervision in observed outcomes rather than model\-predicted values or evaluator judgments\([Sutton, 1988](https://arxiv.org/html/2609.36178#bib.bib22);[Sutton and Barto, 2018](https://arxiv.org/html/2609.36178#bib.bib23)\)\. In long\-horizon reasoning and agentic settings, this can be implemented by branching from intermediate states and sampling multiple continuations\([Kazemnejad et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib12);[Guo et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib19);[Ji et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib17)\)\. However, extending this approach to long\-horizon agentic tasks remains costly since evaluating every candidate state may require repeatedly restoring environment states and executing full continuations\. This motivates selectively identifying potentially consequential decisions and concentrating fine\-grained credit assignment on them, rather than exhaustively evaluating every intermediate state\.
To address this challenge, we introduceProVer\(Propose,Verify, Credit\), a framework for targeting potentially pivotal decisions for fine\-grained credit assignment in agentic reinforcement learning\. The key idea is to use model judgment only to identify a candidate segment, then concentrate outcome verification on its boundaries to estimate its advantage from current\-policy continuations\.ProVercomprises three stages: propose, verify, and credit\. In the propose stage, an LLM\-based agentic judge contrasts successful and failed trajectories from the same GRPO group to localize a potentially pivotal segment in a successful trajectory\. In the verify stage, we restore the environment states immediately before and after the proposed segment and sample fresh continuations from the current policy\. The difference in empirical terminal success rates provides an estimate of the segment’s advantage under the current policy\. In the credit stage, positive estimates are added to the original GRPO advantages of policy tokens within the proposed segment\. By separating LLM\-based proposal from Monte Carlo verification,ProVergrounds additional credit in observed outcomes rather than potentially inaccurate critic scores, while restricting evaluation to the proposed segment’s boundaries\.
We evaluateProVeron ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2609.36178#bib.bib24)\), WebShop\([Yao et al\., 2022](https://arxiv.org/html/2609.36178#bib.bib25)\), and SearchQA using Qwen3\.5\-2B and Qwen3\.5\-4B\.ProVerachieves the strongest average performance at both model scales, yielding relative improvements over GRPO of 9\.91% and 7\.12%, respectively\. Further analyses demonstrate that informed segment selection improves policy learning with modest additional generation overhead, even without a frontier\-scale judge model\. Together, these results support selective outcome verification as an effective and efficient approach to credit assignment in agentic reinforcement learning\.
## 2Related Work
### 2\.1Agentic Reinforcement Learning
Agentic reinforcement learning casts interaction with an environment as a sequential decision problem and optimizes policies over complete environment trajectories\. This paradigm spans web navigation, information seeking, and tool use\([Wei et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib7);[Jin et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib27);[Chen et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib10);[Qian et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib11)\)\. A common foundation is Group Relative Policy Optimization \(GRPO\), which assigns a trajectory\-level outcome advantage to all policy\-generated tokens without training a value model\([Shao et al\., 2024](https://arxiv.org/html/2609.36178#bib.bib2);[Wang et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib3)\)\. GiGPO\([Feng et al\., 2026b](https://arxiv.org/html/2609.36178#bib.bib9)\)refines this supervision through anchor\-state grouping, comparing discounted downstream returns of actions taken from repeated environment states\. This provides step\-level credit without auxiliary models or additional rollouts, although its coverage depends on state recurrence within the sampled group\.ProVerinstead estimates local value changes through fresh continuations from judge\-selected boundary states\.
### 2\.2Critic\-Based Process Supervision for Credit Assignment
LLM critics provide intermediate supervision by assessing reasoning steps or agent actions\. Some methods use natural\-language critiques to identify errors, guide revised attempts, and incorporate critique\-guided improvements into the policy\([Zhang et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib37);[Lin et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib36)\)\. Beyond guiding revisions, LLM critics can directly supply intermediate credit\. CriticSearch uses a frozen retrospective critic to evaluate search actions using completed trajectories and reference answers, converting its judgments into turn\-level feedback\([Zhang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib14)\)\. Related methods assign semantic roles to actions and translate them into local credit corrections\([Xu et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib38)\), or weight retrieval rounds according to their judged contributions\([Wang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib15)\)\. In these credit\-assignment methods, evaluator judgments directly shape local credit\. InProVer, the agentic judge selects a candidate segment, while fresh environment continuations determine its estimated advantage and the resulting additional credit\.
### 2\.3Outcome\-Based Credit Assignment
Outcome\-based credit assignment estimates the contribution of intermediate decisions from downstream task outcomes\. Monte Carlo value estimation provides one way to obtain this signal by averaging returns from continuations sampled at intermediate states\. Prior work organizes outcome\-labeled continuations as implicit prefix trees or explicitly sampled branching structures\([Hou et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib16);[Ji et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib17);[Zhao et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib39)\)\. VinePPO samples continuations at intermediate reasoning\-step boundaries to estimate step\-level advantages\([Kazemnejad et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib12)\), while SPO partitions trajectories into segments and estimates segment\-level advantages through chain\- or tree\-based sampling\([Guo et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib19)\)\.ProVershares their use of continuation\-based value estimation, but uses an agentic judge to select one potentially consequential segment through cross\-trajectory comparison, focusing evaluation on its two boundaries rather than evaluating successive steps or segments throughout a trajectory\.
## 3Method
We consider reinforcement learning for LLM agents that receive verifiable terminal rewards from an environment\. We first introduce the agent–environment setting and GRPO formulation in Section[3\.1](https://arxiv.org/html/2609.36178#S3.SS1), followed by an overview ofProVerin Section[3\.2](https://arxiv.org/html/2609.36178#S3.SS2)\. We then describe its three stages in Sections[3\.3](https://arxiv.org/html/2609.36178#S3.SS3)–[3\.5](https://arxiv.org/html/2609.36178#S3.SS5)\.
### 3\.1Preliminaries
Agent–environment interaction\.Letxxdenote a task instance andπθ\\pi\_\{\\theta\}an LLM agent policy\. At turntt, the complete agent–environment statests\_\{t\}comprises the interaction history available to the policy and the corresponding environment state\. The policy samples a responseat∼πθ\(⋅∣st\)a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\}\), which induces the next statest\+1s\_\{t\+1\}\. Repeating this interaction produces a trajectoryτ=\(s1,a1,…,sT,aT,sT\+1\)\\tau=\(s\_\{1\},a\_\{1\},\\ldots,s\_\{T\},a\_\{T\},s\_\{T\+1\}\)\. After the final turn, the environment returns a terminal rewardR\(τ\)R\(\\tau\)\. We focus on tasks with binary outcomes, whereR\(τ\)∈\{0,1\}R\(\\tau\)\\in\\\{0,1\\\}indicates failure or success, and assume no intermediate rewards\.
Agentic reinforcement learning with GRPO\.For each taskxx, GRPO samples a group ofGGtrajectories𝒢=\{τi\}i=1G\\mathcal\{G\}=\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{G\}from the same policy\. LetRi=R\(τi\)R\_\{i\}=R\(\\tau\_\{i\}\)\. With binary rewards, we use the mean\-centered group\-relative advantage:
AiGRPO=Ri−1G∑j=1GRj\.A\_\{i\}^\{\\mathrm\{GRPO\}\}=R\_\{i\}\-\\frac\{1\}\{G\}\\sum\_\{j=1\}^\{G\}R\_\{j\}\.\(1\)GRPO assignsAiGRPOA\_\{i\}^\{\\mathrm\{GRPO\}\}to every trainable policy token inτi\\tau\_\{i\}, while environment observations and other non\-policy tokens are masked from the optimization objective\. This distinguishes successful from failed trajectories within the group, but it does not distinguish among the decisions that compose each trajectory\.
We denote the successful and failed trajectories by𝒢\+=\{τi∈𝒢∣Ri=1\}\\mathcal\{G\}^\{\+\}=\\\{\\tau\_\{i\}\\in\\mathcal\{G\}\\mid R\_\{i\}=1\\\}and𝒢−=\{τi∈𝒢∣Ri=0\}\\mathcal\{G\}^\{\-\}=\\\{\\tau\_\{i\}\\in\\mathcal\{G\}\\mid R\_\{i\}=0\\\}, respectively\. We call𝒢\\mathcal\{G\}a*mixed\-outcome group*when both subsets are nonempty\.
Figure 1:Overview ofProVer\.\(a\) Propose:An agentic judge contrasts successful and failed trajectories to localize a potentially consequential segment\.\(b\) Verify:Boundary continuations estimate the segment advantageΔ\\Deltafrom the difference in terminal success rates\.\(c\) Credit:Positive estimates augment GRPO advantages only for policy tokens within the selected segment\.
### 3\.2ProVerOverview
Figure[1](https://arxiv.org/html/2609.36178#S3.F1)illustrates the overview ofProVer\. Given a mixed\-outcome GRPO group,ProVeraims to localize potentially pivotal decisions within a successful trajectory and assign them fine\-grained credit without exhaustively evaluating every intermediate state\. We operationalize such decisions as short contiguous segments of agent turns and estimate their contribution from downstream outcomes\.
ProVerproceeds in three stages\. In the propose stage, an LLM agent, called the agentic judge, contrasts successful and failed trajectories to select a potentially consequential segment from a successful trajectory\. In the verify stage,ProVerrestores the agent–environment states immediately before and after the proposed segment and samples fresh current\-policy continuations from both states\. The difference between the two empirical terminal success rates estimates the segment’s advantage\. In the credit stage, when this estimate is positive,ProVeradds it to the original GRPO advantage assigned to each trainable policy token in the selected segment\. This design combines model\-based guidance\([Zhang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib14);[Wang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib15)\)with outcome\-grounded Monte Carlo estimation\([Kazemnejad et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib12);[Guo et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib19)\): the judge identifies a promising segment without directly supplying credit scores, while fresh current\-policy continuations estimate its advantage without requiring exhaustive evaluation of intermediate states\.
### 3\.3Propose: Agentic Judge\-Guided Segment Selection
The propose stage localizes a potentially consequential decision or sequence of decisions within a successful trajectory\. We represent the candidate as a contiguous segment of agent turns, which is subsequently evaluated to determine whether it warrants additional credit\. Selecting such a segment requires understanding how its actions address difficulties encountered elsewhere in the rollout group, rather than assessing each action in isolation\. Because the relevant evidence may be scattered across long trajectories, we use anagentic judgethat can search for recurring failure patterns and selectively inspect the turns needed to assess a candidate segment\. This allows segment selection to draw on cross\-trajectory context without requiring every trajectory to be processed in full\.
Agentic Judge Design\.The judge is implemented as an external LLM agent and receives the task, a complete turn\-indexed successful trajectoryτ\+∈𝒢\+\\tau^\{\+\}\\in\\mathcal\{G\}^\{\+\}, and compact previews of the failed trajectories in𝒢−\\mathcal\{G\}^\{\-\}, each truncated to a fixed length\. It uses the failed trajectories to identify recurrent difficulties and examines how a segment in the successful trajectory avoids or resolves them\. To support this inspection, we provide two tools adapted from[Lee et al\. \(2026\)](https://arxiv.org/html/2609.36178#bib.bib21):
- •search\_trajectory\(query, k\): Searches all failed trajectories using the natural\-languagequeryand returns the top\-kkbounded turn previews with their trajectory identifiers and turn indices\.
- •get\_segment\(traj\_id, start\_turn, end\_turn\): Retrieves the inclusive turn range\[start\_turn,end\_turn\]\[\\texttt\{start\\\_turn\},\\texttt\{end\\\_turn\}\]from the trajectory identified bytraj\_id, including the policy\-visible context before the first retrieved turn\.
The judge can alternate between searching for supporting evidence and inspecting relevant turns to refine its hypothesis about which segment warrants evaluation\. After inspection, it outputs inclusive turn boundaries\(ℓ,r\)\(\\ell,r\)inτ\+\\tau^\{\+\}, defining the proposed segmentτseg=\(sℓ,aℓ,…,sr,ar,sr\+1\)\\tau^\{\\mathrm\{seg\}\}=\(s\_\{\\ell\},a\_\{\\ell\},\\ldots,s\_\{r\},a\_\{r\},s\_\{r\+1\}\)\. The selected segment therefore remains a hypothesis about which decisions contributed to success; the verify stage tests whether it warrants additional credit\.
### 3\.4Verify: Outcome\-Based Segment Evaluation
The propose stage identifies a potentially consequential segment, but it does not determine whether the segment deserves additional credit\. The verify stage makes this distinction explicit: the judge’s proposal only allocates the evaluation budget, while the segment’s advantage is estimated from downstream task outcomes\. Following the Monte Carlo principle of estimating intermediate values from sampled continuations, we compare fresh current\-policy continuations from the states immediately before and after the proposed segment\. The difference between the two empirical terminal success rates provides an outcome\-based estimate of whether, and by how much, executing the segment improves the policy’s probability of success\. We first formalize the segment advantage that this comparison seeks to estimate\.
Segment advantage\.The turn boundaries\(ℓ,r\)\(\\ell,r\)determine the pre\- and post\-segment statesspre=sℓs\_\{\\mathrm\{pre\}\}=s\_\{\\ell\}andspost=sr\+1s\_\{\\mathrm\{post\}\}=s\_\{r\+1\}\. Let𝐚ℓ:r=\(aℓ,…,ar\)\\mathbf\{a\}\_\{\\ell:r\}=\(a\_\{\\ell\},\\ldots,a\_\{r\}\)denote the recorded action sequence withinτseg\\tau^\{\\mathrm\{seg\}\}\. Letπ\\pidenote the policy that generated the source group, held fixed during verification\. We writeVπ\(s\)V^\{\\pi\}\(s\)for the expected discounted return when followingπ\\pifrom statess, andQsegπ\(spre,𝐚ℓ:r\)Q\_\{\\mathrm\{seg\}\}^\{\\pi\}\(s\_\{\\mathrm\{pre\}\},\\mathbf\{a\}\_\{\\ell:r\}\)for the expected discounted return when executing the recorded segment and then followingπ\\pi\.
Treating the recorded segment as a temporally extended action and assuming deterministic segment execution, its current\-policy advantage is given below\. The final equality uses our setting of nonterminal segments with no intermediate rewards andγ=1\\gamma=1:
Asegπ\(spre,𝐚ℓ:r\)\\displaystyle A\_\{\\mathrm\{seg\}\}^\{\\pi\}\(s\_\{\\mathrm\{pre\}\},\\mathbf\{a\}\_\{\\ell:r\}\)=Qsegπ\(spre,𝐚ℓ:r\)−Vπ\(spre\)\\displaystyle=Q\_\{\\mathrm\{seg\}\}^\{\\pi\}\(s\_\{\\mathrm\{pre\}\},\\mathbf\{a\}\_\{\\ell:r\}\)\-V^\{\\pi\}\(s\_\{\\mathrm\{pre\}\}\)\(2\)=∑j=ℓrγj−ℓrj\+γdVπ\(spost\)−Vπ\(spre\)\\displaystyle=\\sum\_\{j=\\ell\}^\{r\}\\gamma^\{j\-\\ell\}r\_\{j\}\+\\gamma^\{d\}V^\{\\pi\}\(s\_\{\\mathrm\{post\}\}\)\-V^\{\\pi\}\(s\_\{\\mathrm\{pre\}\}\)=Vπ\(spost\)−Vπ\(spre\),\\displaystyle=V^\{\\pi\}\(s\_\{\\mathrm\{post\}\}\)\-V^\{\\pi\}\(s\_\{\\mathrm\{pre\}\}\),whererjr\_\{j\}is the reward received at turnjj,d=r−ℓ\+1d=r\-\\ell\+1is the segment duration in turns, andγ\\gammais the per\-turn discount factor\. With binary terminal rewards andγ=1\\gamma=1, the boundary values represent probabilities of eventual success\. Thus, verifying a proposal reduces to estimating the current\-policy values at these two boundaries\.
Boundary rollouts\.We estimate the two boundary values through Monte Carlo policy evaluation\. For each valid proposal, we restorespres\_\{\\mathrm\{pre\}\}andsposts\_\{\\mathrm\{post\}\}, including the interaction history and corresponding environment state\. From each restored state, we independently sampleKKfresh continuations using the fixed policyπ\\pi\. LetRkpreR\_\{k\}^\{\\mathrm\{pre\}\}andRkpostR\_\{k\}^\{\\mathrm\{post\}\}denote their terminal rewards\. Because these rewards are binary, their empirical means estimate the policy’s success probabilities at the two boundaries:
V^preπ=1K∑k=1KRkpre,V^postπ=1K∑k=1KRkpost,Δ^seg=V^postπ−V^preπ,\\widehat\{V\}\_\{\\mathrm\{pre\}\}^\{\\pi\}=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}R\_\{k\}^\{\\mathrm\{pre\}\},\\qquad\\widehat\{V\}\_\{\\mathrm\{post\}\}^\{\\pi\}=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}R\_\{k\}^\{\\mathrm\{post\}\},\\qquad\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}=\\widehat\{V\}\_\{\\mathrm\{post\}\}^\{\\pi\}\-\\widehat\{V\}\_\{\\mathrm\{pre\}\}^\{\\pi\},\(3\)whereΔ^seg\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}is the resulting segment\-advantage estimate\.
Under the stated assumptions,Δ^seg\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}is conditionally unbiased for the segment advantage in Equation[2](https://arxiv.org/html/2609.36178#S3.E2), given the selected segment and its boundary states\.111Appendix[B](https://arxiv.org/html/2609.36178#A2)proves this result and clarifies its scope\.A positiveΔ^seg\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}provides outcome\-based evidence that the segment improves the policy’s probability of success\.
### 3\.5Credit: Segment\-Level Advantage Assignment
Because the boundary\-value comparison estimates the advantage of the segment as a whole, we assign the same verified bonus to all of its trainable policy tokens\. Letuuindex trainable policy tokens inτ\+\\tau^\{\+\}, and letIuseg=𝟙\[ℓ≤turn\(u\)≤r\]I\_\{u\}^\{\\mathrm\{seg\}\}=\\mathbbm\{1\}\[\\ell\\leq\\operatorname\{turn\}\(u\)\\leq r\]indicate whether tokenuubelongs to an action turn within the proposed segment\. For a valid segment\-advantage estimate, we define
AuProVer=\{AGRPO\+λΔ^segIuseg,ifΔ^seg\>0,AGRPO,otherwise,A\_\{u\}^\{\\scriptstyle\\textsc\{\\mbox\{\{ProVer\}\} \}\}=\\begin\{cases\}A^\{\\mathrm\{GRPO\}\}\+\\lambda\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}I\_\{u\}^\{\\mathrm\{seg\}\},&\\text\{if \}\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\>0,\\\\ A^\{\\mathrm\{GRPO\}\},&\\text\{otherwise\},\\end\{cases\}\(4\)whereAGRPOA^\{\\mathrm\{GRPO\}\}is the trajectory\-level advantage ofτ\+\\tau^\{\+\}andλ≥0\\lambda\\geq 0controls the strength of the segment\-level bonus\. Invalid proposals or measurements receive no additional credit\. We substituteAuProVerA\_\{u\}^\{\\scriptstyle\\textsc\{\\mbox\{\{ProVer\}\} \}\}for the ordinary GRPO advantage onτ\+\\tau^\{\+\}while leaving the source trajectories, policy ratios, clipping rule, and other optimization terms unchanged\. All other trajectories retain their original GRPO advantages\. Boundary continuations are used only for advantage estimation and are not included in the policy\-training batch\.
Overall, the judge selects potentially pivotal segments for evaluation, while boundary\-rollout outcomes determine whether and how much additional credit they receive\. This separation enablesProVerto concentrate outcome\-based evaluation on a small part of the trajectory and augment GRPO’s trajectory\-level supervision with selectively verified local credit\.
## 4Experimental Setup
### 4\.1Agent Environments
We evaluateProVeron three long\-horizon agent environments: \(1\)ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2609.36178#bib.bib24)\): a text\-based household simulator in which the agent observes its surroundings and executes natural\-language actions to manipulate objects and complete a task; \(2\)WebShop\([Yao et al\., 2022](https://arxiv.org/html/2609.36178#bib.bib25)\): a simulated e\-commerce website in which the agent issues searches and navigates product pages through web actions before selecting a product that satisfies a natural\-language goal; and \(3\)SearchQA: a corpus\-backed retrieval environment in which the agent iteratively submits search queries, examines retrieved evidence, and returns an answer to a single\- or multi\-hop question\. Further details on the environments, interaction protocols, and prompt templates are provided in Appendix[A\.1](https://arxiv.org/html/2609.36178#A1.SS1)\.
### 4\.2Baselines
We compareProVerwith six baselines: \(1\)GRPO\([Shao et al\., 2024](https://arxiv.org/html/2609.36178#bib.bib2)\), the standard group\-relative policy optimization baseline; \(2\)Budget\-Matched GRPO, which augments GRPO with additional policy rollouts to match the rollout budget used byProVerfor segment\-boundary value estimation; \(3\)GiGPO\([Feng et al\., 2026b](https://arxiv.org/html/2609.36178#bib.bib9)\), which augments GRPO with step\-level advantages derived from discounted returns at repeated environment states; \(4\)SPO\-chain\([Guo et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib19)\), which assigns segment\-level credit using differences between Monte Carlo value estimates at segment boundaries; \(5\)SPO\-tree\([Guo et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib19)\), which uses tree\-structured sampling to share trajectory prefixes and estimate sibling\-relative advantages; \(6\)CriticSearch\([Zhang et al\., 2026](https://arxiv.org/html/2609.36178#bib.bib14)\), which augments GRPO with binary action\-level scores from a retrospective LLM critic\. We use the same underlying LLM for this critic and our agentic judge\.
These baselines examine complementary aspects of credit assignment\. Budget\-Matched GRPO tests whether additional policy rollouts alone account for the gains\. GiGPO provides a comparison with step\-level credit inferred from repeatedly visited states\. SPO\-chain and SPO\-tree compare targeted segment selection with fixed segment boundaries and tree\-structured sampling, respectively\. CriticSearch compares credit derived from sampled downstream outcomes with credit derived directly from an LLM critic’s action\-level judgments\. Complete training and method\-specific implementation details are provided in Appendix[A\.3](https://arxiv.org/html/2609.36178#A1.SS3)\.
### 4\.3Implementation Details
We use Prime\-RL\([Prime Intellect, 2025](https://arxiv.org/html/2609.36178#bib.bib40)\)to train Qwen3\.5\-2B and Qwen3\.5\-4B for 100 optimizer steps\. The standard GRPO configuration uses 16 groups ofG=8G=8trajectories per update\. By default,ProVerusesgpt\-5\.4\-minias its agentic judge, matching the underlying LLM used by the CriticSearch critic\. We applyProVerto groups with success rates strictly between 0 and0\.50\.5, selecting one successful trajectory and proposing one nonterminal segment of at most four action turns per eligible group\. Verification usesK=8K=8continuations per boundary, and positive segment advantages are incorporated with weightλ=1\\lambda=1\. We evaluate task success rate on ALFWorld and WebShop and semantic answer accuracy on SearchQA, usinggpt\-5\.4\-minito judge answer correctness\. Additional training and segment\-selection details are provided in Appendix[A\.2](https://arxiv.org/html/2609.36178#A1.SS2)\.
## 5Experimental Results
Table 1:Performance of the baselines andProVerwith Qwen3\.5\-2B and Qwen3\.5\-4B on three benchmarks\. We report the mean and sample standard deviation over three evaluation runs\.### 5\.1Main Results
ProVerdelivers the strongest overall performance\.Table[1](https://arxiv.org/html/2609.36178#S5.T1)shows thatProVerachieves the highest average score at both model scales and the best result in four of the six model–environment settings\. Relative to GRPO,ProVerimproves average performance by 9\.91% with Qwen3\.5\-2B and 7\.12% with Qwen3\.5\-4B, with gains in all six settings\. These results demonstrate the effectiveness of augmenting trajectory\-level GRPO with fine\-grained credit targeted toward potentially consequential decisions\.
Additional group\-level rollouts alone do not explain the gains\.Budget\-Matched GRPO matches the number of additional policy rollouts used byProVerbut retains trajectory\-level credit\.ProVeroutperforms this control in every setting, with relative improvements ranging from 5\.34% to 22\.23%\. Moreover, Budget\-Matched GRPO does not consistently outperform standard GRPO\. These results support concentrating the additional rollouts on evaluating potentially consequential segments, rather than simply collecting more trajectories for GRPO\.
Targeting consequential decisions outperforms structured rollout baselines\.ProVeroutperforms SPO\-tree across all six settings and SPO\-chain in five of the six settings\. The gains over SPO\-chain are particularly pronounced on SearchQA, reaching 43\.95% with Qwen3\.5\-2B and 17\.74% with Qwen3\.5\-4B\. While SPO\-chain evaluates fixed segment boundaries and SPO\-tree uses a predetermined branching structure,ProVerevaluates the boundaries of a segment selected by the agentic judge\. These comparisons support the effectiveness of using judge\-guided localization to concentrate continuation\-based credit estimation on potentially consequential decisions\.
Judge\-guided verification outperforms direct critic\-based credit\.CriticSearch assigns action\-level credit directly from binary labels produced by a retrospective LLM critic, whereasProVeruses an agentic judge to select a segment and estimates its advantage from downstream outcomes\. Both use the same underlying LLM for the critic or judge\.ProVerachieves a higher average score at both model scales and outperforms CriticSearch in five of the six settings\. These results support the proposed separation between segment selection and outcome\-based credit estimation, compared with directly translating critic judgments into credit\.
### 5\.2Training Cost Analysis
Having established the performance gains ofProVer, we next examine the additional computation required to obtain them\. Table[2](https://arxiv.org/html/2609.36178#S5.T2)reports generated policy tokens, judge API cost, and optimizer\-step wall time for Qwen3\.5\-4B training\.
Selective evaluation of potentially consequential decisions limits additional policy generation\.ProVerselects one short segment from one successful trajectory per eligible group and evaluates only its two boundary states\. SPO\-chain instead evaluates multiple boundaries for every eligible successful trajectory, while CriticSearch scores every scorable turn in all eight trajectories\. Relative to GRPO,ProVerincreases generated policy tokens per step by 2\.4% on ALFWorld, 11\.6% on WebShop, and 16\.8% on SearchQA\. Together with the strongest average performance in Table[1](https://arxiv.org/html/2609.36178#S5.T1), these results show that selectively evaluating one potentially consequential segment provides fine\-grained credit with modest additional policy\-generation overhead\.
ProVerachieves the lowest per\-step wall time among the evaluated fine\-grained methods\.Table[2](https://arxiv.org/html/2609.36178#S5.T2)shows thatProVerhas lower per\-step wall time than SPO\-tree, SPO\-chain, and CriticSearch across all three environments\. This advantage is not explained by generation volume alone: SPO\-tree generates fewer policy tokens than GRPO on ALFWorld and WebShop but takes substantially longer per optimizer step, consistent with additional execution overhead from sequential branching and environment\-state restoration\. SPO\-chain incurs additional rollout cost by evaluating multiple segment boundaries in every eligible successful trajectory, whereasProVerevaluates only two boundaries of one selected segment\. Similarly, CriticSearch scores every scorable action across all eight trajectories in an eligible group, whereas the agentic judge inProVerselectively inspects trajectory evidence to propose one segment, with 38\.4% lower judge API cost on average across the three environments\. Together, these comparisons show that concentrating evaluation on a single proposed segment can provide segment\-level credit with lower per\-step wall time than the evaluated fine\-grained alternatives\.
Table 2:Per\-step training cost for Qwen3\.5\-4B on ALFWorld \(ALF\), WebShop \(WS\), and SearchQA \(SQA\)\. Generated tokens count policy\-rollout outputs\. Judge cost is calculated from recordedgpt\-5\.4\-miniAPI usage using OpenAI’s published pricing as of August 2026\. Time is the mean wall\-clock time per optimizer step in minutes\.
### 5\.3Effect of Judge Model Choice
A natural question is whetherProVerrequires an expensive frontier\-scale model as its external judge\. We compare three models from the GPT\-5\.4 family and a locally served Qwen3\.5\-9B judge while holding the Qwen3\.5\-4B policy backbone, training procedure, and outcome\-verification mechanism fixed\. We also include aRandombaseline that replaces the judge with random sampling of a valid nonterminal segment from the same selected successful trajectory, while retaining the verification budget and credit\-assignment rule\.222The random sampling procedure is detailed in Appendix[A\.3](https://arxiv.org/html/2609.36178#A1.SS3)\.Table[3](https://arxiv.org/html/2609.36178#S5.T3)reports policy performance across the three environments, together with proposal acceptance rate \(Δ^seg\>0\\hat\{\\Delta\}\_\{\\mathrm\{seg\}\}\>0\) and mean segment length on SearchQA\.
Informed segment selection concentrates evaluation on more promising candidate segments\.All four agentic judges outperform random selection across all three environments\. On SearchQA, random selection yields an acceptance rate of 19\.5%, compared with 58\.5–72\.4% for the agentic judges, despite proposing substantially longer segments on average \(2\.38 vs\. 1\.03–1\.45 turns\)\. This suggests that the judges more precisely localize potentially consequential segments, rather than obtaining positive estimates simply by covering more actions\. Together with the policy results, this supports the benefit of judge\-guided localization over allocating fine\-grained credit to arbitrarily selected segments\.
Within the GPT\-5\.4 family, the lower acceptance rates ofgpt\-5\.4\-miniandgpt\-5\.4accompany shorter proposals than those ofgpt\-5\.4\-nano\. A plausible explanation is that these models follow the minimal\-segment instruction more closely, which requires more precise localization, whereas longer segments span more environment interactions and observations and may be more likely to yield a positiveΔ^seg\\hat\{\\Delta\}\_\{\\mathrm\{seg\}\}\. However, lower acceptance does not necessarily imply less effective selection, asgpt\-5\.4\-miniproduces stronger downstream policies thangpt\-5\.4\-nanoacross all three environments\.
Effective segment proposal does not require a frontier\-scale judge\.Although informed selection matters, using a frontier\-scale model does not necessarily improve the resulting policy\. Bothgpt\-5\.4\-miniandgpt\-5\.4\-nanoimprove over standard GRPO across all three environments, andgpt\-5\.4\-miniachieves the highest scores on WebShop and SearchQA\. The locally served Qwen3\.5\-9B judge achieves the highest reported ALFWorld score and exceeds GRPO on SearchQA, although it falls below GRPO on WebShop\. In contrast,gpt\-5\.4does not achieve the highest score on any environment\. Together, these findings support the practical design ofProVer: smaller or locally served models can identify useful candidate segments, while environment rollouts supply the numerical credit signal\. Judge choice remains task\-dependent, but frontier\-scale capability is unnecessary for strong performance in the evaluated settings\.
Table 3:Effect of segment selection and judge model choice on Qwen3\.5\-4B policy training\.Randomreplaces judge\-guided selection with random segment sampling while retaining the same outcome\-verification and credit\-assignment procedure\. We additionally report proposal diagnostics on SearchQA: acceptance rate \(Δ^seg\>0\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\>0\) and mean segment length in agent turns\.
## 6Conclusion
We presentProVer, a framework for targeting potentially pivotal decisions for fine\-grained credit assignment in agentic reinforcement learning\. An agentic judge localizes a candidate segment, while outcome verification estimates its advantage from fresh current\-policy continuations\. Incorporating positive segment\-advantage estimates into GRPO yields the strongest average performance among the evaluated baselines at both model scales across three agent environments, with modest additional generation overhead\. Our analyses show that informed segment selection contributes to these gains without requiring a frontier\-scale judge\. Together, these findings support a practical role for LLM judgment in agentic reinforcement learning: identifying where fine\-grained credit is most useful, while leaving the numerical credit signal to downstream environment outcomes\.
## AI use statement
We used generative AI tools to assist with experiment implementation and debugging, quantitative result analysis, and manuscript revision\. In particular, these tools helped implement and test training and evaluation infrastructure, inspect experimental artifacts, and improve the clarity and structure of the paper\. All AI\-assisted code was reviewed and tested against the intended behavior, and all reported results and numerical claims were checked against the underlying experiment artifacts or independently recomputed by the authors\. We did not use generative AI to create the benchmark datasets or ground\-truth task outcomes\. The authors reviewed all AI\-assisted work and take responsibility for the final text, claims, code, and artifacts presented in this paper\.
## References
- Bahdanauet al\.\(2017\)D\. Bahdanau, P\. Brakel, K\. Xu, A\. Goyal, R\. Lowe, J\. Pineau, A\. Courville, and Y\. BengioAn actor\-critic algorithm for sequence prediction\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SJDaqqveg)Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
- Chenet al\.\(2025\)M\. Chen, L\. Sun, T\. Li, H\. Sun, Y\. Zhou, C\. Zhu, H\. Wang, J\. Z\. Pan, W\. Zhang, H\. Chen,et al\.Learning to reason with search for llms via reinforcement learning\.arXiv preprint arXiv:2503\.19470\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1)\.
- Cuiet al\.\(2026\)G\. Cui, L\. Yuan, Z\. Wang, H\. Wang, Y\. Zhang, J\. Chen, W\. Li, B\. He, Y\. Fan, T\. Yu, Q\. Xu, W\. Chen, J\. Yuan, H\. Chen, K\. Zhang, X\. Lv, S\. Wang, Y\. Yao, X\. Han, H\. Peng, Y\. Cheng, Z\. Liu, M\. Sun, B\. Zhou, and N\. DingProcess reinforcement through implicit rewards\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=9SkkifLopZ)Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
- Douzeet al\.\(2024\)M\. Douze, A\. Guzhva, C\. Deng, J\. Johnson, G\. Szilvasy, P\. Mazaré, M\. Lomeli, L\. Hosseini, and H\. JégouThe faiss library\.External Links:2401\.08281Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p1.1)\.
- Fenget al\.\(2026a\)J\. Feng, S\. Huang, X\. Qu, G\. Zhang, Y\. Qin, B\. Zhong, C\. Jiang, J\. Chi, and W\. ZhongRetool: reinforcement learning for strategic tool use in llms\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 37909–37926\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1)\.
- Fenget al\.\(2026b\)L\. Feng, Z\. Xue, T\. Liu, and B\. AnGroup\-in\-group policy optimization for llm agent training\.Advances in Neural Information Processing Systems38,pp\. 46375–46408\.Cited by:[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.36178#S4.SS2.p1.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1)\.
- Guoet al\.\(2026\)Y\. Guo, L\. Xu, J\. Liu, D\. Ye, and S\. QiuSegment policy optimization: effective segment\-level credit assignment in rl for large language models\.Advances in Neural Information Processing Systems38,pp\. 114399–114431\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.36178#S2.SS3.p1.1),[§3\.2](https://arxiv.org/html/2609.36178#S3.SS2.p2.1),[§4\.2](https://arxiv.org/html/2609.36178#S4.SS2.p1.1)\.
- Hoet al\.\(2020\)X\. Ho, A\. D\. Nguyen, S\. Sugawara, and A\. AizawaConstructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps\.InProceedings of the 28th International Conference on Computational Linguistics,pp\. 6609–6625\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Houet al\.\(2025\)Z\. Hou, Z\. Hu, Y\. Li, R\. Lu, J\. Tang, and Y\. DongTreerl: llm reinforcement learning with on\-policy tree search\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12355–12369\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.36178#S2.SS3.p1.1)\.
- Huanget al\.\(2024\)T\. Huang, D\. Jung, V\. Kumar, M\. Kachuee, X\. Li, P\. Xu, and M\. ChenPlanning and editing what you retrieve for enhanced tool learning\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 975–988\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1)\.
- Jiet al\.\(2026\)Y\. Ji, Z\. Ma, Y\. Wang, G\. Chen, X\. Chu, and L\. WuTree search for llm agent reinforcement learning\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 87362–87388\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.36178#S2.SS3.p1.1)\.
- Jinet al\.\(2025\)B\. Jin, H\. Zeng, Z\. Yue, J\. Yoon, S\. O\. Arik, D\. Wang, H\. Zamani, and J\. HanSearch\-r1: training LLMs to reason and leverage search engines with reinforcement learning\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Rwhi91ideu)Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p1.1),[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1),[§1](https://arxiv.org/html/2609.36178#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1)\.
- Joshiet al\.\(2017\)M\. Joshi, E\. Choi, D\. S\. Weld, and L\. ZettlemoyerTriviaQA: a large scale distantly supervised challenge dataset for reading comprehension\.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1601–1611\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Kazemnejadet al\.\(2025\)A\. Kazemnejad, M\. Aghajohari, E\. Portelance, A\. Sordoni, S\. Reddy, A\. Courville, and N\. L\. RouxVinePPO: refining credit assignment in RL training of LLMs\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=Myx2kJFzAn)Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.36178#S2.SS3.p1.1),[§3\.2](https://arxiv.org/html/2609.36178#S3.SS2.p2.1)\.
- Kwiatkowskiet al\.\(2019\)T\. Kwiatkowski, J\. Palomaki, O\. Redfield, M\. Collins, A\. Parikh, C\. Alberti, D\. Epstein, I\. Polosukhin, J\. Devlin, K\. Lee,et al\.Natural questions: a benchmark for question answering research\.Transactions of the Association for Computational Linguistics7,pp\. 452–466\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Leeet al\.\(2026\)Y\. Lee, H\. Yen, X\. Ye, and D\. ChenAgentic aggregation for parallel scaling of long\-horizon agentic tasks\.InSecond Workshop on Agents in the Wild: Safety, Security, and Beyond,External Links:[Link](https://openreview.net/forum?id=hXzAocijH7)Cited by:[§3\.3](https://arxiv.org/html/2609.36178#S3.SS3.p2.1)\.
- Lightmanet al\.\(2024\)H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. CobbeLet’s verify step by step\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 39578–39601\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
- Linet al\.\(2026\)J\. Lin, X\. Yu, Y\. Xin, Y\. Guo, Z\. Jiang, Z\. Yue, W\. Wang, H\. Zou, C\. Qin, and H\. XiongICRL: learning to internalize self\-critique with reinforcement learning\.arXiv preprint arXiv:2605\.15224\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1),[§2\.2](https://arxiv.org/html/2609.36178#S2.SS2.p1.1)\.
- Liuet al\.\(2026\)H\. Liu, D\. Yu, S\. Lu, Y\. Zhou, R\. Liu, Z\. Liang, H\. Mi, C\. Wei, and D\. YuSave the good prefix: precise error penalization via process\-supervised rl to enhance llm reasoning\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 35450–35477\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
- Mallenet al\.\(2023\)A\. Mallen, A\. Asai, V\. Zhong, R\. Das, D\. Khashabi, and H\. HajishirziWhen not to trust language models: investigating effectiveness of parametric and non\-parametric memories\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 9802–9822\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Presset al\.\(2023\)O\. Press, M\. Zhang, S\. Min, L\. Schmidt, N\. A\. Smith, and M\. LewisMeasuring and narrowing the compositionality gap in language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 5687–5711\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Prime Intellect \(2025\)Prime IntellectPrime\-rl\.External Links:[Link](https://github.com/PrimeIntellect-ai/prime-rl)Cited by:[§4\.3](https://arxiv.org/html/2609.36178#S4.SS3.p1.1)\.
- Qianet al\.\(2026\)C\. Qian, E\. C\. Acikgoz, Q\. He, H\. Wang, X\. Chen, D\. Hakkani\-Tur, G\. Tur, and H\. JiToolrl: reward is all tool learning needs\.Advances in Neural Information Processing Systems38,pp\. 105523–105553\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.36178#S4.SS2.p1.1)\.
- Shridharet al\.\(2021\)M\. Shridhar, X\. Yuan, M\. Cote, Y\. Bisk, A\. Trischler, and M\. Hausknecht\{ALFW\}orld: aligning text and embodied environments for interactive learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0IOX0YcCdTn)Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.36178#S1.p4.1),[§4\.1](https://arxiv.org/html/2609.36178#S4.SS1.p1.1)\.
- Sutton and Barto \(2018\)R\. S\. Sutton and A\. G\. BartoReinforcement learning: an introduction\.2 edition,MIT Press\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
- Sutton \(1988\)R\. S\. SuttonLearning to predict by the methods of temporal differences\.Machine learning3\(1\),pp\. 9–44\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
- Trivediet al\.\(2022\)H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. SabharwalMuSiQue: multi\-hop questions via single\-hop question composition\.Transactions of the Association for Computational Linguistics10,pp\. 539–554\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Wanget al\.\(2026\)J\. Wang, Z\. Xi, Y\. Yang, H\. Luo, S\. Dou, T\. Gui, and Q\. ZhangEnhancing llm\-based search agents via contribution weighted group relative policy optimization\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 31704–31718\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.36178#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.36178#S3.SS2.p2.1)\.
- Wanget al\.\(2025\)Z\. Wang, K\. Wang, Q\. Wang, P\. Zhang, L\. Li, Z\. Yang, X\. Jin, K\. Yu, M\. N\. Nguyen, L\. Liu,et al\.Ragen: understanding self\-evolution in llm agents via multi\-turn reinforcement learning\.arXiv preprint arXiv:2504\.20073\.Cited by:[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1)\.
- Weiet al\.\(2025\)Z\. Wei, W\. Yao, Y\. Liu, W\. Zhang, Q\. Lu, L\. Qiu, C\. Yu, P\. Xu, C\. Zhang, B\. Yin,et al\.Webagent\-r1: training web agents via end\-to\-end multi\-turn reinforcement learning\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 7920–7939\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36178#S2.SS1.p1.1)\.
- Xiet al\.\(2025\)Z\. Xi, Y\. Ding, W\. Chen, B\. Hong, H\. Guo, J\. Wang, X\. Guo, D\. Yang, C\. Liao, W\. He,et al\.Agentgym: evaluating and training large language model\-based agents across diverse environments\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 27914–27961\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.p1.1)\.
- Xuet al\.\(2026\)Y\. Xu, Z\. Zhou, H\. Sang, X\. Li, J\. Zhang, X\. Du, S\. Na, Z\. Wang, and A\. GeramifardTRIAGE: role\-typed credit assignment for agentic reinforcement learning\.arXiv preprint arXiv:2606\.32017\.Cited by:[§2\.2](https://arxiv.org/html/2609.36178#S2.SS2.p1.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,pp\. 2369–2380\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px3.p2.1)\.
- Yaoet al\.\(2022\)S\. Yao, H\. Chen, J\. Yang, and K\. NarasimhanWebshop: towards scalable real\-world web interaction with grounded language agents\.Advances in Neural Information Processing Systems35,pp\. 20744–20757\.Cited by:[§A\.1](https://arxiv.org/html/2609.36178#A1.SS1.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.36178#S1.p4.1),[§4\.1](https://arxiv.org/html/2609.36178#S4.SS1.p1.1)\.
- Zhanget al\.\(2025\)X\. Zhang, Y\. Zhang, H\. Sun, K\. Feng, C\. Lu, C\. Yang, and H\. MengCritique\-grpo: advancing llm reasoning with natural language and numerical feedback\.arXiv preprint arXiv:2506\.03106\.Cited by:[§2\.2](https://arxiv.org/html/2609.36178#S2.SS2.p1.1)\.
- Zhanget al\.\(2026\)Y\. Zhang, H\. Huang, Z\. Song, Z\. Zhao, Q\. Zhang, Y\. Zhu, and D\. ZhaoCriticsearch: fine\-grained credit assignment for search agents via a retrospective critic\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 12272–12290\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.36178#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.36178#S3.SS2.p2.1),[§4\.2](https://arxiv.org/html/2609.36178#S4.SS2.p1.1)\.
- Zhaoet al\.\(2026\)Z\. Zhao, Z\. Ren, J\. Zou, L\. Yang, Z\. Xu, X\. Ge, Z\. Chen, X\. Ma, D\. Shi, S\. Wang,et al\.Reinforced efficient reasoning via semantically diverse exploration\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 47994–48007\.Cited by:[§2\.3](https://arxiv.org/html/2609.36178#S2.SS3.p1.1)\.
- Zhouet al\.\(2024a\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p1.1)\.
- Zhouet al\.\(2024b\)Y\. Zhou, A\. Zanette, J\. Pan, S\. Levine, and A\. KumarArcher: training language model agents via hierarchical multi\-turn rl\.arXiv preprint arXiv:2402\.19446\.Cited by:[§1](https://arxiv.org/html/2609.36178#S1.p2.1)\.
## Appendix AExperiment Details
### A\.1Agent Environment Details
We use the AgentGym implementations of ALFWorld, WebShop, and SearchQA\([Xi et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib26)\)\. The system and user prompts below are reproduced verbatim from our implementation; brace\-delimited fields indicate dynamic content in the environment\-observation templates\. The policy responds through native tool calling with exactly oneact\(action=\.\.\.\)call per turn\.
#### ALFWorld\.
ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2609.36178#bib.bib24)\)is a text\-based embodied environment for household tasks\. The initial observation contains the task description, current location, visible objects, and exact available actions\. Each subsequent observation reports the result of the previous action and an updated action set\. The agent must choose an exact available action string, which may navigate to a location or manipulate an object through operations such as taking, opening, placing, heating, cooling, or cleaning\. We use the text\-world version with a maximum of 50 turns\. We train on the canonical 2,620 training games and evaluate on all 134 valid\-unseen games\. The evaluation metric is task success accuracy\.
ALFWorld Prompt TemplateSYSTEMYou are controlling an agent in the ALFWorld household environment\.The initial task, current room, and exact AVAILABLE ACTIONS are provided before your first turn\.Call act with exactly one currently available action on every turn\.Do not invent actions\. Continue until the environment reports success or the turn budget ends\.USERComplete the ALFWorld task\.ENVIRONMENT OBSERVATION TEMPLATEOBSERVATION:\{observation\}AVAILABLE ACTIONS:\- \{action\_1\}\- \{action\_2\}\.\.\.DONE: \{done\}REWARD: \{reward\}
#### WebShop\.
WebShop\([Yao et al\., 2022](https://arxiv.org/html/2609.36178#bib.bib25)\)presents a natural\-language shopping instruction and a rendered e\-commerce page\. Each observation contains the visible page text and the currently available interaction affordances\. The agent may issuesearch\[query\]when a search bar is available or select an exact visible target withclick\[target\]\. Through these actions, it navigates search results and product pages, chooses product options, and completes a purchase\. We allow at most 15 turns\. We train on the 10,587 canonical training goals using binary task success as the reward\. We evaluate on all 500 canonical test goals and report exact binary task success metrics\.
WebShop Prompt TemplateSYSTEMYou are controlling an agent in the WebShop environment\.The shopping instruction, current page, and available search/click actions are provided before your first turn\.Call act exactly once on every turn\. Use search\[non\-empty query\] only when search is available, or click\[target\] with an exact current target\.Continue until a purchase ends the episode or the turn budget ends\.USERComplete the WebShop task\.ENVIRONMENT OBSERVATION TEMPLATEOBSERVATION:\{observation\}AVAILABLE ACTIONS:\- \{visible\_action\_1\}\- \{visible\_action\_2\}\.\.\.DONE: \{done\}SUCCESS: \{success\}SCORE: \{native\_reward\}
#### SearchQA\.
SearchQA presents a question and allows the agent to alternate between evidence retrieval and answer submission\. A<search\>query</search\>action returns retrieved passages inside the next observation, which the agent can use to formulate later queries\. A terminal<answer\>answer</answer\>action submits a concise final answer\. We utilize a FAISS index\([Douze et al\., 2024](https://arxiv.org/html/2609.36178#bib.bib35)\)over the 2018 Wikipedia corpus with E5\-base\-v2 query embeddings used by[Jin et al\. \(2025\)](https://arxiv.org/html/2609.36178#bib.bib27)\. Agents are allowed at most 30 turns to complete a task\.
The training pool contains 169,615 questions from the training splits of\([Jin et al\., 2025](https://arxiv.org/html/2609.36178#bib.bib27)\)consisting of NQ\([Kwiatkowski et al\., 2019](https://arxiv.org/html/2609.36178#bib.bib30)\)and HotpotQA\([Yang et al\., 2018](https://arxiv.org/html/2609.36178#bib.bib28)\)\. Evaluation uses a fixed stratified manifest of 400 held\-out questions used by[Lin et al\. \(2026\)](https://arxiv.org/html/2609.36178#bib.bib36): 50 each from NQ, TriviaQA\([Joshi et al\., 2017](https://arxiv.org/html/2609.36178#bib.bib29)\), PopQA\([Mallen et al\., 2023](https://arxiv.org/html/2609.36178#bib.bib31)\), HotpotQA, 2WikiMultiHopQA\([Ho et al\., 2020](https://arxiv.org/html/2609.36178#bib.bib32)\), and Bamboogle\([Press et al\., 2023](https://arxiv.org/html/2609.36178#bib.bib33)\), and 100 from MuSiQue\([Trivedi et al\., 2022](https://arxiv.org/html/2609.36178#bib.bib34)\)\. Training uses a binary normalized exact\-match reward\. For evaluation, we report semantic\-equivalence accuracy using GPT\-5\.4\-mini as an LLM judge to recognize correct answer variants that exact matching may miss\.
SearchQA Prompt TemplateSYSTEMYou are answering a question with the SearchQA environment\.The exact question is provided before your first turn\.Call act exactly once on every turn with exactly one argument named action\.The action value must be one complete string with both tags: use <search\>non\-empty query</search\> to retrieve evidence or <answer\>concise final answer</answer\> to finish\.Never omit </search\> or </answer\>, and never send query or answer as a separate argument\.Base the final answer on retrieved evidence and do not invent search results\.USERAnswer the SearchQA question\.ENVIRONMENT OBSERVATION TEMPLATEOBSERVATION:\{observation\}AVAILABLE ACTIONS:\- <search\>non\-empty query</search\>\- <answer\>concise final answer</answer\>DONE: \{done\}REWARD: \{reward\}
### A\.2Additional Implementation Details
Training configuration\.The standard GRPO configuration has a batch size of 128 trajectories, organized into 16 groups of eight trajectories\. We use AdamW with a learning rate of10−610^\{\-6\}and weight decay of 0 on ALFWorld and WebShop and 0\.1 on SearchQA\. Training uses a sampling temperature of 1 and a maximum completion length per turn of 512 tokens on ALFWorld and WebShop and 256 tokens on SearchQA\. Each run uses four A100 SXM GPUs for policy inference and four for training\. The baselines use the same policy backbones and shared optimization settings, with method\-specific changes to credit assignment and rollout allocation detailed in Appendix[A\.3](https://arxiv.org/html/2609.36178#A1.SS3)\.
Segment selection and verification\.ForG=8G=8, the eligibility criterion selects groups containing one, two, or three successful trajectories\. From each eligible group, we select the successful trajectory with the shortest horizon and ask the judge to identify a contiguous nonterminal segment of at most four action turns\. We restore the exact agent–environment states immediately before and after the segment and sampleK=8K=8current\-policy continuations from each boundary\. The difference between their mean terminal rewards estimates the segment advantage\. When this estimate is positive, we add it to the GRPO advantage of each trainable policy token within the segment, usingλ=1\\lambda=1in Equation[4](https://arxiv.org/html/2609.36178#S3.E4)\. All other source tokens retain their GRPO advantages, and ineligible groups retain standard GRPO updates\. Boundary continuations are used only for estimation and are not included in the policy\-training batch\.
### A\.3Baseline Implementation
#### GRPO\.
GRPO assigns Equation[1](https://arxiv.org/html/2609.36178#S3.E1)uniformly to all trainable assistant tokens in a trajectory\. It uses neither localized credit nor auxiliary policy rollouts and therefore provides the outcome\-only reference\.
#### Budget\-Matched GRPO\.
Budget\-Matched GRPO uses the same outcome\-only credit in Equation[1](https://arxiv.org/html/2609.36178#S3.E1)but replacesProVer’s boundary continuations with additional rollouts\. Each update contains 16 groups, and we use two phase\-specific group sizes to match the correspondingProVerrun’s realized rollout budget over 100 steps\. Table[4](https://arxiv.org/html/2609.36178#A1.T4)reports the number of trajectories sampled per group in each phase\. Thus, Budget\-Matched GRPO controls for additional policy sampling without using boundary restoration or localized credit\.
Table 4:Number of trajectories sampled per group in each phase of Budget\-Matched GRPO\. Every update contains 16 groups\.
#### GiGPO\.
GiGPO augments trajectory\-level GRPO advantages with step\-level advantages obtained by grouping action turns that share the same environment state and comparing their discounted downstream returns\. This provides local credit from state recurrence within the sampled rollout group, without auxiliary continuation rollouts or an external judge\. We use the same policy backbones and shared training configuration as the other baselines\.
#### SPO\-chain\.
SPO\-chain assigns segmentjjthe advantageAj=γVj−Vj−1A\_\{j\}=\\gamma V\_\{j\}\-V\_\{j\-1\}, whereVj−1V\_\{j\-1\}andVjV\_\{j\}are Monte Carlo values at the segment’s start and end boundaries andγ\\gammais the discount factor\. It masks tokens with generating\-policy probability at least0\.90\.9\. Applying the original method to every trajectory in every group is impractical for long, stateful agent trajectories\. We therefore apply it only to successful trajectories in groups containing one, two, or three successes followingProVer’s setup and divide their action turns into up to three contiguous, approximately equal segments; all other trajectories retain GRPO credit\. LetV0V\_\{0\}be the original group’s mean reward,V1V\_\{1\}andV2V\_\{2\}the internal\-boundary values, andV3=RiV\_\{3\}=R\_\{i\}the trajectory reward\. With no temporal discounting \(γ=1\\gamma=1\), the three advantages areA1=V1−V0A\_\{1\}=V\_\{1\}\-V\_\{0\},A2=V2−V1A\_\{2\}=V\_\{2\}\-V\_\{1\}, andA3=Ri−V2A\_\{3\}=R\_\{i\}\-V\_\{2\}\. Each internal value usesK=8K=8exact\-state continuations\. Fixed thirds keep the budget comparable toProVer, although SPO\-chain remains more expensive because it measures every successful trajectory rather than one selected success\.
#### SPO\-tree\.
SPO\-tree constructs a three\-level binary tree with branching factors22–22–22, giving 14 sampled branches and eight terminal leaves per training group\. For each childccof parentpp, it assigns the branch advantageA\(c\)=V\(c\)−V\(p\)A\(c\)=V\(c\)\-V\(p\)with no temporal discounting, where a leaf value is its binary terminal reward and each internal value is the mean reward of its descendant leaves\. The first two levels advance by four environment actions per level in ALFWorld and one action per level in WebShop and SearchQA, while the final level continues to the terminal condition or episode horizon\. Each child resumes from the parent’s exact environment state\. We pack the 14 branches into eight leaf containers and assign each branch to exactly one container, preventing duplicated credit for shared prefixes\. As in SPO\-chain, tokens with generating\-policy probability at least0\.90\.9are masked from the loss\. If tree collection fails, we discard the incomplete tree and sample a fresh group of eight standard GRPO trajectories\.
#### CriticSearch\.
CriticSearch uses the same frozengpt\-5\.4\-minias its retrospective critic\. Because labeling every trajectory in every group would require many costly critic calls, we apply CriticSearch only to groups containing one, two, or three successes, matchingProVer’s setup; all other groups use GRPO\. For each eligible group, the critic labels every trajectory, and CriticSearch uses
At=0\.75AiGRPO\+0\.25Atcritic\.A\_\{t\}=0\.75A\_\{i\}^\{\\mathrm\{GRPO\}\}\+0\.25A\_\{t\}^\{\\mathrm\{critic\}\}\.\(5\)Because CriticSearch labels all eight trajectories in each eligible group, it uses eight critic calls per group, whereasProVersends only one selected successful trajectory to its agentic judge\. SearchQA critics receive reference answers and label valid search actions\. ALFWorld and WebShop critics receive only policy\-visible trajectories and terminal success, never hidden environment state\. Invalid critic outputs or token alignment produce a whole\-group GRPO fallback\.
Random segment selection\.For the random\-selection baseline introduced in Section[5\.3](https://arxiv.org/html/2609.36178#S5.SS3), we replace the agentic judge with random selection of a starting turn and a segment length between one and four turns from the selected successful trajectory\. Intervals that violate the proposal\-validity constraints are discarded and receive no additional credit\. All other settings, including outcome verification and credit assignment, remain identical toProVer\.
### A\.4Agentic Judge Prompt
The agentic judge receives the complete turn\-indexed successful trajectory and compact indexes of the failed trajectories from the same GRPO group\. The prompt asks it to diagnose a recurrent failure mode and return the shortest non\-terminal expert segment that solves or avoids that failure\. The judge is given trajectory search and bounded segment\-inspection tools\. Its output contains only the inclusive segment bounds and a rationale\. We deterministically validate the output schema, turn bounds, action validity, and nonterminal endpoint, allowing up to three additional turns to correct any validation errors\.
Agentic Judge Prompt TemplateSYSTEMYou are an agentic judge identifying a short segment of a verified successful EXPERT trajectory that will receive additional fine\-grained positive credit on top of the GRPO terminal reward\.Diagnose one recurrent failure mode from the failed trajectories, then select the shortest EXPERT segment that solves or avoids that diagnosed failure mode\.You receive the complete EXPERT trajectory directly in the user prompt and compact indexes for all failed trajectories from the same task\. Each failed index gives its stable traj\_id and inclusive start\_turn/end\_turn bounds for get\_segment\.Diagnose a recurrent failure mode:The search\_trajectory and get\_segment tools are optional and available if useful\. Identify a failure mode that recurs in multiple failed trajectories\. Treat different actions as equivalent when they reach equally useful states, and do not group failures that require different fixes\.Select the EXPERT segment that solves the failure mode:From the complete EXPERT, find the shortest contiguous segment whose actions solve or avoid the diagnosed failure mode\. A failure that already performs the proposed EXPERT transition is not valid contrastive evidence merely because it fails elsewhere\.The policy tokens in every selected turn will themselves receive additional positive credit\. Select only actions whose own behavior should be reinforced\. Judge each action together with the transition it causes, not merely by whether a later state is useful\. A malformed, unavailable, unsuccessful, or error\-producing action must not be selected merely because recovering from its error later helped the EXPERT succeed\. A valid corrective action after an error may be selected when that corrective action itself is useful, but the preceding bad action must be excluded\.Segment\-selection rules:1\. Select the shortest contiguous segment that directly solves or avoids the diagnosed failure mode\. Every selected turn must be necessary to that transition and worth reinforcing; exclude merely contextual or adjacent turns\.2\. Select behavior that materially advances the task, not behavior that is merely late in a successful trajectory\. Do not select continued retrieval or repeated activity when the prior state is already sufficient unless it produces a necessary new state change\.3\. Treat actions as equivalent when they reach equally useful states, even if their routes, queries, object identities, or wording differ\.4\. The state after segment\_end\_turn must be nonterminal, and the terminal action must not be selected\.You do not write a hint, propose new actions, or assign reward\. You only select an existing EXPERT segment\. Use only recorded observations and affordances visible to the rollout policy\. Deterministic validation separately checks the output schema, bounds, action grammar and affordances, recorded action outcomes, terminality, and sampled tokens; your responsibility is the semantic and causal judgment\.If you use search previews, treat them as bounded evidence rather than complete local behavior; use get\_segment when more context would help\. Failed trajectories are contrastive evidence, not demonstrations\.Return only one JSON object with exactly these fields:\{"segment\_start\_turn": <inclusive integer\>,"segment\_end\_turn": <inclusive integer strictly before the EXPERT’s final recorded turn, whose resulting state is nonterminal\>,"rationale": "<the recurrent contrast, visible state before and after the selected segment, why each selected action itself deserves additional positive credit, and why this is the minimal complete causal transition\>"\}Do not include markdown or any text outside the JSON object\.USER\{"task": \{task\},"expert": \{complete\_turn\_indexed\_expert\_trajectory\},"failed\_trajectory\_indexes": \{compact\_failed\_trajectory\_indexes\},"instructions": "The expert is provided in full turn order above\. Select one expert segment satisfying the system prompt\. The trajectory tools are optional\."\}TOOLSsearch\_trajectory\(query, k=10\): Globally rank bounded turn previews across all failed trajectories\.get\_segment\(traj\_id, start\_turn, end\_turn\): Inspect a bounded inclusive turn interval from the expert or a failed trajectory\.
## Appendix BConditional Unbiasedness of the Segment Estimator
This section proves the conditional unbiasedness claim in Section[3\.4](https://arxiv.org/html/2609.36178#S3.SS4)and characterizes the finite\-sample variance of the raw pre/post estimator\. The result does not apply to the positive part\[Δ^seg\]\+\[\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\]\_\{\+\}or to the complete judge\-selected training update\.
#### Conditional unbiasedness\.
Letℱ\\mathcal\{F\}contain the source rollout group and all information used to selectτseg\\tau^\{\\mathrm\{seg\}\}\. Under exact boundary reconstruction, a fixed policyπ\\pi, and conditionally independent continuation samples, the selected segment and its boundary states are fixed after conditioning onℱ\\mathcal\{F\}\. Each empirical boundary value is an unbiased sample mean:
𝔼\[V^preπ∣ℱ\]=Vπ\(spre\),𝔼\[V^postπ∣ℱ\]=Vπ\(spost\)\.\\mathbb\{E\}\\\!\\left\[\\widehat\{V\}\_\{\\mathrm\{pre\}\}^\{\\pi\}\\mid\\mathcal\{F\}\\right\]=V^\{\\pi\}\(s\_\{\\mathrm\{pre\}\}\),\\qquad\\mathbb\{E\}\\\!\\left\[\\widehat\{V\}\_\{\\mathrm\{post\}\}^\{\\pi\}\\mid\\mathcal\{F\}\\right\]=V^\{\\pi\}\(s\_\{\\mathrm\{post\}\}\)\.Linearity of expectation and Equation[2](https://arxiv.org/html/2609.36178#S3.E2)therefore give
𝔼\[Δ^seg∣ℱ\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\\mid\\mathcal\{F\}\\right\]=Vπ\(spost\)−Vπ\(spre\)\\displaystyle=V^\{\\pi\}\(s\_\{\\mathrm\{post\}\}\)\-V^\{\\pi\}\(s\_\{\\mathrm\{pre\}\}\)=Asegπ\(spre,𝐚ℓ:r\)\.\\displaystyle=A\_\{\\mathrm\{seg\}\}^\{\\pi\}\(s\_\{\\mathrm\{pre\}\},\\mathbf\{a\}\_\{\\ell:r\}\)\.Thus, the rawΔ^seg\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}is conditionally unbiased for the selected segment’s advantage\.□\\square
#### Finite\-sample variance\.
For binary terminal outcomes, letppre=Vπ\(spre\)p\_\{\\mathrm\{pre\}\}=V^\{\\pi\}\(s\_\{\\mathrm\{pre\}\}\)andppost=Vπ\(spost\)p\_\{\\mathrm\{post\}\}=V^\{\\pi\}\(s\_\{\\mathrm\{post\}\}\)\. The two sample means have conditional variancesppre\(1−ppre\)/Kp\_\{\\mathrm\{pre\}\}\(1\-p\_\{\\mathrm\{pre\}\}\)/Kandppost\(1−ppost\)/Kp\_\{\\mathrm\{post\}\}\(1\-p\_\{\\mathrm\{post\}\}\)/K\. Their conditional independence therefore gives
Var\(Δ^seg∣ℱ\)=ppost\(1−ppost\)\+ppre\(1−ppre\)K\.\\operatorname\{Var\}\(\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\\mid\\mathcal\{F\}\)=\\frac\{p\_\{\\mathrm\{post\}\}\(1\-p\_\{\\mathrm\{post\}\}\)\+p\_\{\\mathrm\{pre\}\}\(1\-p\_\{\\mathrm\{pre\}\}\)\}\{K\}\.Thus, the estimator becomes more precise asKKincreases, with variance decreasing at rate1/K1/K\. For finiteKK, however, sampling noise can change the sign ofΔ^seg\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}, so a positive estimate provides empirical evidence of benefit rather than a conclusive guarantee\.
#### Scope of the unbiasedness guarantee\.
The conditional\-unbiasedness guarantee applies only to the raw estimator for a fixed segment after conditioning on the judge’s selection\.ProVerinstead uses its positive part,
\[Δ^seg\]\+=max\(Δ^seg,0\),\[\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\]\_\{\+\}=\\max\(\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\},0\),which is generally biased because negative sampling errors are replaced by zero while positive sampling errors are retained\. Consequently,
𝔼\[\[Δ^seg\]\+∣ℱ\]≠\[Asegπ\(spre,𝐚ℓ:r\)\]\+\\mathbb\{E\}\\\!\\left\[\[\\widehat\{\\Delta\}\_\{\\mathrm\{seg\}\}\]\_\{\+\}\\mid\\mathcal\{F\}\\right\]\\neq\[A\_\{\\mathrm\{seg\}\}^\{\\pi\}\(s\_\{\\mathrm\{pre\}\},\\mathbf\{a\}\_\{\\ell:r\}\)\]\_\{\+\}in general\. For example, even when the true segment advantage is zero, finite\-sample estimates can be positive or negative; positive filtering discards the negative estimates but retains the positive ones, yielding a positive expected credit bonus\. The completeProVerupdate additionally includes judge\-guided segment selection, positive filtering, assigning the resulting bonus only to tokens within the selected segment, scaling byλ\\lambda, and adding the bonus to the GRPO trajectory advantage\. Therefore, conditional unbiasedness of the raw Monte Carlo estimator does not imply that the completeProVerupdate is an unbiased estimator of the standard policy gradient\. Nor does the boundary comparison identify separate causal effects for individual actions or tokens within the segment\.相似文章
@SharonYixuanLi:扩展基于结果的强化学习无法解决长周期智能体任务。信用分配是瓶颈,而轮次级奖励…
TRACE 提出了一种轮次级奖励分配方法,利用冻结参考模型的对数概率和时间差分学习来解决长周期智能体任务中的信用分配问题,在没有评论家或过程标签的情况下,在搜索基准测试中取得了显著改进。
PGPO:面向多轮智能体任务的势引导策略优化
PGPO 提出势引导策略优化用于多轮智能体任务,使得在 LLM 后训练中实现更细粒度的信用分配,并在 ALFWorld 和 WebShop 基准测试上展示出强劲结果。
面向进度与可靠性的智能体强化学习组策略优化
ProGPO是一种免学习评论器的方法,用于LLM智能体基于组的RL中的步骤级优势估计,它使用精确前缀动作比较和基于rollout的状态势,以改善长视界任务上的信用分配。在ALFWorld和WebShop上使用Qwen2.5模型的实验表明,它优于现有的智能体RL基线。
GAGPO:广义优势分组策略优化
GAGPO提出了一种无评论家的强化学习方法,在多方交互的自主任务中,利用非参数分组价值代理进行步级信用分配,在ALFWorld和WebShop上超越了强基线模型。
AgentV-RL:用智能体验证器扩展奖励建模
AgentV-RL引入了智能体验证器框架,通过具有工具增强的前向和后向智能体进行双向验证来增强奖励建模,相比最先进的ORM实现了25.2%的性能提升。该方法通过将多轮深思熟虑过程与强化学习相结合,解决了验证器在复杂推理任务中的误差传播和基础性不足等问题。