From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
Summary
Introduces Hindsight Policy Optimization (HPO), a novel policy gradient method that uses an intent space and Wasserstein distance to reduce variance in long-horizon language agent training, showing improved stability over GRPO and PPO.
View Cached Full Text
Cached at: 07/21/26, 06:48 AM
# Leveraging Hindsight for Long-Horizon Language Agent Training
Source: [https://arxiv.org/html/2607.16257](https://arxiv.org/html/2607.16257)
## From Outcomes to Actions: Leveraging Hindsight for Long\-Horizon Language Agent Training
Tingyun LiJinyi HanXinyi WangSihang JiangYizhou YingXiaojun MengJiansheng WeiJiaqing LiangYanghua Xiao
###### Abstract
Reinforcement learning \(RL\) has become a widely adopted technique for improving large language models \(LLMs\) on complex tasks\. Despite this progress, existing RL methods still face challenges in training agents with longer\-horizon interactions\. One major bottleneck is distinguishing the contribution of different actions in long\-horizon interaction, leading to high optimization variance\. To address this, we introduce a novel policy gradient method,HindsightPolicyOptimization \(HPO\), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low\-variance learning signals from the Wasserstein distance between them\. We theoretically and empirically show that aggregating semantically similar states and actions in the intent space yields a bounded\-variance estimator and improves policy performance stably\. Our code is available online111https://github\.com/Jiangzs1028/hpo\.
Machine Learning, ICML
## 1Introduction
Reinforcement learning \(RL\) has become a widely adopted technique for improving large language models \(LLMs\) on complex decision\-making tasks\. Compared to earlier prompt\-based approaches, RL allows LLM\-based agents to actively explore and learn from interaction with the environment, and has shown strong potential in applications such as tool use\(Jinet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib1); Caoet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib24)\), smart device operation\(Luoet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib28)\), and interactive environments\([Abdulhaiet al\.,](https://arxiv.org/html/2607.16257#bib.bib8); Huet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib29)\)\.
Figure 1:Existing RL methods struggle in long\-horizon interactions\. GRPO becomes unstable with longer\-horizon interactions, while PPO remains stable but converges slowly due to critic warmup\.Despite this progress, existing RL algorithms still face challenges in training agents for longer\-horizon interactions\. As illustrated in Figure[1](https://arxiv.org/html/2607.16257#S1.F1), GRPO\(Shaoet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib22)\)can maintain stable training under shorter interaction horizons, but becomes unstable as the horizon length increases, often exhibiting sudden reward collapse\(Jinet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib1); Xueet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib2); Sunet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib27)\)\. In contrast, PPO\(Schulmanet al\.,[2017](https://arxiv.org/html/2607.16257#bib.bib21)\)achieves more stable training behavior, but typically converges more slowly due to the additional requirement of training a critic model\.
One major bottleneck is distinguishing the contribution of different actions in long\-horizon interaction, leading to high optimization variance\. Unlike single\-step tasks, multi\-turn decision\-making requires attributing outcomes to a sequence of interdependent actions, where effective and ineffective actions are often interleaved within the same trajectory\(Zhanget al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib33)\)\. However, RL methods that rely solely on outcome rewards, such as GRPO, typically propagate the final outcome uniformly across all actions, which may mistakenly reinforce some ineffective actions and thereby destabilize training\(Zhanget al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib33)\)\. In contrast, PPO uses a critic to better distinguish actions, leading to more stable training, but it converges more slowly due to the warm\-up required by the critic\(Jinet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib1)\)\.
We define the hindsight distribution as the optimal action distribution induced by an agent retrospectively reselecting its early actions after observing the final outcome, and show that existing RL methods with Monte Carlo estimation can be interpreted as apoint\-wisecomparison of thediscrepancy between the current policy distribution and this hindsight distributionin a discrete state–action space\. This point\-wise comparisonignores the underlying semantic structure of the language action space, resulting in unbounded variance in exponentially large language action spaces\.
Although the action space of large language models is vast, the set of underlying intents they can express is comparatively limited\.For example,Who is the president of the United States?andI want to know who the U\.S\. president iscorrespond to distinct language actions but convey nearly identical semantic intent\.
Based on this insight, we introduce a novel policy gradient method,HindsightPolicyOptimization \(HPO\), that projects both the current policy distribution and the hindsight distribution into an intent space and measures their discrepancy using the Wasserstein distance\(Villani and others,[2008](https://arxiv.org/html/2607.16257#bib.bib31)\)\. We theoretically show that aggregating semantically similar states and actions in this lower\-dimensional intent space improves statistical efficiency and yields a bounded\-variance estimator\.
Our contributions are as follows:
1. 1\.We reformulate existing Monte Carlo policy gradient methods from a hindsight distribution perspective, and show that point\-wise comparisons in the discrete language action space can lead to unbounded variance due to ignoring the underlying semantic structure\.
2. 2\.We propose HPO, a principled approach that constructs hindsight distributions in an intent space and derives bounded\-variance learning signals for smoother and more sample\-efficient policy optimization, with negligible computational overhead\.
3. 3\.We evaluate HPO on multiple long\-horizon tasks and show that it substantially reduces training noise and stabilizes optimization\. Further experiments show that such learning signals are interpretable and remain sufficient to drive effective policy learning even without outcome rewards\.
## 2Preliminaries
### 2\.1Task Formulation
We formulate language action tasks as a Markov decision process \(MDP\)ℳ=\(𝒮,𝒜,𝒯,ℛ,γ\)\\mathcal\{M\}=\(\\mathcal\{S\},\\mathcal\{A\},\\mathcal\{T\},\\mathcal\{R\},\\gamma\), where𝒮\\mathcal\{S\}denotes the state space,𝒜\\mathcal\{A\}is the action space of natural language actions,𝒯:𝒮×𝒜→𝒮\\mathcal\{T\}:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathcal\{S\}is the transition function,ℛ\\mathcal\{R\}is the reward function, andγ∈\(0,1\]\\gamma\\in\(0,1\]is the discount factor\. In the RLVR setting, rewards are sparse and only provided at the end of each episode and the discount factor is typically set toγ=1\\gamma=1\. When the context is clear, we omitγ\\gammafor notational simplicity\. An agent interacts with the environment over discrete time steps\. At each steptt, the agent observes the current statest∈𝒮s\_\{t\}\\in\\mathcal\{S\}and selects an actionat∼πθ\(⋅∣st\)a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\}\)\. The environment then transitions to the next statest\+1s\_\{t\+1\}according to𝒯\(st\+1∣st,at\)\\mathcal\{T\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\}\)\. The objective is to maximize the expected cumulative reward under the policyπθ\\pi\_\{\\theta\}\.
### 2\.2Value Function
ForMπ=\(S0,A0,S1,A1,⋯\)M\_\{\\pi\}=\(S\_\{0\},A\_\{0\},S\_\{1\},A\_\{1\},\\cdots\), there is an occupancy measureρπ\\rho\_\{\\pi\}which satisfiesρπ\(s,a\)=\(1−γ\)∑t=0∞γtℙ\(St=s,At=a\)\\rho\_\{\\pi\}\(s,a\)=\(1\-\\gamma\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}\(S\_\{t\}=s,A\_\{t\}=a\)\. Moreover, there is a one\-to\-one correspondence between those measures and the policies\. The state\-value function under policyπ\\piis defined asVπ\(s\)≜𝔼π\[∑t=0∞γtrt∣S0=s\]V\_\{\\pi\}\(s\)\\triangleq\\mathbb\{E\}\_\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{t\}\\mid S\_\{0\}=s\\right\], which measures the expected discounted return starting from statess\. Similarly, the action\-value function is defined asQπ\(s,a\)≜𝔼π\[∑t=0∞γtrt∣S0=s,A0=a\]Q\_\{\\pi\}\(s,a\)\\triangleq\\mathbb\{E\}\_\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{t\}\\mid S\_\{0\}=s,A\_\{0\}=a\\right\]\.
### 2\.3Wasserstein Distance
Wasserstein distance, originating from optimal transport theory\(Villani and others,[2008](https://arxiv.org/html/2607.16257#bib.bib31)\), measures the discrepancy between two probability distributions by the minimum cost of transporting mass from one distribution to the other\. Given two probability measuresμA\\mu\_\{A\}andμB\\mu\_\{B\}on a metric space\(𝒳,dis\)\(\\mathcal\{X\},\\text\{dis\}\), the 1\-Wasserstein distance is defined as
W1\(μA,μB\)=infπ∈Π\(μA,μB\)∫𝒳×𝒳dis\(xA,xB\)𝑑π\(xA,xB\),W\_\{1\}\(\\mu\_\{A\},\\mu\_\{B\}\)=\\inf\_\{\\pi\\in\\Pi\(\\mu\_\{A\},\\mu\_\{B\}\)\}\\int\_\{\\mathcal\{X\}\\times\\mathcal\{X\}\}\\text\{dis\}\(x\_\{A\},x\_\{B\}\)\\,d\\pi\(x\_\{A\},x\_\{B\}\),\(1\)whereΠ\(μA,μB\)\\Pi\(\\mu\_\{A\},\\mu\_\{B\}\)is the set of couplings whose marginals areμA\\mu\_\{A\}andμB\\mu\_\{B\}\. Intuitively, each coupling can be viewed as a possible transportation plan that describes how mass fromμA\\mu\_\{A\}is assigned to mass inμB\\mu\_\{B\}\. In practice,W1W\_\{1\}is often computed via its dual formulation:
W1\(μA,μB\)=sup‖f‖L≤1\(𝔼x∼μA\[f\(x\)\]−𝔼x∼μB\[f\(x\)\]\),W\_\{1\}\(\\mu\_\{A\},\\mu\_\{B\}\)=\\sup\_\{\\\|f\\\|\_\{L\}\\leq 1\}\\left\(\\mathbb\{E\}\_\{x\\sim\\mu\_\{A\}\}\[f\(x\)\]\-\\mathbb\{E\}\_\{x\\sim\\mu\_\{B\}\}\[f\(x\)\]\\right\),\(2\)
where the maximization is over all 1\-Lipschitz functionsff\. The optimal functionf⋆f^\{\\star\}in this dual problem is referred to as the*Kantorovich potential*\(Villani and others,[2008](https://arxiv.org/html/2607.16257#bib.bib31)\)\.
## 3The Challenge of Long\-Horizon Language Agent Reinforcement Learning
Existing studies\(Jinet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib1); Xueet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib2); Sunet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib27)\)have found that RL methods become increasingly unstable as the interaction horizon grows\. In this section, we show that this instability arises from the high\-variance gradient noise induced by long\-horizon interactions, which free\-critic methods struggle to mitigate and may even further exacerbate\.
### 3\.1Rethinking the Objective in RL for LLMs
For a unified analysis, common policy optimization methods used in RL for LLMs, including GRPO, REINFORCE\+\+\(Huet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib23)\), and PPO, can all be expressed in the following generic policy gradient form like REINFORCE\(Williams,[1992](https://arxiv.org/html/2607.16257#bib.bib36)\):
∇θ𝒥\(θ\)=𝔼τ∼πθ\[∑t=0∞∇θlogπθ\(at∣st\)\(Q^\(st,at\)−bt\)\]\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\\tau\\sim\\pi\_\{\\theta\}\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)\(\\hat\{Q\}\(s\_\{t\},a\_\{t\}\)\-b\_\{t\}\)\\right\]\(3\)
whereQ^\(st,at\)\\hat\{Q\}\(s\_\{t\},a\_\{t\}\)denotes an estimator ofQπ\(st,at\)Q\_\{\\pi\}\(s\_\{t\},a\_\{t\}\)\.
For example, the objective in GRPO can be interpreted as repeatedly samplingGGtrajectories from a shared initial states0s\_\{0\}\. For analytical convenience, we assume equal return variance across differents0s\_\{0\}\. Concretely, GRPO optimizes
∇θ𝒥GRPO\(θ\)=𝔼\{τi\}i=1G∼πθ\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{GRPO\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{G\}\\sim\\pi\_\{\\theta\}\}\[1G∑i=1G∑t=0∞∇θlogπθ\(ai,t∣si,t\)\(Q^\(si,t,ai,t\)−bi,t\)\],\\displaystyle\\left\[\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\sum\_\{t=0\}^\{\\infty\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\_\{i,t\}\\mid s\_\{i,t\}\)\\big\(\\hat\{Q\}\(s\_\{i,t\},a\_\{i,t\}\)\-b\_\{i,t\}\\big\)\\right\],\(4\)
whereQ^\(si,t,ai,t\)=R\(τi\)\\hat\{Q\}\(s\_\{i,t\},a\_\{i,t\}\)=R\(\\tau\_\{i\}\)andbi,t=1G∑j=1GR\(τj\)b\_\{i,t\}=\\frac\{1\}\{G\}\\sum\_\{j=1\}^\{G\}R\(\\tau\_\{j\}\)\. The latter can be viewed as a Monte Carlo \(MC\) estimateV¯π\(s0\)\\bar\{V\}\_\{\\pi\}\(s\_\{0\}\)ofVπ\(s0\)V\_\{\\pi\}\(s\_\{0\}\)\.
Although MC estimation provides an unbiased estimate ofQπQ\_\{\\pi\}, it suffers from excessive variance by accumulating stochasticity from future actions\(Suttonet al\.,[1998](https://arxiv.org/html/2607.16257#bib.bib34)\)\. For example, even if the agent takes a correct action at the current step, the final return still depends on every later sampled action: a wrong one can turn an otherwise successful trajectory into a failure\. Therefore, the same current action may receive different terminal returns across rollouts, and this variance can grow exponentially with the horizon\.
It is well known that excessive variance can hinder effective training updates, so existing RL algorithms typically seek to reduce training variance\. Algorithms such as GRPO introduce a fixed baseline to reduce variance without introducing bias, and have achieved strong performance on single\-turn tasks such as mathematical reasoning\. However, we show below that the effectiveness of this approach diminishes significantly in long\-horizon interactions\.
###### Lemma 3\.1\(Variance Decomposition\)\.
Consider the REINFORCE gradient estimatorG^b\\hat\{G\}\_\{b\}with the baselinebb,
G^b=g\(a\)\(Q^\(s,a\)−b\(s\)\),\\hat\{G\}\_\{b\}\\;=\\;g\(a\)\\,\\big\(\\hat\{Q\}\(s,a\)\-b\(s\)\\big\),\(5\)wherea∼π\(⋅∣s\)a\\sim\\pi\(\\cdot\\mid s\)andg\(a\)=∇θlogπθ\(a∣s\)g\(a\)=\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\.
Let∥⋅∥\\\|\\cdot\\\|denotes the Euclidean norm\. Then its conditional variance admits the decomposition
Var\(G^b∣s\)=𝔼\[‖g\(a\)‖2∣s\]⋅Var\(Q^\(s,a\)−b\(s\)∣s\)\\displaystyle\\mathrm\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)\\\!=\\\!\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\\\!\\cdot\\\!\\mathrm\{Var\}\\\!\\left\(\\hat\{Q\}\(s,a\)\-b\(s\)\\mid s\\right\)\(6\)
#### The gradient norm\.
The gradient\-norm factor in Eq\. \([6](https://arxiv.org/html/2607.16257#S3.E6)\) can magnify optimization noise\. With rounds of external tool feedback, some low\-probability actions are increasingly sampled \(e\.g\., meaningless or repetitive tokens\)\(Xueet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib2)\)\. Due to the logarithmic form of the policy gradient, these rare actions can induce disproportionately large, and even unbounded gradient norms, which further amplify the noise variance caused by state mismatch in Eq\. \([6](https://arxiv.org/html/2607.16257#S3.E6)\)\. This perspective also helps explain why filtering strategies that remove low\-probability samples can mitigate collapse\(Xueet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib2)\)\.
#### The variance of the baseline\-adjustedQQsignal\.
The second factor of the gradient variance is the variance of the baseline\-adjustedQQsignal,Q^\(s,a\)−b\(s\)\\hat\{Q\}\(s,a\)\-b\(s\)\. In principle, there exists an optimal baseline that minimizes this variance, but it is generally difficult to compute in practice\. Existing methods therefore approximate it in different ways: GRPO uses a fixed baseline, while PPO learns a critic\-based baseline\.
###### Theorem 3\.3\(Optimal Baseline and Excess Variance\)\.
Among all scalar baselinesb\(s\)b\(s\)independent of the sampled actionaa, the conditional varianceVar\(G^b∣s\)\\mathrm\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)is minimized by
b∗\(s\)=𝔼\[‖g\(a\)‖2Q^\(s,a\)∣s\]𝔼\[‖g\(a\)‖2∣s\]\.b^\{\*\}\(s\)=\\frac\{\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\hat\{Q\}\(s,a\)\\mid s\\right\]\}\{\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\}\.\(7\)Moreover, for any alternative baselineb~\(s\)\\tilde\{b\}\(s\), the increase in variance admits the exact expression
Var\(G^b~∣s\)\\displaystyle\\mathrm\{Var\}\(\\hat\{G\}\_\{\\tilde\{b\}\}\\mid s\)−Var\(G^b∗∣s\),=\\displaystyle\\,\-\\,\\mathrm\{Var\}\(\\hat\{G\}\_\{b^\{\*\}\}\\mid s\),\\,=\\,𝔼\[‖g\(a\)‖2∣s\]\(b~\(s\)−b∗\(s\)\)2\.\\displaystyle\\mathbb\{E\}\\\!\\left\[\\,\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\(\\tilde\{b\}\(s\)\-b^\{\*\}\(s\)\)^\{2\}\.\(8\)
We remark that even though the optimal state\-only baseline is known, it is rarely used in practice due to its computational intractability\(Liet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib49); Duanet al\.,[2016](https://arxiv.org/html/2607.16257#bib.bib50)\)\. Rather, for both computational and conceptual benefit, the choice of𝔼\[Q^\(s,a\)∣s\]\\mathbb\{E\}\[\\hat\{Q\}\(s,a\)\\mid s\]is often used\(Schulmanet al\.,[2017](https://arxiv.org/html/2607.16257#bib.bib21); Wuet al\.,[2018](https://arxiv.org/html/2607.16257#bib.bib51)\)\. In particular, when‖g\(a\)‖2\\\|g\(a\)\\\|^\{2\}andQ^\(s,a\)\\hat\{Q\}\(s,a\)are loosely correlated, this baseline is close to the optimal baseline\. For ease of analysis and empirical evaluation, we follow this setting\.
Under this proxy, the minimum\-variance baseline for MC returns is the state\-dependent value functionVπ\(st\)=𝔼\[R∣st\]V\_\{\\pi\}\(s\_\{t\}\)=\\mathbb\{E\}\[R\\mid s\_\{t\}\]\. GRPO approximate it with a fixed baselineV¯π\(s0\)\\bar\{V\}\_\{\\pi\}\(s\_\{0\}\)shared across all subsequent time steps\. This approximation can be effective in single\-turn tasks, but becomes inaccurate in long\-horizon, tool\-augmented interactions, whereVπ\(st\)V\_\{\\pi\}\(s\_\{t\}\)can vary dramatically across states: for instance, after a successful search, the agent is much more likely to find the correct final answer, whereas after a failed search, its chance of success may drop sharply\.
To further analyze how this mismatch changes with the interaction horizon, we conduct a preliminary multi\-turn Search experiment, where states are grouped by identical environment observations to estimateVπ\(st\)V\_\{\\pi\}\(s\_\{t\}\)\. As shown in Figure[2](https://arxiv.org/html/2607.16257#S3.F2), as the interaction horizon increases, the mismatch\(Vπ\(st\)−V¯π\(s0\)\)2\(V\_\{\\pi\}\(s\_\{t\}\)\-\\bar\{V\}\_\{\\pi\}\(s\_\{0\}\)\)^\{2\}can grow large, directly weakening the variance\-reduction effect of the baseline by Theorem[3\.3](https://arxiv.org/html/2607.16257#S3.Thmtheorem3)and, in extreme cases, leading to higher gradient variance than using no baseline at all\.
Figure 2:Multi\-turn Search experiment where states are grouped by identical environment observations\. As the maximum interaction turnsnngrows, the mismatch betweenVπ\(st\)V\_\{\\pi\}\(s\_\{t\}\)and the fixed baselineV¯π\(s0\)\\bar\{V\}\_\{\\pi\}\(s\_\{0\}\)can become increasingly severe\.SummaryAs the interaction horizon increases, the variance of RL training can grow substantially, while existing critic\-free methods struggle to reduce it and may even exacerbate it\. This motivates the need for a estimator that remains low\-variance in long\-horizon interactions\.
Figure 3:Illustration of HPO\. HPO performs semantic projection and compares the policy distribution with the hindsight distribution in the intent space\. By solving the Wasserstein distance between the two distributions, it obtains step\-level advantages, which are combined with episode\-level advantages to update the policy\.
## 4Training LLM Agents with HPO
In this section, we introduce a novel policy gradient estimator\. Unlike Monte Carlo estimators defined over discrete action spaces, it constructs learning signals in a semantically aggregated intent space\. By sharing statistical information at the semantic level, it improves statistical efficiency and substantially reduces estimation variance\.
### 4\.1Policy Gradient from the Hindsight Perspective
It is well known that RL becomes challenging in large action spaces\. Compared with single\-turn tasks, aTT\-turn interaction induces an action\-sequence space whose size scales as\|𝒜\|T\|\\mathcal\{A\}\|^\{T\}, growing exponentially with the interaction horizon\. For language agents, this difficulty is further amplified because each action is a natural\-language output, which can itself be viewed as a joint choice over a sequence of tokens and thus induces an enormous action space\.
To analyze the trajectory distribution, we consider their induced state–action occupancy measures over pairs\(s,a\)\(s,a\), and define the hindsight distribution and show that the policy gradient can be equivalently reformulated from the perspective of the hindsight trajectory distribution\.
###### Definition 4\.1\(Hindsight Policy Distribution\)\.
For any statess, actionaa, and return valuezz, lethzπ\(a∣s\)h\_\{z\}^\{\\pi\}\(a\\mid s\)denote the conditional probability that the first action equalsaawhen sampling a trajectory fromπ\\pistarting atss, given that the realized return equalszz\. In the binary\-reward setting considered in this work, we takez=1z=1and define the corresponding hindsight occupancy measure as
ρπh\(s,a\)=∑t=0∞ℙπ\(\(st,at\)=\(s,a\)\|∑t′=t∞rt′=1\)\.\\rho\_\{\\pi\}^\{h\}\(s,a\)=\\sum\_\{t=0\}^\{\\infty\}\\mathbb\{P\}\_\{\\pi\}\\\!\\left\(\(s\_\{t\},a\_\{t\}\)=\(s,a\)\\;\\middle\|\\;\\sum\_\{t^\{\\prime\}=t\}^\{\\infty\}r\_\{t^\{\\prime\}\}=1\\right\)\.\(9\)
Intuitively, in the binary\-reward setting, this distribution is obtained by retaining only state\-action pairs from trajectories that eventually succeed\. For continuous\-reward settings, we provide the corresponding distribution definition in Appendix[B](https://arxiv.org/html/2607.16257#A2)\.
###### Lemma 4\.2\.
Consider the objective of policy optimization𝒥\(θ\)=𝔼τ∼π\[∑t≥0rt\]\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\\tau\\sim\\pi\}\[\\sum\_\{t\\geq 0\}r\_\{t\}\]\. Its policy gradient admits the following two equivalent forms:
∇θ\\displaystyle\\nabla\_\{\\theta\}𝒥\(θ\)=𝔼\(s,a\)∼ρπ\[Qπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[Q\_\{\\pi\}\(s,a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(10\)=𝔼\(s,a\)∼ρπ\[−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(11\)where−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)=ρπh\(s,a\)ρπ\(s,a\)\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}=\\frac\{\\rho^\{h\}\_\{\\pi\}\(s,a\)\}\{\\rho\_\{\\pi\}\(s,a\)\}denotes the point\-wise variational derivative ofKL\(ρπh\|\|ρπ\)KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)with respect toρπ\(s,a\)\\rho\_\{\\pi\}\(s,a\)\.
A detailed proof is provided in Appendix[D\.3](https://arxiv.org/html/2607.16257#A4.SS3)\. Lemma[4\.2](https://arxiv.org/html/2607.16257#S4.Thmtheorem2)provides an alternative interpretation of policy gradients from the perspective of the hindsight distribution: Monte Carlo policy\-gradient methods can be viewed as comparing the current occupancyρπ\\rho\_\{\\pi\}with the hindsight occupancyρπh\\rho\_\{\\pi\}^\{h\}through the point\-wise ratioρπh\(s,a\)/ρπ\(s,a\)\\rho\_\{\\pi\}^\{h\}\(s,a\)/\\rho\_\{\\pi\}\(s,a\)\.
This point\-wise comparison treats different state–action pairs as unrelated atoms\. Such treatment can be effective in traditional RL problems with small action spaces, where the same state–action pair can be revisited frequently\. However, in long\-horizon language\-agent training, the number of possible atoms grows with both the per\-step language action space and the interaction horizon\. Semantically similar actions with different surface forms are counted as distinct atoms and cannot share statistical evidence, making value estimation for language\-agent actions highly inefficient\.
This suggests that reducing the effective estimation space is necessary for obtaining stable learning signals\. Although the action space of LLMs is exponentially large, the space of intents they can express is significantly lower\-dimensional\. Many distinct actions are semantically similar and lead to comparable outcomes; for example,Who is the president of the United States?andI want to know who the U\.S\. president iscorrespond to different language actions but express nearly identical underlying intent\.
Based on this insight, the following sections introduce how to leverage the semantics of language actions to reduce the effective size of the action space and obtain more stable value estimation\.
### 4\.2Intent Space Construction
Given a state space𝒮\\mathcal\{S\}and an action space𝒜\\mathcal\{A\}, each state–action pair\(s,a\)\(s,a\)is embedded into add\-dimensional Hilbert semantic space via a pretrained encoderΦ:𝒮×𝒜→ℝd\\Phi:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathbb\{R\}^\{d\}, where the Euclidean distance serves as the metric of the embedding space\. In this semantic space, interactions that are lexically different but semantically equivalent are mapped to nearby regions\.
In the binary\-reward setting, the policy distributionρπ\\rho\_\{\\pi\}and hindsight distributionρπh\\rho\_\{\\pi\}^\{h\}are constructed as empirical distributions over embedded state–action pairs in the intent space\.
ρπ=\\displaystyle\\rho\_\{\\pi\}=∑t=0∞𝔼π\[δΦ\(st,at\)\],\\displaystyle\\sum\_\{t=0\}^\{\\infty\}\\mathbb\{E\}\_\{\\pi\}\\\!\\left\[\\delta\_\{\\Phi\(s\_\{t\},a\_\{t\}\)\}\\right\],ρπh=\\displaystyle\\rho\_\{\\pi\}^\{h\}=∑t=0∞𝔼π\[𝟏\(∑t′=t∞rt′=1\)δΦ\(st,at\)\],\\displaystyle\\sum\_\{t=0\}^\{\\infty\}\\mathbb\{E\}\_\{\\pi\}\\\!\\left\[\\mathbf\{1\}\\\!\\left\(\\sum\_\{t^\{\\prime\}=t\}^\{\\infty\}r\_\{t^\{\\prime\}\}=1\\right\)\\;\\delta\_\{\\Phi\(s\_\{t\},a\_\{t\}\)\}\\right\],\(12\)whereδΦ\(st,at\)\\delta\_\{\\Phi\(s\_\{t\},a\_\{t\}\)\}denotes the Dirac measure centered at the intent embeddingΦ\(st,at\)\\Phi\(s\_\{t\},a\_\{t\}\)\.
#### Data for constructing the hindsight distribution\.
Although constructing the hindsight distribution from on\-policy rollouts yields a closer approximation to the on\-policy value functionQπQ\_\{\\pi\}, it suffers from unstable estimation due to fluctuating successful rollouts\. Alternatively, using an offline set of successful trajectories obtained via rejection sampling from the initial policy provides a more stable construction of the hindsight distribution, particularly in sparse\-reward scenarios, where on\-policy rollouts rarely yield successful trajectories\. We refer to these two variants asonline HPOandoffline HPO, respectively\.
### 4\.3Hindsight Policy Optimization
Compared with existing RL estimators that use KL\-induced point\-wise estimates in the discrete state–action space, we seek a new value estimator that preserves the optimal hindsight policy while reducing value\-estimation fluctuations among semantically similar language actions\. In the intent space, this desideratum can be formalized as the following constrained scoring problem\.
###### Proposition 4\.3\(Intent value estimatior\)\.
Let𝒵=Φ\(𝒮×𝒜\)\\mathcal\{Z\}=\\Phi\(\\mathcal\{S\}\\times\\mathcal\{A\}\)be the intent space equipped with metricdd, and letρπ\\rho\_\{\\pi\}andρπh\\rho\_\{\\pi\}^\{h\}denote the current and hindsight occupancy measures on𝒵\\mathcal\{Z\}\. A scoring function that identifies hindsight\-optimal trajectories among trajectories under the semantic smoothness constraint‖v‖L≤1\\\|v\\\|\_\{L\}\\leq 1is given by
sup‖v‖L≤1\(𝔼z∼ρπh\[v\(z\)\]−𝔼z∼ρπ\[v\(z\)\]\)\.\\sup\_\{\\\|v\\\|\_\{L\}\\leq 1\}\\left\(\\mathbb\{E\}\_\{z\\sim\\rho\_\{\\pi\}^\{h\}\}\[v\(z\)\]\-\\mathbb\{E\}\_\{z\\sim\\rho\_\{\\pi\}\}\[v\(z\)\]\\right\)\.\(13\)Any optimizerv∗v^\{\*\}is a Kantorovich potential of the 1\-Wasserstein distance\.
We find that the intent value evaluator is exactly the dual problem of the Wasserstein distance betweenρπh\\rho\_\{\\pi\}^\{h\}andρπ\\rho\_\{\\pi\}introduced in Section[2\.3](https://arxiv.org/html/2607.16257#S2.SS3)\. Under this view, the resulting policy gradient has the same form as Eq\. \([11](https://arxiv.org/html/2607.16257#S4.E11)\), except that the KL divergence is replaced by the Wasserstein distance\.
Letf∗f^\{\*\}denote the Kantorovich potential forW1\(ρπ,ρπh\)W\_\{1\}\(\\rho\_\{\\pi\},\\rho\_\{\\pi\}^\{h\}\)\. The corresponding variational gradient gives the following policy\-gradient form:
∇θ𝒥\(θ\)=𝔼\(s,a\)∼ρπ\[−δW1\(ρπ\|\|ρπh\)δρπ\(s,a\)∇θlogπθ\(a∣s\)\]\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[\-\\frac\{\\delta W\_\{1\}\(\\rho\_\{\\pi\}\|\|\\rho^\{h\}\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(14\)whereδW1\(ρπ\|\|ρπh\)δρπ\(s,a\)=f∗\(s,a\)\\frac\{\\delta W\_\{1\}\(\\rho\_\{\\pi\}\|\|\\rho^\{h\}\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}=f^\{\*\}\(s,a\)denotes the point\-wise variational derivative ofW1\(ρπ\|\|ρπh\)W\_\{1\}\(\\rho\_\{\\pi\}\|\|\\rho^\{h\}\_\{\\pi\}\)with respect toρπ\(s,a\)\\rho\_\{\\pi\}\(s,a\), i\.e\., the Kantorovich potential in Section[2\.3](https://arxiv.org/html/2607.16257#S2.SS3)\.
f∗\(s,a\)f^\{\*\}\(s,a\)possesses several desirable properties\. Notably, it satisfies the 1\-Lipschitz continuity condition in the intent space, implying that behaviors mapped to similar intents yield close potential values\. Furthermore, it is also computationally efficient\(Ambrosioet al\.,[2005](https://arxiv.org/html/2607.16257#bib.bib52); Alvarez\-Melis and Fusi,[2021](https://arxiv.org/html/2607.16257#bib.bib38)\)in thatf∗\(s,a\)f^\{\*\}\(s,a\)can be recovered directly from the dual solution of the Wasserstein distance calculation at virtually no additional cost\(Alvarez\-Melis and Fusi,[2021](https://arxiv.org/html/2607.16257#bib.bib38); Justet al\.,[2023](https://arxiv.org/html/2607.16257#bib.bib40)\)\.
As proved in Section[4\.4](https://arxiv.org/html/2607.16257#S4.SS4),−δW1\(ρπ∥ρπh\)δρπ\(s,a\)=−f∗\(s,a\)\-\\frac\{\\delta W\_\{1\}\(\\rho\_\{\\pi\}\\,\\\|\\,\\rho^\{h\}\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}=\-f^\{\*\}\(s,a\), which is defined in the intent space, yields a learning signal with lower variance, at the cost of introducing a certain degree of bias\. In contrast, the learning signal−δKL\(ρπh∥ρπ\)δρπ\(s,a\)=ρπh\(s,a\)ρπ\(s,a\)\-\\frac\{\\delta\\mathrm\{KL\}\(\\rho^\{h\}\_\{\\pi\}\\,\\\|\\,\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}=\\frac\{\\rho^\{h\}\_\{\\pi\}\(s,a\)\}\{\\rho\_\{\\pi\}\(s,a\)\}, which is defined over the discrete action space, is unbiased but typically exhibits much higher variance\.
Since GRPO is currently one of the most widely adopted algorithms in RLVR, we choose it as the base optimization algorithm and apply the proposed hindsight distribution to its policy update\. Concretely, the optimization terms induced by the KL divergence and the Wasserstein distance admit intuitive interpretations as*episode\-level*and*step\-level*advantages in GRPO, respectively\.
#### KL term \(Episode\-level advantage\)\.
The KL\-based optimization term corresponds to a Monte Carlo estimate ofQπ\(s,a\)Q\_\{\\pi\}\(s,a\)\. Therefore, we directly reuse the episode\-level advantage originally defined in GRPO as the learning signal for the KL term:
AE\(τi\)=Ri−mean\(\{Ri\}i=1G\)std\(\{Ri\}i=1G\),A\_\{E\}\(\\tau\_\{i\}\)=\\frac\{R\_\{i\}\-\\operatorname\{mean\}\\\!\\left\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\right\)\}\{\\operatorname\{std\}\\\!\\left\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\right\)\},\(15\)
#### Wasserstein term \(Step\-level advantage\)\.
For the Wasserstein\-based optimization term, we obtain the Kantorovich potential−δW1\(ρπ∥ρπh\)δρπ\(s,a\)=−f∗\(st,at\)\-\\frac\{\\delta W\_\{1\}\(\\rho\_\{\\pi\}\\,\\\|\\,\\rho^\{h\}\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}=\-f^\{\*\}\(s\_\{t\},a\_\{t\}\)defined in the intent space\. To ensure compatibility in scale with the episode\-level advantage and to enable a meaningful interpolation between the two signals, we normalize the potential into a step\-level advantage:
AS\(st,at\)=−f∗\(st,at\)−mean\(−f∗\)std\(f∗\)\.A\_\{S\}\(s\_\{t\},a\_\{t\}\)=\\frac\{\-f^\{\*\}\(s\_\{t\},a\_\{t\}\)\-\\text\{mean\}\(\-f^\{\*\}\)\}\{\\operatorname\{std\}\(f^\{\*\}\)\}\.\(16\)
#### Final advantage formulation\.
We define the final advantage as a weighted interpolation between the episode\-level and step\-level advantages, to achieve a better bias–variance trade\-off:
A^\(st\(i\),at\(i\)\)=11\+ω\[AE\(τi\)\+ω⋅AS\(st\(i\),at\(i\)\)\],\\hat\{A\}\(s^\{\(i\)\}\_\{t\},a^\{\(i\)\}\_\{t\}\)=\\frac\{1\}\{1\+\\omega\}\\left\[A\_\{E\}\(\\tau\_\{i\}\)\+\\omega\\cdot A\_\{S\}\(s^\{\(i\)\}\_\{t\},a^\{\(i\)\}\_\{t\}\)\\right\],\(17\)whereω\>0\\omega\>0controls the relative contribution of the step\-level hindsight signal\. By tuningω\\omega, this formulation provides an explicit mechanism to trade off variance reduction against bias, enabling more stable and effective long\-horizon optimization\.
Let𝒟\\mathcal\{D\}denote the training prompt distribution, and letq∼𝒟q\\sim\\mathcal\{D\}denote a sampled task prompt\. Then the policy optimization objective of HPO is:
𝒥\\displaystyle\\mathcal\{J\}\(θ\)HPO=𝔼q∼𝒟,\{τ\(i\)\}i=1G∼πold\(⋅∣q\)1G∑i=1G1\|τ\(i\)\|∑t=1\|τ\(i\)\|\{\}\_\{\\mathrm\{HPO\}\}\(\\theta\)=\\mathbb\{E\}\_\{q\\sim\\mathcal\{D\},\\left\\\{\\tau^\{\(i\)\}\\right\\\}\_\{i=1\}^\{G\}\\sim\\pi\_\{\\mathrm\{old\}\}\(\\cdot\\mid q\)\}\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\frac\{1\}\{\\left\|\\tau^\{\(i\)\}\\right\|\}\\sum\_\{t=1\}^\{\\left\|\\tau^\{\(i\)\}\\right\|\}min\[wt\(i\)\(θ\)A^t\(i\),clip\(wt\(i\)\(θ\),1±ϵ\)A^t\(i\)\],\\displaystyle\\min\\left\[w\_\{t\}^\{\(i\)\}\(\\theta\)\\hat\{A\}\_\{t\}^\{\(i\)\},\\operatorname\{clip\}\\left\(w\_\{t\}^\{\(i\)\}\(\\theta\),1\\pm\\epsilon\\right\)\\hat\{A\}\_\{t\}^\{\(i\)\}\\right\],\(18\)wherewt\(i\)=πθ\(yt\(i\)\|q,y<t\(i\)\)πold\(yt\(i\)\|q,y<t\(i\)\)w\_\{t\}^\{\(i\)\}=\\frac\{\\pi\_\{\\theta\}\(y\_\{t\}^\{\(i\)\}\|q,y\_\{<t\}^\{\(i\)\}\)\}\{\\pi\_\{\\text\{old\}\}\(y\_\{t\}^\{\(i\)\}\|q,y\_\{<t\}^\{\(i\)\}\)\}is the importance sampling ratio between the current policy and the behavior policy\.
The overall workflow ofHPOis illustrated in Figure[3](https://arxiv.org/html/2607.16257#S3.F3), and the complete algorithm is summarized in Algorithm[1](https://arxiv.org/html/2607.16257#alg1)\.
Algorithm 1Hindsight Policy OptimizationInput:initial policy model
πθinit\\pi\_\{\\theta\_\{\\text\{init\}\}\}; task prompts
𝒟\\mathcal\{D\}; hyperparameters
ε,β,μ,ω\\varepsilon,\\beta,\\mu,\\omega
policy model
πθ←πθinit\\pi\_\{\\theta\}\\leftarrow\\pi\_\{\\theta\_\{\\text\{init\}\}\}
forstep
=1,…,M=1,\\ldots,Mdo
Sample a batch
𝒟b\\mathcal\{D\}\_\{b\}from
𝒟\\mathcal\{D\}
Update the old policy model
πθold←πθ\\pi\_\{\\theta\_\{\\text\{old\}\}\}\\leftarrow\\pi\_\{\\theta\}
Sample outputs
\{oi\}i=1G∼πθold\(⋅∣q\)\\\{o\_\{i\}\\\}\_\{i=1\}^\{G\}\\sim\\pi\_\{\\theta\_\{\\text\{old\}\}\}\(\\cdot\\mid q\)for each
q∈𝒟bq\\in\\mathcal\{D\}\_\{b\}
Compute rewards
\{ri\}i=1G\\\{r\_\{i\}\\\}\_\{i=1\}^\{G\}for each output
oio\_\{i\}
Get embeddings by encoding
\(s,a\)∈\{τi\}i=1G\(s,a\)\\in\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{G\}with
Φ\\Phi
Construct distributions
ρπ\\rho\_\{\\pi\},
ρπh\\rho^\{h\}\_\{\\pi\}using Eq\. \([4\.2](https://arxiv.org/html/2607.16257#S4.Ex3)\)
Get
f∗\{f^\{\*\}\}from computing
𝒲\(ρπ,ρπh\)\\mathcal\{W\}\(\\rho\_\{\\pi\},\\rho^\{h\}\_\{\\pi\}\)
Compute
A^t\(i\)\\hat\{A\}^\{\(i\)\}\_\{t\}according to Eq\. \([17](https://arxiv.org/html/2607.16257#S4.E17)\)
foriteration
=1,…,μ=1,\\ldots,\\mudo
Update the policy
πθ\\pi\_\{\\theta\}with the objective in Eq\. \([4\.3](https://arxiv.org/html/2607.16257#S4.Ex4)\)\.
endfor
endfor
Output:
πθ\\pi\_\{\\theta\}
### 4\.4Theoretical Analysis
###### Lemma 4\.4\.
Letρπ\\rho\_\{\\pi\}be a state\-action occupancy measure, and defineD=sup\(s,a\),\(s′,a′\)∈ρπd\(\(s,a\),\(s′,a′\)\)D=\\sup\_\{\(s,a\),\(s^\{\\prime\},a^\{\\prime\}\)\\in\\rho\_\{\\pi\}\}d\\big\(\(s,a\),\(s^\{\\prime\},a^\{\\prime\}\)\\big\)\. Letf∗\(s,a\)f^\{\*\}\(s,a\)denote the−δW1\(ρπ\|\|ρπh\)δρπ\(s,a\)\-\\frac\{\\delta W\_\{1\}\(\\rho\_\{\\pi\}\|\|\\rho^\{h\}\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}, andg\(s,a\)g\(s,a\)denote the−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}, then we have
Varρπ\(f∗\(s,a\)\)≤D2/4\.\\displaystyle\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(f^\{\*\}\(s,a\)\)\\leq D^\{2\}/4\.\(19\)Varρπ\(g\(s,a\)\)=∫ρπh\(s,a\)2ρπ\(s,a\)d\(s,a\)−1=χ2\(ρπh∥ρπ\)\.\\displaystyle\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(g\(s,a\)\)=\\int\\frac\{\\rho\_\{\\pi\}^\{h\}\(s,a\)^\{2\}\}\{\\rho\_\{\\pi\}\(s,a\)\}\\,d\(s,a\)\-1=\\chi^\{2\}\(\\rho\_\{\\pi\}^\{h\}\\\|\\rho\_\{\\pi\}\)\.\(20\)whereχ2\(⋅∥⋅\)\\chi^\{2\}\(\\cdot\\\|\\cdot\)denotes the chi\-squared divergence\.
The lemma highlights a sharp contrast between the learning signals induced by Wasserstein and KL\. Thef∗\(s,a\)f^\{\*\}\(s,a\)is 1\-Lipschitz over the intent space and thus admits avariance boundVarρπ\(f∗\)≤D2/4\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(f^\{\*\}\)\\leq D^\{2\}/4\. In contrast, theg\(s,a\)=ρπh\(s,a\)/ρπ\(s,a\)g\(s,a\)=\\rho\_\{\\pi\}^\{h\}\(s,a\)/\\rho\_\{\\pi\}\(s,a\)has variance characterized byχ2\(ρπh∥ρπ\)\\chi^\{2\}\(\\rho\_\{\\pi\}^\{h\}\\\|\\rho\_\{\\pi\}\), which can bearbitrarily large\.
This contrast explains the instability of KL\-based signals in large discrete action spaces\. The KL gradient depends on the pointwise ratioρπh\(s,a\)/ρπ\(s,a\)\\rho\_\{\\pi\}^\{h\}\(s,a\)/\\rho\_\{\\pi\}\(s,a\), whose variance is governed byχ2\(ρπh\|ρπ\)\\chi^\{2\}\(\\rho\_\{\\pi\}^\{h\}\|\\rho\_\{\\pi\}\)and can become arbitrarily large whenρπ\(s,a\)\\rho\_\{\\pi\}\(s,a\)is small\. Since discrete actions are treated as isolated atoms, no statistical sharing occurs across actions, making reliable estimation of such ratios sample\-inefficient, especially in the presence of rare but important events\. In contrast, the Wasserstein signal leverages the geometry of the intent space to compare distributions via mass transport, allowing probability to be matched across nearby intents without pointwise division\. This yields a smoother learning signal with variance bounded by the metric diameterDD, providing a principled justification forHPO’s intent\-space formulation\.
Table 1:Performance on SearchQA\. Results are averaged over 3 random seeds\.Compared with other baselines, HPO achieves consistent improvements\.Table 2:The results on Textcraft benchmark\.
## 5Experiments
### 5\.1Setup
#### Tasks and Benchmarks\.
We evaluate HPO on two long\-horizon interactive tasks: SearchQA and TextCraft\(Prasadet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib9)\)\. For SearchQA, we consider seven benchmarks, including single\-hop QA \(NQ\(Kwiatkowskiet al\.,[2019](https://arxiv.org/html/2607.16257#bib.bib41)\), TriviaQA\(Joshiet al\.,[2017](https://arxiv.org/html/2607.16257#bib.bib42)\), and PopQA\(Mallenet al\.,[2022](https://arxiv.org/html/2607.16257#bib.bib43)\)\) and multi\-hop QA \(HotpotQA\(Yanget al\.,[2018](https://arxiv.org/html/2607.16257#bib.bib44)\), 2Wiki\(Hoet al\.,[2020](https://arxiv.org/html/2607.16257#bib.bib45)\), Musique\(Trivediet al\.,[2023](https://arxiv.org/html/2607.16257#bib.bib46)\), and Bamboogle\(Presset al\.,[2023](https://arxiv.org/html/2607.16257#bib.bib47)\)\)\. For TextCraft, we report the average success rate \(%\) for each query\. Full settings details are provided in Appendix[A](https://arxiv.org/html/2607.16257#A1)\.
#### Models and Baselines\.
We use Qwen2\.5\-3B/7B\-Instruct\(Qwenet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib48)\)as our base models\. For HPO, we use Qwen3\-Embedding\-0\.6B as encoder\. We compare our approach against a diverse set of strong baselines\. Specifically, we includePPO, a widely adopted actor\-critic method that relies on an additional value model, as well as critic\-free methods such asGRPO, which estimate advantages over groups of trajectories\. For the SearchQA task, we further evaluate recent competitive models trained with different RL algorithms, includingSearch\-R1,ZeroSearch\. Full training settings and hyperparameter details are provided in Appendix[A](https://arxiv.org/html/2607.16257#A1)\.
### 5\.2Main Result
#### HPO significantly improves policy performance, especially on multi\-hop tasks\.
As shown in Table[1](https://arxiv.org/html/2607.16257#S4.T1), HPO improves average performance by 4%\-7% over other baselines\. Moreover, the improvements are especially pronounced on multi\-hop benchmarks such as HotpotQA and Bamboogle\. We analyze that these gains likely stem from HPO’s fine\-grained, step\-level advantages, which enable more accurate credit assignment by identifying and reinforcing truly valuable actions; this in turn translates into higher final answer accuracy on multi\-hop questions\.
#### Offline HPO achieves better performance than online HPO\.
As shown in Table[1](https://arxiv.org/html/2607.16257#S4.T1), offline HPO is more stable, with smaller performance variance across benchmarks\. Although online HPO can construct an on\-policy hindsight distribution from newly collected rollouts, the number of successful samples used for constructing hindsight distribution fluctuates under stochastic sampling, leading to unstable estimation\. Offline HPO instead constructs the hindsight from a fixed offline dataset, and it further mitigates the issue of sparse rewards: when no correct solution is found, both GRPO and online HPO cannot update policy effectively\. In contrast, offline HPO can still update by constructing the hindsight from a fixed set of successful trajectories in the offline data, thus providing informative step\-level advantages\.
#### HPO consistently improves performance across long\-horizon agent tasks\.
Table[2](https://arxiv.org/html/2607.16257#S4.T2)reports results on the TextCraft benchmark\. HPO achieves consistent improvements in average success rate over all baselines, demonstrating that HPO is effective across different task domains and interaction dynamics, highlighting its robustness and broad applicability to long\-horizon agent settings\.
### 5\.3Training Dynamics
Figure[4](https://arxiv.org/html/2607.16257#S5.F4)shows the training reward dynamics on Qwen2\.5\-7B\. GRPO exhibits unstable long\-horizon training with abrupt reward collapses, as it fails to distinguish successful actions from failed ones within trajectories\. PPO achieves more stable training by introducing a critic to identify valuable actions, but converges more slowly because the critic requires substantial warm\-up to produce reliable value estimates\. In contrast, HPO achieves fast and stable convergence by identifying valuable actions without training an additional critic model\. Notably, offline HPO demonstrates the fastest early\-stage reward growth, which can be attributed to its step\-level advantage in identifying correct actions within failed trajectories, thereby enabling meaningful policy updates even in the absence of fully correct samples\.
Figure 4:The reward dynamics of HPO and other baseline methods with Qwen2\.5\-7B over SearchQA, where curves and uncertainty ranges are computed using the mean and standard deviation over three random seeds, respectively\.GRPO exhibits unstable training, PPO converges more slowly, while HPO achieves fast and stable convergence\.
### 5\.4Interpretability Analysis ofASA\_\{S\}
#### ASA\_\{S\}alone is sufficient to drive effective policy updates\.
To isolate the contribution ofAEA\_\{E\}, we ablate it and re\-run off\-policy HPO over Qwen2\.5\-3B\. In this setting, policy optimization is driven solely by the learning signal fromASA\_\{S\}, while outcome rewards are computed only for monitoring and reporting training dynamics\. As shown in Figure[5](https://arxiv.org/html/2607.16257#S5.F5), optimizing withASA\_\{S\}alone yields stable and effective policy updates: the monitored training reward increases steadily throughout optimization, without exhibiting reward collapse\. The final model is only marginally worse than the standard outcome\-reward–optimized model, suggesting that the bias introduced byASA\_\{S\}is limited and practically acceptable\. Overall, these results highlight the practical utility ofASA\_\{S\}, particularly for open\-ended tasks where outcome rewards are difficult to specify, or for online settings where outcome evaluation is costly, sinceASA\_\{S\}can be learned from offline data that has already been evaluated\. Thecase studyin the Appendix[E\.2](https://arxiv.org/html/2607.16257#A5.SS2)further confirms thatASA\_\{S\}reliably identifies valuable actions and provides sufficiently informative step\-level signals for policy optimization\.
Figure 5:Training reward dynamics of HPO using only step\-level advantages with Qwen2\.5\-3B\. The dark blue curve is EMA\-smoothed from the raw data\.Without episode\-level rewards, training remains stable and effective, indicating that step\-level advantages provide a reliable learning signal\.
### 5\.5Efficiency Analysis
We analyze the computational budget of HPO\. HPO shares the same core architecture as GRPO, including multi\-turn rollouts, computation of old and reference probabilities, and policy updates\. The primary additions introduced by HPO are the step\-relative advantage estimation components\. To evaluate cost, we train a SearchQA LLM agent using Qwen2\.5\-7B\-Instruct with Qwen3\-0\.6B\-Embedding as the encoder, and record a per\-iteration time breakdown\.
Figure 6:Per\-iteration training time breakdown of HPO with Qwen2\.5\-7B\. Blue bars denote components shared with GRPO, while orange bars denote HPO\-specific additions\.The added overhead in HPO is negligible \(<0\.8%<0\.8\\%\)\.As shown in Figure[6](https://arxiv.org/html/2607.16257#S5.F6), the added components introduce negligible overhead\. Compared with GRPO, HPO requires only an additional 2\.03s to compute the step\-level advantage, accounting for approximately 0\.8% of the total time per step, which demonstrates that HPO achieves computational efficiency comparable to that of GRPO\.
## 6Related Work
#### LLMs as decision\-making agents\.
Large language models \(LLMs\) have been increasingly used as autonomous agents for reasoning, planning, and action execution across diverse domains, including code generation\(Zhanget al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib3)\), smart device control\(Zhang and Zhang,[2024](https://arxiv.org/html/2607.16257#bib.bib4);[Guret al\.,](https://arxiv.org/html/2607.16257#bib.bib5); Heet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib6); Honget al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib7)\), and interactive games\([Abdulhaiet al\.,](https://arxiv.org/html/2607.16257#bib.bib8); Prasadet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib9)\)\. Early approaches primarily relied on prompting\-based frameworks such as ReAct\(Yaoet al\.,[2022](https://arxiv.org/html/2607.16257#bib.bib12)\)and Self\-Refine\(Madaanet al\.,[2023](https://arxiv.org/html/2607.16257#bib.bib13)\), which depend on strong proprietary models \(e\.g\., OpenAI o3\) and do not endow models with intrinsic agentic capabilities\. More recent work shifts toward supervised fine\-tuning \(SFT\)\(Chenet al\.,[2023](https://arxiv.org/html/2607.16257#bib.bib15)\)and reinforcement learning \(RL\)\(Zhaiet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib14)\), allowing agents to acquire behaviors directly through data or environment interaction rather than handcrafted prompts\.
#### Reinforcement learning for LLM agents\.
Reinforcement learning has emerged as a key post\-training paradigm for LLMs, supporting both improved reasoning\(Guoet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib16); Jaechet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib17); Penninoet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib18)\)and interactive decision\-making\(Xiet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib19); Peiyuanet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib20)\)\. Common algorithms include PPO\(Schulmanet al\.,[2017](https://arxiv.org/html/2607.16257#bib.bib21)\), GRPO\(Shaoet al\.,[2024](https://arxiv.org/html/2607.16257#bib.bib22)\), and REINFORCE\+\+\(Huet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib23)\)\. However, many representative works \(e\.g\., DeepSeek\-R1\) focus on single\-turn settings and do not apply to long\-horizon interaction\. Although recent studies explore multi\-step RL for agent training\(Caoet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib24); Jinet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib1);[Qiet al\.,](https://arxiv.org/html/2607.16257#bib.bib25); Wanget al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib26)\), they often face challenges in scalability and optimization stability when applied to complex and diverse environments\(Jinet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib1); Xueet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib2); Sunet al\.,[2025](https://arxiv.org/html/2607.16257#bib.bib27)\)\.
## 7Conclusion
This work proposes a novel policy gradient estimator that compares the current policy with the hindsight distribution in an intent space, yielding low\-variance and smooth learning signals that substantially improve the stability and efficiency of long\-horizon training, and offering a new pathway toward stable critic\-free reinforcement learning\. Future work will explore more fine\-grained constructions of the intent space and extend this framework to more complex multimodal and open\-ended environments\.
## Acknowledgments
Thanks for the kind suggestions and support from Huawei Large Model Data Technology Lab\.
## Impact Statement
This paper presents work whose goal is to advance the field of Machine Learning\. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here\.
## References
- \[1\]M\. Abdulhai, I\. White, C\. V\. Snell, C\. Sun, J\. Hong, Y\. Zhai, K\. Xu, and S\. LevineLMRL gym: benchmarks for multi\-turn reinforcement learning with language models\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- D\. Alvarez\-Melis and N\. Fusi \(2021\)Dataset dynamics via gradient flows in probability space\.InInternational conference on machine learning,pp\. 219–230\.Cited by:[§4\.3](https://arxiv.org/html/2607.16257#S4.SS3.p4.2)\.
- L\. Ambrosio, N\. Gigli, and G\. Savare \(2005\)Gradient flows in metric spaces and in the wasserstein space of probability measures\.Lectures in Mathematics, ETH Zurich, Birkhäuser\.Cited by:[§4\.3](https://arxiv.org/html/2607.16257#S4.SS3.p4.2)\.
- S\. Cao, S\. Hegde, D\. Li, T\. Griggs, S\. Liu, E\. Tang, J\. Pan, X\. Wang, A\. Malik, G\. Neubig,et al\.\(2025\)SkyRL\-v0: train real\-world long\-horizon agents via reinforcement learning\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- B\. Chen, C\. Shu, E\. Shareghi, N\. Collier, K\. Narasimhan, and S\. Yao \(2023\)Fireact: toward language agent fine\-tuning\.arXiv preprint arXiv:2310\.05915\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- Y\. Duan, X\. Chen, R\. Houthooft, J\. Schulman, and P\. Abbeel \(2016\)Benchmarking deep reinforcement learning for continuous control\.InInternational conference on machine learning,pp\. 1329–1338\.Cited by:[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.SSS0.Px2.p2.3)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- \[8\]I\. Gur, H\. Furuta, A\. V\. Huang, M\. Safdari, Y\. Matsuo, D\. Eck, and A\. FaustA real\-world webagent with planning, long context understanding, and program synthesis\.InThe Twelfth International Conference on Learning Representations,Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- H\. He, W\. Yao, K\. Ma, W\. Yu, Y\. Dai, H\. Zhang, Z\. Lan, and D\. Yu \(2024\)WebVoyager: building an end\-to\-end web agent with large multimodal models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6864–6890\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- X\. Ho, A\. D\. Nguyen, S\. Sugawara, and A\. Aizawa \(2020\)Constructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps\.arXiv preprint arXiv:2011\.01060\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- W\. Hong, W\. Wang, Q\. Lv, J\. Xu, W\. Yu, J\. Ji, Y\. Wang, Z\. Wang, Y\. Dong, M\. Ding,et al\.\(2024\)Cogagent: a visual language model for gui agents\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 14281–14290\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- J\. Hu, J\. K\. Liu, H\. Xu, and W\. Shen \(2025\)Reinforce\+\+: stabilizing critic\-free policy optimization with global advantage normalization\.arXiv preprint arXiv:2501\.03262\.Cited by:[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- S\. Hu, T\. Huang, G\. Liu, R\. R\. Kompella, F\. Ilhan, S\. F\. Tekin, Y\. Xu, Z\. Yahn, and L\. Liu \(2024\)A survey on large language model\-based game agents\.arXiv preprint arXiv:2404\.02039\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p1.1)\.
- A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney,et al\.\(2024\)Openai o1 system card\.arXiv preprint arXiv:2412\.16720\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- B\. Jin, H\. Zeng, Z\. Yue, J\. Yoon, S\. Arik, D\. Wang, H\. Zamani, and J\. Han \(2025\)Search\-r1: training llms to reason and leverage search engines with reinforcement learning\.arXiv preprint arXiv:2503\.09516\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p1.1),[§1](https://arxiv.org/html/2607.16257#S1.p2.1),[§1](https://arxiv.org/html/2607.16257#S1.p3.1),[§3](https://arxiv.org/html/2607.16257#S3.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- M\. Joshi, E\. Choi, D\. S\. Weld, and L\. Zettlemoyer \(2017\)Triviaqa: a large scale distantly supervised challenge dataset for reading comprehension\.arXiv preprint arXiv:1705\.03551\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- H\. A\. Just, F\. Kang, J\. T\. Wang, Y\. Zeng, M\. Ko, M\. Jin, and R\. Jia \(2023\)Lava: data valuation without pre\-specified learning algorithms\.arXiv preprint arXiv:2305\.00054\.Cited by:[§4\.3](https://arxiv.org/html/2607.16257#S4.SS3.p4.2)\.
- T\. Kwiatkowski, J\. Palomaki, O\. Redfield, M\. Collins, A\. Parikh, C\. Alberti, D\. Epstein, I\. Polosukhin, J\. Devlin, K\. Lee,et al\.\(2019\)Natural questions: a benchmark for question answering research\.Transactions of the Association for Computational Linguistics7,pp\. 453–466\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- Z\. Li, T\. Xu, Y\. Zhang, Z\. Lin, Y\. Yu, R\. Sun, and Z\. Luo \(2024\)ReMax: a simple, effective, and efficient reinforcement learning method for aligning large language models\.InInternational Conference on Machine Learning,pp\. 29128–29163\.Cited by:[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.SSS0.Px2.p2.3)\.
- R\. Luo, L\. Wang, W\. He, L\. Chen, J\. Li, and X\. Xia \(2025\)Gui\-r1: a generalist r1\-style vision\-language action model for gui agents\.arXiv preprint arXiv:2504\.10458\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p1.1)\.
- A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.\(2023\)Self\-refine: iterative refinement with self\-feedback\.Advances in Neural Information Processing Systems36,pp\. 46534–46594\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- A\. Mallen, A\. Asai, V\. Zhong, R\. Das, H\. Hajishirzi, and D\. Khashabi \(2022\)When not to trust language models: investigating effectiveness and limitations of parametric and non\-parametric memories\.arXiv preprint arXiv:2212\.105117\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- F\. Peiyuan, Y\. He, G\. Huang, Y\. Lin, H\. Zhang, Y\. Zhang, and H\. Li \(2024\)Agile: a novel reinforcement learning framework of llm agents\.Advances in Neural Information Processing Systems37,pp\. 5244–5284\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- F\. Pennino, B\. Raimondi, M\. Rondelli, A\. Gurioli, and M\. Gabbrielli \(2025\)From reasoning to code: grpo optimization for underrepresented languages\.arXiv preprint arXiv:2506\.11027\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- A\. Prasad, A\. Koller, M\. Hartmann, P\. Clark, A\. Sabharwal, M\. Bansal, and T\. Khot \(2024\)Adapt: as\-needed decomposition and planning with language models\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 4226–4252\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- O\. Press, M\. Zhang, S\. Min, L\. Schmidt, N\. A\. Smith, and M\. Lewis \(2023\)Measuring and narrowing the compositionality gap in language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 5687–5711\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- \[27\]Z\. Qi, X\. Liu, I\. L\. Iong, H\. Lai, X\. Sun, J\. Sun, X\. Yang, Y\. Yang, S\. Yao, W\. Xu,et al\.WebRL: training llm web agents via self\-evolving online curriculum reinforcement learning\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, and et al\. \(2025\)Qwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px2.p1.1)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.SSS0.Px2.p2.3),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p2.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- H\. Sun, Z\. Qiao, J\. Guo, X\. Fan, Y\. Hou, Y\. Jiang, P\. Xie, Y\. Zhang, F\. Huang, and J\. Zhou \(2025\)Zerosearch: incentivize the search capability of llms without searching\.arXiv preprint arXiv:2505\.04588\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p2.1),[§3](https://arxiv.org/html/2607.16257#S3.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- R\. S\. Sutton, A\. G\. Barto,et al\.\(1998\)Reinforcement learning: an introduction\.Vol\.1,MIT press Cambridge\.Cited by:[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.p7.1)\.
- H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. Sabharwal \(2023\)Interleaving retrieval with chain\-of\-thought reasoning for knowledge\-intensive multi\-step questions\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 10014–10037\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- C\. Villaniet al\.\(2008\)Optimal transport: old and new\.Vol\.338,Springer\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p6.1),[§2\.3](https://arxiv.org/html/2607.16257#S2.SS3.p1.3),[§2\.3](https://arxiv.org/html/2607.16257#S2.SS3.p3.2)\.
- Z\. Wang, K\. Wang, Q\. Wang, P\. Zhang, L\. Li, Z\. Yang, X\. Jin, K\. Yu, M\. N\. Nguyen, L\. Liu,et al\.\(2025\)Ragen: understanding self\-evolution in llm agents via multi\-turn reinforcement learning\.arXiv preprint arXiv:2504\.20073\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- R\. J\. Williams \(1992\)Simple statistical gradient\-following algorithms for connectionist reinforcement learning\.Machine learning8\(3\),pp\. 229–256\.Cited by:[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.p1.1)\.
- C\. Wu, A\. Rajeswaran, Y\. Duan, V\. Kumar, A\. M\. Bayen, S\. Kakade, I\. Mordatch, and P\. Abbeel \(2018\)Variance reduction for policy gradient with action\-dependent factorized baselines\.InInternational Conference on Learning Representations,Cited by:[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.SSS0.Px2.p2.3)\.
- Z\. Xi, J\. Huang, C\. Liao, B\. Huang, H\. Guo, J\. Liu, R\. Zheng, J\. Ye, J\. Zhang, W\. Chen,et al\.\(2025\)Agentgym\-rl: training llm agents for long\-horizon decision making through multi\-turn reinforcement learning\.arXiv preprint arXiv:2509\.08755\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- Z\. Xue, L\. Zheng, Q\. Liu, Y\. Li, X\. Zheng, Z\. MA, and B\. An \(2025\)SimpleTIR: end\-to\-end reinforcement learning for multi\-turn tool\-integrated reasoning\.InNeurIPS 2025 Fourth Workshop on Deep Learning for Code,Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.16257#S3.SS1.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2607.16257#S3.p1.1),[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px2.p1.1)\.
- Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. Manning \(2018\)HotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 2369–2380\.Cited by:[§5\.1](https://arxiv.org/html/2607.16257#S5.SS1.SSS0.Px1.p1.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. Cao \(2022\)React: synergizing reasoning and acting in language models\.InThe eleventh international conference on learning representations,Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- S\. Zhai, H\. Bai, Z\. Lin, J\. Pan, P\. Tong, Y\. Zhou, A\. Suhr, S\. Xie, Y\. LeCun, Y\. Ma,et al\.\(2024\)Fine\-tuning large vision\-language models as decision\-making agents via reinforcement learning\.Advances in neural information processing systems37,pp\. 110935–110971\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- G\. Zhang, H\. Geng, X\. Yu, Z\. Yin, Z\. Zhang, Z\. Tan, H\. Zhou, Z\. Li, X\. Xue, Y\. Li,et al\.\(2025\)The landscape of agentic reinforcement learning for llms: a survey\.arXiv preprint arXiv:2509\.02547\.Cited by:[§1](https://arxiv.org/html/2607.16257#S1.p3.1)\.
- K\. Zhang, J\. Li, G\. Li, X\. Shi, and Z\. Jin \(2024\)CodeAgent: enhancing code generation with tool\-integrated agent systems for real\-world repo\-level coding challenges\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13643–13658\.Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
- Z\. Zhang and A\. Zhang \(2024\)You only look at screens: multimodal chain\-of\-action agents\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 3132–3149\.External Links:[Link](https://aclanthology.org/2024.findings-acl.186/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.186)Cited by:[§6](https://arxiv.org/html/2607.16257#S6.SS0.SSS0.Px1.p1.1)\.
## Appendix
## Appendix AExperimental Details
### A\.1Task Description
#### SearchQA\.
The deep search senario features a search engine–based environment equipped with specialized tools and APIs supporting the interaction with search engines\. These APIs enable agents to dynamically generate search queries during the reasoning process, retrieve relevant information from external sources, and incorporate the retrieved information into subsequent reasoning steps\. This setting allows agents to engage in complex reasoning processes that involve iterative searching and information integration, thereby enhancing their capability to solve intricate problems where external knowledge is essential\.
#### TextCraft\.
TextCraft is a text\-based game environment mirroring Minecraft\. The APIs in TextCraft include crafting, inventory management, and dynamic narrative generation\. These APIs allow agents to execute predefined crafting recipes, manipulate inventory contents, navigate virtual spaces\.
### A\.2Experimental Settings\.
#### Settings for SearchQA\.
The maximum prompt length is 1024 tokens, and the maximum response length is 512 tokens\. The max turn is set to 4\. The learning rate is1×10−61\\times 10^\{\-6\}for the actor\. The reward consists of an outcome reward and a format reward, weighted in a ratio of 8:2\. We set the train batch size to 64 and use a group size of 8\. Rollout and validation temperatures are set to 1\.0 and 0\.0, respectively\. The mini\-batch size is 32, and the KL\-divergence loss coefficient is set to 0\.001\. We use E5 as the retriever\. The weighting coefficientω\\omegais set to0\.50\.5\. For the training data, we use a mixture of NQ and HotpotQA\. In addition, for offline HPO, we use 8 offline correct trajectories per query\. We use Qwen3\-Embedding\-0\.6B as encoder\. During training, we save a model checkpoint every 50 steps; for the main experimental results, we test these checkpoints and report the best testing performance\.
#### Settings for TextCraft\.
The maximum prompt length is 1024 tokens, and the maximum response length is 512 tokens\. The max turn is set to 30\. The learning rate is1×10−61\\times 10^\{\-6\}for the actor\. We adopt a outcome\-based reward, assigning a reward of 1 for success and 0 for failure\. We set the train batch size to 64 and use a group size of 8\. Rollout and validation temperatures are set to 1\.0 and 0\.0, respectively\. The mini\-batch size is 32, and the KL\-divergence loss coefficient is set to 0\.001\. The weighting coefficientω\\omegais set to0\.50\.5\. For offline HPO, we use 8 offline correct trajectories per query\. We use Qwen3\-Embedding\-0\.6B as encoder\. During training, we save a model checkpoint every 50 steps; for the main experimental results, we test these checkpoints and report the best testing performance\.
## Appendix BGeneral Hindsight Distribution
We first revisit the binary\-reward definition used in the main text\. LetRt=∑t′=t∞rt′R\_\{t\}=\\sum\_\{t^\{\\prime\}=t\}^\{\\infty\}r\_\{t^\{\\prime\}\}denote the future return from steptt\. In the binary setting, the hindsight distribution is defined by conditioning on successful future outcomes:
ρπh\(s,a\)=∑t=0∞ℙπ\(\(st,at\)=\(s,a\)∣Rt=1\)\.\\rho\_\{\\pi\}^\{h\}\(s,a\)=\\sum\_\{t=0\}^\{\\infty\}\\mathbb\{P\}\_\{\\pi\}\\\!\\left\(\(s\_\{t\},a\_\{t\}\)=\(s,a\)\\mid R\_\{t\}=1\\right\)\.\(21\)By Bayes’ rule, this can be rewritten as
ρπh\(s,a\)\\displaystyle\\rho\_\{\\pi\}^\{h\}\(s,a\)=1Z∑t=0∞ℙπ\(\(st,at\)=\(s,a\)\)ℙπ\(Rt=1∣st=s,at=a\)\\displaystyle=\\frac\{1\}\{Z\}\\sum\_\{t=0\}^\{\\infty\}\\mathbb\{P\}\_\{\\pi\}\\\!\\left\(\(s\_\{t\},a\_\{t\}\)=\(s,a\)\\right\)\\mathbb\{P\}\_\{\\pi\}\\\!\\left\(R\_\{t\}=1\\mid s\_\{t\}=s,a\_\{t\}=a\\right\)\(22\)=ρπ\(s,a\)Qπ\(s,a\)Z,\\displaystyle=\\frac\{\\rho\_\{\\pi\}\(s,a\)Q\_\{\\pi\}\(s,a\)\}\{Z\},\(23\)whereQπ\(s,a\)=ℙπ\(Rt=1∣st=s,at=a\)Q\_\{\\pi\}\(s,a\)=\\mathbb\{P\}\_\{\\pi\}\(R\_\{t\}=1\\mid s\_\{t\}=s,a\_\{t\}=a\)in the binary\-reward case andZZis the normalizing constant\.
This form suggests a direct extension beyond binary rewards: instead of conditioning on the eventRt=1R\_\{t\}=1, we construct the hindsight distribution by reweighting probability mass according to realized future returns\. Under continuous or dense rewards, the hindsight distribution should therefore take the following form:
ρπh\(s,a\)=1Z∑t=0∞𝔼π\[Rt𝟏\(\(st,at\)=\(s,a\)\)\],\\rho\_\{\\pi\}^\{h\}\(s,a\)=\\frac\{1\}\{Z\}\\sum\_\{t=0\}^\{\\infty\}\\mathbb\{E\}\_\{\\pi\}\\\!\\left\[R\_\{t\}\\mathbf\{1\}\\\!\\left\(\(s\_\{t\},a\_\{t\}\)=\(s,a\)\\right\)\\right\],\(24\)whereZZnormalizes the distribution\. Equivalently, this can be written asρπh\(s,a\)=ρπ\(s,a\)Qπ\(s,a\)/Z\\rho\_\{\\pi\}^\{h\}\(s,a\)=\\rho\_\{\\pi\}\(s,a\)Q\_\{\\pi\}\(s,a\)/Z, whereQπ\(s,a\)=𝔼π\[Rt∣st=s,at=a\]Q\_\{\\pi\}\(s,a\)=\\mathbb\{E\}\_\{\\pi\}\[R\_\{t\}\\mid s\_\{t\}=s,a\_\{t\}=a\]\. If returns can be negative or have different scales, one can use a nonnegative normalized or shifted return as the weighting function\.
In summary, the two cases are unified by the same reweighting view\.Binary rewards \(0/1\):the distribution is obtained via rejection sampling over trajectories, retaining only those with return equal to 1\.Continuous rewards:the distribution is constructed by reweighting probability mass according to trajectory returns\.
## Appendix CEmbedding Model Ablation
HPO constructs intent distributions in a semantic embedding space, so a natural question is whether its step\-level advantage estimates are sensitive to the particular embedding model used to represent state–action pairs\. To examine this, we conduct an ablation over the choice of encoder\. In our main experiments, we useQwen3\-Embedding\-0\.6Bas the default encoder\. We compare it with two alternatives:Qwen3\-Embedding\-4B, which has a larger model scale within the same embedding family, andall\-MiniLM\-L12\-v2, which provides a smaller encoder with a different architecture\.
For each encoder, we recompute the HPO step\-level advantages on the same set of 256 samples\. We then compare the resulting advantage rankings using Kendall’sτ\\tau, which measures whether different encoders assign consistent relative importance to the same steps\. Let m1, m2, and m3 denote Qwen3\-Embedding\-0\.6B, Qwen3\-Embedding\-4B, and all\-MiniLM\-L12\-v2, respectively\. The rank correlations are highly significant across encoder choices: m1–m2 yieldsp<2\.2×10−308p<2\.2\\times 10^\{\-308\}, and m1–m3 yieldsp=6\.8×10−212p=6\.8\\times 10^\{\-212\}\.
These results show that the advantage estimates produced by HPO are strongly aligned across embedding models, including both a larger encoder from the same family and an encoder with a different architecture\. This suggests that HPO mainly relies on stable semantic structure in the embedding space rather than idiosyncrasies of a particular encoder, indicating that the method is not sensitive to the specific embedding model choice\.
## Appendix DProof
### D\.1Proof of Lemma 3\.1
###### Lemma D\.1\(Variance Decomposition\)\.
Consider the REINFORCE gradient estimatorG^b\\hat\{G\}\_\{b\}with the baselinebb,
G^b=g\(a\)\(Q^\(s,a\)−b\(s\)\),\\hat\{G\}\_\{b\}\\;=\\;g\(a\)\\,\\big\(\\hat\{Q\}\(s,a\)\-b\(s\)\\big\),\(25\)wherea∼π\(⋅∣s\)a\\sim\\pi\(\\cdot\\mid s\)andg\(a\)=∇θlogπθ\(a∣s\)g\(a\)=\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\.
Let∥⋅∥\\\|\\cdot\\\|denotes the Euclidean norm\. Then its conditional variance admits the decomposition
Var\(G^b∣s\)=𝔼\[‖g\(a\)‖2∣s\]⋅Var\(Q^\(s,a\)−b\(s\)∣s\)\\displaystyle\\mathrm\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)\\\!=\\\!\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\\\!\\cdot\\\!\\mathrm\{Var\}\\\!\\left\(\\hat\{Q\}\(s,a\)\-b\(s\)\\mid s\\right\)\(26\)
###### Proof\.
Condition on the statessand write
G^b=g\(a\)X,X:=Q^\(s,a\)−b\(s\),a∼π\(⋅∣s\)\.\\hat\{G\}\_\{b\}=g\(a\)\\,X,\\qquad X:=\\hat\{Q\}\(s,a\)\-b\(s\),\\qquad a\\sim\\pi\(\\cdot\\mid s\)\.Under the standard simplifying assumption thatg\(a\)g\(a\)andQ^\(s,a\)\\hat\{Q\}\(s,a\)are \(conditionally\) uncorrelated givenss, we have
Var\(G^b∣s\)=𝔼\[∥g\(a\)∥2X2∣s\]−∥𝔼\[g\(a\)X∣s\]∥2\.\\text\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}X^\{2\}\\mid s\\right\]\-\\left\\\|\\mathbb\{E\}\[g\(a\)X\\mid s\]\\right\\\|^\{2\}\.Using conditional independence,
𝔼\[‖g\(a\)‖2X2∣s\]=𝔼\[‖g\(a\)‖2∣s\]𝔼\[X2∣s\],𝔼\[g\(a\)X∣s\]=𝔼\[g\(a\)∣s\]𝔼\[X∣s\]\.\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}X^\{2\}\\mid s\\right\]=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\\mathbb\{E\}\\\!\\left\[X^\{2\}\\mid s\\right\],\\qquad\\mathbb\{E\}\[g\(a\)X\\mid s\]=\\mathbb\{E\}\[g\(a\)\\mid s\]\\mathbb\{E\}\[X\\mid s\]\.Moreover, for score functions one typically has𝔼\[g\(a\)∣s\]=𝔼\[∇θlogπθ\(a∣s\)∣s\]=0\\mathbb\{E\}\[g\(a\)\\mid s\]=\\mathbb\{E\}\[\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\\mid s\]=0\. Therefore,
Var\(G^b∣s\)=𝔼\[‖g\(a\)‖2∣s\]𝔼\[X2∣s\]\.\\text\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\\mathbb\{E\}\\\!\\left\[X^\{2\}\\mid s\\right\]\.Finally, if the baseline is chosen as the conditional meanb\(s\)=𝔼\[Q^\(s,a\)∣s\]b\(s\)=\\mathbb\{E\}\[\\hat\{Q\}\(s,a\)\\mid s\], then𝔼\[X∣s\]=0\\mathbb\{E\}\[X\\mid s\]=0and thus𝔼\[X2∣s\]=Var\(X∣s\)=Var\(Q^\(s,a\)−b\(s\)∣s\)\\mathbb\{E\}\[X^\{2\}\\mid s\]=\\text\{Var\}\(X\\mid s\)=\\text\{Var\}\(\\hat\{Q\}\(s,a\)\-b\(s\)\\mid s\), yielding
Var\(G^b∣s\)=𝔼\[‖g\(a\)‖2∣s\]Var\(Q^\(s,a\)−b\(s\)∣s\),\\text\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\\text\{Var\}\\\!\\left\(\\hat\{Q\}\(s,a\)\-b\(s\)\\mid s\\right\),∎
### D\.2Proof of Theorem 3\.3
###### Theorem D\.2\(Optimal Baseline and Excess Variance\)\.
Among all scalar baselinesb\(s\)b\(s\)independent of the sampled actionaa, the conditional varianceVar\(G^b∣s\)\\mathrm\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)is minimized by
b∗\(s\)=𝔼\[‖g\(a\)‖2Q^\(s,a\)∣s\]𝔼\[‖g\(a\)‖2∣s\]\.b^\{\*\}\(s\)=\\frac\{\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\hat\{Q\}\(s,a\)\\mid s\\right\]\}\{\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\}\.\(27\)Moreover, for any alternative baselineb~\(s\)\\tilde\{b\}\(s\), the increase in variance admits the exact expression
Var\(G^b~∣s\)\\displaystyle\\mathrm\{Var\}\(\\hat\{G\}\_\{\\tilde\{b\}\}\\mid s\)−Var\(G^b∗∣s\),=\\displaystyle\\,\-\\,\\mathrm\{Var\}\(\\hat\{G\}\_\{b^\{\*\}\}\\mid s\),\\,=\\,𝔼\[‖g\(a\)‖2∣s\]\(b~\(s\)−b∗\(s\)\)2\.\\displaystyle\\mathbb\{E\}\\\!\\left\[\\,\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\(\\tilde\{b\}\(s\)\-b^\{\*\}\(s\)\)^\{2\}\.\(28\)
###### Proof\.
Fixssand writeX=Q^\(s,a\)X=\\hat\{Q\}\(s,a\)\. For any scalar baselineb=b\(s\)b=b\(s\)independent ofaa,
Var\(G^b∣s\)=𝔼\[∥g\(a\)∥2\(X−b\)2∣s\]−∥𝔼\[g\(a\)X∣s\]∥2,\\text\{Var\}\(\\hat\{G\}\_\{b\}\\mid s\)=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-b\)^\{2\}\\mid s\\right\]\-\\left\\\|\\mathbb\{E\}\[g\(a\)X\\mid s\]\\right\\\|^\{2\},where we use𝔼\[g\(a\)∣s\]=0\\mathbb\{E\}\[g\(a\)\\mid s\]=0, so subtracting an action\-independent baseline does not change the conditional mean\. Since the second term does not depend onbb, the variance is minimized by
b∗\(s\)=argminb𝔼\[‖g\(a\)‖2\(X−b\)2∣s\]=𝔼\[‖g\(a\)‖2X∣s\]𝔼\[‖g\(a\)‖2∣s\]\.b^\{\*\}\(s\)=\\arg\\min\_\{b\}\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-b\)^\{2\}\\mid s\\right\]=\\frac\{\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}X\\mid s\\right\]\}\{\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\}\.For any alternativeb~\(s\)\\tilde\{b\}\(s\),
𝔼\[‖g\(a\)‖2\(X−b~\)2∣s\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-\\tilde\{b\}\)^\{2\}\\mid s\\right\]=𝔼\[‖g\(a\)‖2\(X−b∗\+b∗−b~\)2∣s\]\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-b^\{\*\}\+b^\{\*\}\-\\tilde\{b\}\)^\{2\}\\mid s\\right\]=𝔼\[‖g\(a\)‖2\(X−b∗\)2∣s\]−2\(b~−b∗\)𝔼\[‖g\(a\)‖2\(X−b∗\)∣s\]\+𝔼\[‖g\(a\)‖2∣s\]\(b~−b∗\)2\.\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-b^\{\*\}\)^\{2\}\\mid s\\right\]\-2\(\\tilde\{b\}\-b^\{\*\}\)\\,\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-b^\{\*\}\)\\mid s\\right\]\+\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\(\\tilde\{b\}\-b^\{\*\}\)^\{2\}\.Sinceb∗=𝔼\[‖g\(a\)‖2X∣s\]/𝔼\[‖g\(a\)‖2∣s\]b^\{\*\}=\\mathbb\{E\}\[\\\|g\(a\)\\\|^\{2\}X\\mid s\]/\\mathbb\{E\}\[\\\|g\(a\)\\\|^\{2\}\\mid s\], the cross term vanishes,𝔼\[‖g\(a\)‖2\(X−b∗\)∣s\]=0\\mathbb\{E\}\[\\\|g\(a\)\\\|^\{2\}\(X\-b^\{\*\}\)\\mid s\]=0, and therefore
𝔼\[‖g\(a\)‖2\(X−b~\)2∣s\]−𝔼\[‖g\(a\)‖2\(X−b∗\)2∣s\]=𝔼\[‖g\(a\)‖2∣s\]\(b~−b∗\)2\.\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-\\tilde\{b\}\)^\{2\}\\mid s\\right\]\-\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\(X\-b^\{\*\}\)^\{2\}\\mid s\\right\]=\\mathbb\{E\}\\\!\\left\[\\\|g\(a\)\\\|^\{2\}\\mid s\\right\]\(\\tilde\{b\}\-b^\{\*\}\)^\{2\}\.
which immediately gives
Var\(G^b~∣s\)−Var\(G^b∗∣s\)=𝔼\[‖g\(a\)‖2∣s\]\(b~\(s\)−b∗\(s\)\)2\.\\text\{Var\}\(\\hat\{G\}\_\{\\tilde\{b\}\}\\mid s\)\-\\text\{Var\}\(\\hat\{G\}\_\{b^\{\*\}\}\\mid s\)=\\mathbb\{E\}\[\\\|g\(a\)\\\|^\{2\}\\mid s\]\\,\(\\tilde\{b\}\(s\)\-b^\{\*\}\(s\)\)^\{2\}\.∎
### D\.3Proof of Lemma 4\.2
###### Lemma D\.3\.
Consider the policy objective𝒥\(θ\)=𝔼τ∼π\[∑t≥0rt\]\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\\tau\\sim\\pi\}\[\\sum\_\{t\\geq 0\}r\_\{t\}\]\. Its policy gradient admits the following two equivalent forms:
∇θ𝒥\(θ\)\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=𝔼\(s,a\)∼ρπ\[Qπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[Q\_\{\\pi\}\(s,a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(29\)=𝔼\(s,a\)∼ρπ\[−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(30\)where−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)=ρπh\(s,a\)ρπ\(s,a\)\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}=\\frac\{\\rho^\{h\}\_\{\\pi\}\(s,a\)\}\{\\rho\_\{\\pi\}\(s,a\)\}denotes the pointwise variational derivative ofKL\(ρπh\|\|ρπ\)KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)with respect toρπ\(s,a\)\\rho\_\{\\pi\}\(s,a\)\.
###### Proof\.
For𝒥\(θ\)=𝔼τ∼π\[∑t≥0rt\]\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\\tau\\sim\\pi\}\[\\sum\_\{t\\geq 0\}r\_\{t\}\], by the policy gradient theorem, its gradient admits the following equivalent form
∇θ𝒥\(θ\)=𝔼\(s,a\)∼ρπ\[Qπ\(s,a\)∇θlogπθ\(a∣s\)\]\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[Q\_\{\\pi\}\(s,a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(31\)Writing the expectation explicitly, we obtain
∇θ𝒥\(θ\)=∑s,aρπ\(s,a\)Qπ\(s,a\)∇θlogπθ\(a∣s\)\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=\\sum\_\{s,a\}\\rho\_\{\\pi\}\(s,a\)Q\_\{\\pi\}\(s,a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\(32\)LetZ=∑s,aρπ\(s,a\)Qπ\(s,a\)Z=\\sum\_\{s,a\}\\rho\_\{\\pi\}\(s,a\)Q\_\{\\pi\}\(s,a\)\. Defineρπh\(s,a\)=ρπ\(s,a\)Qπ\(s,a\)/Z\\rho\_\{\\pi\}^\{h\}\(s,a\)=\\rho\_\{\\pi\}\(s,a\)Q\_\{\\pi\}\(s,a\)/Z\. Then the above becomes
∇θ𝒥\(θ\)\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=∑s,aρπ\(s,a\)Qπ\(s,a\)∇θlogπθ\(a∣s\)\\displaystyle=\\sum\_\{s,a\}\\rho\_\{\\pi\}\(s,a\)Q\_\{\\pi\}\(s,a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\(33\)=Z∑s,aρπ\(s,a\)ρπh\(s,a\)ρπ\(s,a\)∇θlogπθ\(a∣s\)\\displaystyle=Z\\sum\_\{s,a\}\\rho\_\{\\pi\}\(s,a\)\\frac\{\\rho^\{h\}\_\{\\pi\}\(s,a\)\}\{\\rho\_\{\\pi\}\(s,a\)\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\(34\)=Z𝔼\(s,a\)∼ρπ\[ρπh\(s,a\)ρπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle=Z\\,\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[\\frac\{\\rho^\{h\}\_\{\\pi\}\(s,a\)\}\{\\rho\_\{\\pi\}\(s,a\)\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(35\)where the ratioρπh\(s,a\)ρπ\(s,a\)\\frac\{\\rho^\{h\}\_\{\\pi\}\(s,a\)\}\{\\rho\_\{\\pi\}\(s,a\)\}can be interpreted as the pointwise variational derivative ofKL\(ρπh\|\|ρπ\)KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)with respect toρπ\\rho\_\{\\pi\}, i\.e\.,−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}\.
SinceZZis a positive constant, it does not affect the direction of the gradient\. Therefore, the policy gradient admits the following two equivalent forms:
∇θ𝒥\(θ\)\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{J\}\(\\theta\)=𝔼\(s,a\)∼ρπ\[Qπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[Q\_\{\\pi\}\(s,a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(36\)=𝔼\(s,a\)∼ρπ\[−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)∇θlogπθ\(a∣s\)\]\\displaystyle=\\mathbb\{E\}\_\{\(s,a\)\\sim\\rho\_\{\\pi\}\}\[\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\\mid s\)\]\(37\)∎
### D\.4Proof of Lemma 4\.3
###### Lemma D\.4\.
DefineD=sup\(s,a\),\(s′,a′\)∈ρπd\(\(s,a\),\(s′,a′\)\)D=\\sup\_\{\(s,a\),\(s^\{\\prime\},a^\{\\prime\}\)\\in\\rho\_\{\\pi\}\}d\\big\(\(s,a\),\(s^\{\\prime\},a^\{\\prime\}\)\\big\)\. Letf∗\(s,a\)f^\{\*\}\(s,a\)denote the−δW1\(ρπ\|\|ρπh\)δρπ\(s,a\)\-\\frac\{\\delta W\_\{1\}\(\\rho\_\{\\pi\}\|\|\\rho^\{h\}\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}, andg\(s,a\)g\(s,a\)denote the−δKL\(ρπh\|\|ρπ\)δρπ\(s,a\)\-\\frac\{\\delta KL\(\\rho^\{h\}\_\{\\pi\}\|\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(s,a\)\}, then
Varρπ\(f∗\(s,a\)\)≤D2/4\.\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(f^\{\*\}\(s,a\)\)\\leq D^\{2\}/4\.\(38\)and
Varρπ\(g\(s,a\)\)=∫ρπh\(s,a\)2ρπ\(s,a\)d\(s,a\)−1=χ2\(ρπh∥ρπ\)\.\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(g\(s,a\)\)=\\int\\frac\{\\rho\_\{\\pi\}^\{h\}\(s,a\)^\{2\}\}\{\\rho\_\{\\pi\}\(s,a\)\}\\,d\(s,a\)\-1=\\chi^\{2\}\(\\rho\_\{\\pi\}^\{h\}\\\|\\rho\_\{\\pi\}\)\.\(39\)
###### Proof\.
For theW1W\_\{1\}term, recall the Kantorovich–Rubinstein dual:
W1\(ρπh,ρπ\)=sup‖f‖Lip≤1𝔼x∼ρπh\[f\(x\)\]−𝔼x∼ρπ\[f\(x\)\],x=\(s,a\)\.W\_\{1\}\(\\rho\_\{\\pi\}^\{h\},\\rho\_\{\\pi\}\)=\\sup\_\{\\\|f\\\|\_\{\\mathrm\{Lip\}\}\\leq 1\}\\ \\mathbb\{E\}\_\{x\\sim\\rho\_\{\\pi\}^\{h\}\}\[f\(x\)\]\-\\mathbb\{E\}\_\{x\\sim\\rho\_\{\\pi\}\}\[f\(x\)\],\\quad x=\(s,a\)\.\(40\)
Letf∗f^\{\*\}be the Kantorovich potential, which is11\-Lipschitz on the support ofρπ\\rho\_\{\\pi\}\(up to an additive constant\)\. By definition of the diameterD=supx,x′∈supp\(ρπ\)d\(x,x′\)D=\\sup\_\{x,x^\{\\prime\}\\in\\mathrm\{supp\}\(\\rho\_\{\\pi\}\)\}d\(x,x^\{\\prime\}\), any11\-Lipschitz function satisfies, for allx,x′∈supp\(ρπ\)x,x^\{\\prime\}\\in\\mathrm\{supp\}\(\\rho\_\{\\pi\}\),\|f∗\(x\)−f∗\(x′\)\|≤d\(x,x′\)≤D\|f^\{\*\}\(x\)\-f^\{\*\}\(x^\{\\prime\}\)\|\\leq d\(x,x^\{\\prime\}\)\\leq D,
Therefore, by Popoviciu’s inequality on variances,
Varρπ\(f∗\(x\)\)≤\(supf∗−inff∗\)24≤D24\.\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(f^\{\*\}\(x\)\)\\leq\\frac\{\(\\sup f^\{\*\}\-\\inf f^\{\*\}\)^\{2\}\}\{4\}\\leq\\frac\{D^\{2\}\}\{4\}\.\(41\)
For theKLKLterm, by the stated definition,
g\(x\)=−δKL\(ρπh∥ρπ\)δρπ\(x\)=ρπh\(x\)ρπ\(x\),g\(x\)\\;=\\;\-\\frac\{\\delta KL\(\\rho\_\{\\pi\}^\{h\}\\\|\\rho\_\{\\pi\}\)\}\{\\delta\\rho\_\{\\pi\}\(x\)\}\\;=\\;\\frac\{\\rho\_\{\\pi\}^\{h\}\(x\)\}\{\\rho\_\{\\pi\}\(x\)\},\(42\)
assumingρπ\(x\)\>0\\rho\_\{\\pi\}\(x\)\>0wheneverρπh\(x\)\>0\\rho\_\{\\pi\}^\{h\}\(x\)\>0\. Then
𝔼x∼ρπ\[g\(x\)\]=∫ρπ\(x\)ρπh\(x\)ρπ\(x\)𝑑x=∫ρπh\(x\)𝑑x=1,\\mathbb\{E\}\_\{x\\sim\\rho\_\{\\pi\}\}\[g\(x\)\]=\\int\\rho\_\{\\pi\}\(x\)\\frac\{\\rho\_\{\\pi\}^\{h\}\(x\)\}\{\\rho\_\{\\pi\}\(x\)\}\\,dx=\\int\\rho\_\{\\pi\}^\{h\}\(x\)\\,dx=1,\(43\)and
Varρπ\(g\(x\)\)=𝔼ρπ\[g\(x\)2\]−\(𝔼ρπ\[g\(x\)\]\)2=∫ρπ\(x\)\(ρπh\(x\)ρπ\(x\)\)2𝑑x−1=∫ρπh\(x\)2ρπ\(x\)𝑑x−1\.\\mathrm\{Var\}\_\{\\rho\_\{\\pi\}\}\(g\(x\)\)=\\mathbb\{E\}\_\{\\rho\_\{\\pi\}\}\[g\(x\)^\{2\}\]\-\\Big\(\\mathbb\{E\}\_\{\\rho\_\{\\pi\}\}\[g\(x\)\]\\Big\)^\{2\}=\\int\\rho\_\{\\pi\}\(x\)\\Big\(\\frac\{\\rho\_\{\\pi\}^\{h\}\(x\)\}\{\\rho\_\{\\pi\}\(x\)\}\\Big\)^\{2\}\\,dx\-1=\\int\\frac\{\\rho\_\{\\pi\}^\{h\}\(x\)^\{2\}\}\{\\rho\_\{\\pi\}\(x\)\}\\,dx\-1\.\(44\)Finally, noting that
χ2\(ρπh∥ρπ\)=∫\(ρπh\(x\)−ρπ\(x\)\)2ρπ\(x\)𝑑x=∫ρπh\(x\)2ρπ\(x\)𝑑x−2∫ρπh\(x\)𝑑x\+∫ρπ\(x\)𝑑x=∫ρπh\(x\)2ρπ\(x\)𝑑x−1,\\chi^\{2\}\(\\rho\_\{\\pi\}^\{h\}\\\|\\rho\_\{\\pi\}\)=\\int\\frac\{\(\\rho\_\{\\pi\}^\{h\}\(x\)\-\\rho\_\{\\pi\}\(x\)\)^\{2\}\}\{\\rho\_\{\\pi\}\(x\)\}\\,dx=\\int\\frac\{\\rho\_\{\\pi\}^\{h\}\(x\)^\{2\}\}\{\\rho\_\{\\pi\}\(x\)\}\\,dx\-2\\int\\rho\_\{\\pi\}^\{h\}\(x\)\\,dx\+\\int\\rho\_\{\\pi\}\(x\)\\,dx=\\int\\frac\{\\rho\_\{\\pi\}^\{h\}\(x\)^\{2\}\}\{\\rho\_\{\\pi\}\(x\)\}\\,dx\-1,\(45\)∎
## Appendix EAdditional Analysis
### E\.1Analysis of the weightω\\omega
To further investigate the effect of the step\-level advantage weightω\\omega, multiple repeated HPO\-off experiments were conducted with differentω\\omegasettings\. Figure[7](https://arxiv.org/html/2607.16257#A5.F7)illustrates the training reward trajectories for varying values ofω\\omega\. The results show that introducing a largerω\\omegacan lead to more stable optimization, while this effect diminishes asω\\omegadecreases\. In addition, larger values ofω\\omegayield faster reward improvement in the early training stages but converge to slightly lower final performance\. A plausible explanation is that a largerω\\omegaenables more effective learning from sparse rewards during early training, whereas the increased optimization bias introduced by an overly largeω\\omegamay mildly degrade asymptotic performance\. Overall, these results indicate that adjustingω\\omegaprovides a practical mechanism for balancing bias and variance during training\.
Figure 7:Training reward dynamics of Qwen2\.5\-3B\-Instruct with differentω\\omegasettings\.
### E\.2Case Study
We further conduct a qualitative analysis ofASA\_\{S\}across representative samples\. We find thatASA\_\{S\}exhibits strong step\-level discrimination: it consistently assigns higher scores to informative and goal\-progressing actions, while down\-weighting failed, redundant, or uninformative steps\. This behavior suggests thatASA\_\{S\}captures meaningful local learning signals and also explains why optimizing withASA\_\{S\}alone can still drive effective policy updates\.
Example 1User:In what year was the earlier of these two documentaries produced: The Life and Times of Hank Greenberg or Burden of Dreams?#### Assistant: <think\><\\text\{think\}\>The question is asking about the production year of two documentaries, and we need to find out which one is earlier and then provide the year of that documentary\.</think\></\\text\{think\}\><search\><\\text\{search\}\>which documentary was produced earlier, The Life and Times of Hank Greenberg or Burden of Dreams</search\></\\text\{search\}\> \(𝐀𝐬=\+0\.424\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.424\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… The Life and Times of Hank Greenberg is a 1998 documentary film written, directed, and produced by Aviva Kempner about Hank Greenberg, first baseman of the Detroit Tigers, who was inducted into the Baseball Hall of Fame\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that The Life and Times of Hank Greenberg is a 1998 documentary\. Now let’s find out about the production year of Burden of Dreams\.</think\></\\text\{think\}\><search\><\\text\{search\}\>when was Burden of Dreams produced</search\></\\text\{search\}\> \(𝐀𝐬=\+0\.218\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.218\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… Burden of Dreams is a 1982 “making\-of” documentary film directed by Les Blank, shot during and about the chaotic production of Werner Herzog’s film Fitzcarraldo\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\. \(𝐀𝐬=−0\.472\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{1,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{1,0,0\}\\mathbf\{\-0\.472\}\}\) #### User: Your previous action is invalid\. If you want to search, you should put the query between<search\><\\text\{search\}\>and</search\></\\text\{search\}\>\. If you want to give the final answer, you should put the answer between<answer\><\\text\{answer\}\>and</answer\></\\text\{answer\}\>\. Please try again\. #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\.</think\></\\text\{think\}\><answer\><\\text\{answer\}\>1982</answer\></\\text\{answer\}\> \(𝐀𝐬=\+0\.377\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.377\}\}\) #### User: Congratulations\! You have answered the question correctly\!\!\!
Example 2User:Bone carving was used in the casket now on display in what institution?#### Assistant: <think\><\\text\{think\}\>I need to find out which institution has the casket with bone carving that is now on display\.</think\></\\text\{think\}\><search\><\\text\{search\}\>Bone carving was used in the casket now on display in what institution?</search\></\\text\{search\}\> \(𝐀𝐬=\+1\.20\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+1\.20\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… ivory casket … now in the Walters Art Museum … Doc 2… a small Anglo\-Saxon whale’s bone … chest … densely decorated with knife\-cut … scenes … now in theBritish Museum\. … Doc 3… </information\></\\text\{information\}\> #### Assistant: <think\><\\text\{think\}\>From the search results, it is clear that the casket with bone carving now on display is in the Walters Art Museum, Baltimore, Maryland\.</think\></\\text\{think\}\><answer\><\\text\{answer\}\>Walters Art Museum</answer\></\\text\{answer\}\> \(𝐀𝐬=−0\.574\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{1,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{1,0,0\}\\mathbf\{\-0\.574\}\}\) #### User: Sorry, your answer is incorrect\.
Example 1: the advantage of GRPOUser:In what year was the earlier of these two documentaries produced: The Life and Times of Hank Greenberg or Burden of Dreams?#### Assistant: <think\><\\text\{think\}\>The question is asking about the production year of two documentaries, and we need to find out which one is earlier and then provide the year of that documentary\.</think\></\\text\{think\}\><search\><\\text\{search\}\>which documentary was produced earlier, The Life and Times of Hank Greenberg or Burden of Dreams</search\></\\text\{search\}\> \(𝐀𝐬=\+1\.29\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+1\.29\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… The Life and Times of Hank Greenberg is a 1998 documentary film written, directed, and produced by Aviva Kempner about Hank Greenberg, first baseman of the Detroit Tigers, who was inducted into the Baseball Hall of Fame\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that The Life and Times of Hank Greenberg is a 1998 documentary\. Now let’s find out about the production year of Burden of Dreams\.</think\></\\text\{think\}\><search\><\\text\{search\}\>when was Burden of Dreams produced</search\></\\text\{search\}\> \(𝐀𝐬=\+1\.29\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+1\.29\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… Burden of Dreams is a 1982 “making\-of” documentary film directed by Les Blank, shot during and about the chaotic production of Werner Herzog’s film Fitzcarraldo\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\. \(𝐀𝐬=\+1\.29\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+1\.29\}\}\) #### User: Your previous action is invalid\. If you want to search, you should put the query between<search\><\\text\{search\}\>and</search\></\\text\{search\}\>\. If you want to give the final answer, you should put the answer between<answer\><\\text\{answer\}\>and</answer\></\\text\{answer\}\>\. Please try again\. #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\.</think\></\\text\{think\}\><answer\><\\text\{answer\}\>1982</answer\></\\text\{answer\}\> \(𝐀𝐬=\+1\.29\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+1\.29\}\}\) #### User: Congratulations\! You have answered the question correctly\!\!\!
Example 2: the advantage of PPOUser:In what year was the earlier of these two documentaries produced: The Life and Times of Hank Greenberg or Burden of Dreams?#### Assistant: <think\><\\text\{think\}\>The question is asking about the production year of two documentaries, and we need to find out which one is earlier and then provide the year of that documentary\.</think\></\\text\{think\}\><search\><\\text\{search\}\>which documentary was produced earlier, The Life and Times of Hank Greenberg or Burden of Dreams</search\></\\text\{search\}\> \(𝐀𝐬=\+0\.821\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.821\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… The Life and Times of Hank Greenberg is a 1998 documentary film written, directed, and produced by Aviva Kempner about Hank Greenberg, first baseman of the Detroit Tigers, who was inducted into the Baseball Hall of Fame\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that The Life and Times of Hank Greenberg is a 1998 documentary\. Now let’s find out about the production year of Burden of Dreams\.</think\></\\text\{think\}\><search\><\\text\{search\}\>when was Burden of Dreams produced</search\></\\text\{search\}\> \(𝐀𝐬=\+0\.762\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.762\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… Burden of Dreams is a 1982 “making\-of” documentary film directed by Les Blank, shot during and about the chaotic production of Werner Herzog’s film Fitzcarraldo\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\. \(𝐀𝐬=\+0\.684\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.684\}\}\) #### User: Your previous action is invalid\. If you want to search, you should put the query between<search\><\\text\{search\}\>and</search\></\\text\{search\}\>\. If you want to give the final answer, you should put the answer between<answer\><\\text\{answer\}\>and</answer\></\\text\{answer\}\>\. Please try again\. #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\.</think\></\\text\{think\}\><answer\><\\text\{answer\}\>1982</answer\></\\text\{answer\}\> \(𝐀𝐬=\+0\.707\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.707\}\}\) #### User: Congratulations\! You have answered the question correctly\!\!\!
Example 3: the advantage of HPOUser:In what year was the earlier of these two documentaries produced: The Life and Times of Hank Greenberg or Burden of Dreams?#### Assistant: <think\><\\text\{think\}\>The question is asking about the production year of two documentaries, and we need to find out which one is earlier and then provide the year of that documentary\.</think\></\\text\{think\}\><search\><\\text\{search\}\>which documentary was produced earlier, The Life and Times of Hank Greenberg or Burden of Dreams</search\></\\text\{search\}\> \(𝐀𝐬=\+0\.424\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.424\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… The Life and Times of Hank Greenberg is a 1998 documentary film written, directed, and produced by Aviva Kempner about Hank Greenberg, first baseman of the Detroit Tigers, who was inducted into the Baseball Hall of Fame\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that The Life and Times of Hank Greenberg is a 1998 documentary\. Now let’s find out about the production year of Burden of Dreams\.</think\></\\text\{think\}\><search\><\\text\{search\}\>when was Burden of Dreams produced</search\></\\text\{search\}\> \(𝐀𝐬=\+0\.218\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.218\}\}\) #### User: <information\><\\text\{information\}\> Doc 1… Burden of Dreams is a 1982 “making\-of” documentary film directed by Les Blank, shot during and about the chaotic production of Werner Herzog’s film Fitzcarraldo\. … Doc 2… Doc 3… </information\></\\text\{information\}\> #### Assistant: From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\. \(𝐀𝐬=−0\.472\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{1,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{1,0,0\}\\mathbf\{\-0\.472\}\}\) #### User: Your previous action is invalid\. If you want to search, you should put the query between<search\><\\text\{search\}\>and</search\></\\text\{search\}\>\. If you want to give the final answer, you should put the answer between<answer\><\\text\{answer\}\>and</answer\></\\text\{answer\}\>\. Please try again\. #### Assistant: <think\><\\text\{think\}\>From the information provided, it is clear that Burden of Dreams was produced in 1982\. Comparing the production years of the two documentaries, The Life and Times of Hank Greenberg \(1998\) and Burden of Dreams \(1982\), it is clear that Burden of Dreams was produced earlier\. Therefore, the answer is 1982\.</think\></\\text\{think\}\><answer\><\\text\{answer\}\>1982</answer\></\\text\{answer\}\> \(𝐀𝐬=\+0\.377\\mathbf\{A\_\{s\}\}=\{\\color\[rgb\]\{0,0\.5,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.5,0\}\\mathbf\{\+0\.377\}\}\) #### User: Congratulations\! You have answered the question correctly\!\!\!Similar Articles
Gradient Extrapolation-Based Policy Optimization
The article introduces Gradient Extrapolation-Based Policy Optimization (GXPO), a method that approximates multi-step lookahead in RL training for LLMs using only three backward passes. It demonstrates improved reasoning performance on math benchmarks over standard GRPO while maintaining fixed active-phase costs.
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Introduces Surrogate Latent Policy Optimization (SLPO) to apply outcome-reward RL to autoregressive latent reasoners, enabling test-time scaling and variable-horizon policies that improve accuracy on harder instances.
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
PPO-HSC introduces a High-order Sampling Coverage reward to encourage exploration of diverse reasoning patterns in RL fine-tuning of LLMs, improving solution diversity and state-space coverage on math and code tasks.
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning
StepPO introduces a step-centric paradigm for agentic reinforcement learning that aligns policy optimization with agent decision granularity, outperforming token-centric methods in multi-turn interaction tasks.
Learning More from Less: Reinforcement Learning from Hindsight
Introduces Learning from Hindsight (LfH), a method that applies hindsight relabeling to RL post-training of vision-language-action models. By relabeling failed robot rollouts with the tasks they actually achieved, LfH achieves 5x improvement in sample efficiency on out-of-distribution manipulation tasks.