面向目标条件强化学习的多时间尺度学习方法

arXiv cs.AI 论文

摘要

论文提出了 Generalized Implicit Temporal Abstraction(GITA),这是一种用于离线目标条件强化学习的方法,在多个时间抽象尺度上训练单个 k 条件价值函数,从而解决了长程价值信号与局部分辨率之间的权衡问题。GITA 在 OGBench 上超越了 HIQL 和 OTA 等离线 GCRL 基线方法,相比 HIQL 将平均成功率提升了 25 个百分点。

arXiv:2610.00849v1 Announce Type: new Abstract: Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
查看原文
查看缓存全文

缓存时间: 2026/10/02 09:47

# Learning Multiple Timescales forGoal-Conditioned Reinforcement Learning
Source: [https://arxiv.org/html/2610.00849](https://arxiv.org/html/2610.00849)
Pedro Robles Dutenhefner1Dikshant Shehmar2,3Wagner Meira Jr\.1Marlos C\. Machado2,3,4 1Universidade Federal de Minas Gerais,2University of Alberta 3Alberta Machine Intelligence Institute \(Amii\),4Canada CIFAR AI Chair

###### Abstract

Existing approaches to offline goal\-conditioned reinforcement learning \(GCRL\) struggle with long\-horizon tasks\. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states\. Temporal abstraction—treatingkkenvironment steps as a single transition—restores this signal at long range, but no single fixedkksuits all state\-goal distances: largekkpreserves value differences across long temporal distances while collapsing distinctions between nearby states, and smallkkdoes the reverse\. We make this trade\-off explicit and introduce Generalized Implicit Temporal Abstraction \(GITA\), which conditions a single value function onkk\. GITA trains one policy by aggregating advantage\-weighted supervision across multiplekkvalues, so scales assigning larger positive advantages to a state\-goal pair contribute more strongly to its update\. GITA does not need to choose between local resolution and long\-range signal; it retains both without committing to a singlekk\. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points \(73% relative improvement\) over HIQL\. It also improves over the strongest fixed\-kkmethod, OTA, by 7 percentage points \(14% relative\)\.111Code is available at[https://github\.com/pedroroblesduten/GITA](https://github.com/pedroroblesduten/GITA)\.

## 1Introduction

Offline goal\-conditioned reinforcement learning \(GCRL\) aims to learn policies that reach arbitrary goals using only previously collected data\([Levine et al\., 2020](https://arxiv.org/html/2610.00849#bib.bib11)\)\. They typically do so by learning a goal\-conditioned value function\([Schaul et al\., 2015](https://arxiv.org/html/2610.00849#bib.bib1)\)\. In long\-horizon tasks, discounting shrinks the value difference between neighboring states as the goal moves farther away, and once that difference falls below the function approximator error, the learned values can no longer rank actions correctly\([Opryshko et al\., 2026](https://arxiv.org/html/2610.00849#bib.bib10)\)\.

Temporal abstraction\([Sutton et al\., 1999](https://arxiv.org/html/2610.00849#bib.bib8)\)mitigates this issue by treating multiple primitive steps as a single transition\. Even the simplest instantiation, fixing a stride ofkksteps, shortens the value\-propagation chain for a given state\-goal distance by a factor ofkk\. However, a largekkthat restores signal for distant goals also collapses distinctions between nearby states, as nearby states whose remaining distances fall within the same abstract\-horizon bin become indistinguishable under the abstract transition\. The rightkk, therefore, depends on goal distance, and no fixed choice serves every state\-goal pair\.

![Refer to caption](https://arxiv.org/html/2610.00849v1/diagrama_bonitao.png)Figure 1:Overview of the proposed multi\-scale temporal abstraction framework\.Left: a shared value function conditioned onkklearns multiple abstraction scales\.Right: scale\-specific advantages set the relative contribution of eachkkvalue, inducing an implicit soft assignment when training the shared high\-level policy\.In this paper, we introduce Generalized Implicit Temporal Abstraction \(GITA\), a GCRL method that learns a singlekk\-conditioned value function spanning multiple temporal abstraction scales\. From this value function, GITA derives an advantage at every scale and trains one policy on their aggregate via Advantage\-Weighted Regression\([Peng et al\., 2019](https://arxiv.org/html/2610.00849#bib.bib12), AWR;\)\. Because AWR weights each sample by its exponentiated advantage, scales assigning larger positive advantages at a given state\-goal pair contribute more strongly to the update\. GITA therefore combines multiple effective horizons*per state\-goal pair*with no explicit selection mechanism—the sense in which the abstraction is implicit\. It lifts the constraint that every fixed\-scale method inherits by having nearby pairs resolved at smallkk, distant ones at largekk, within one value function\.

Empirically, we characterize howkkgoverns value discrimination and demonstrate GITA’s effectiveness on OGBench\([Park et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib6)\)\. In control environments where we can cleanly vary the distance between states, we show that distinguishing the values of two states at distanceddrequireskkto grow withdd: below a threshold, the values of all sufficiently distant states collapse together\. Conversely, a large enoughkksaturates the values of states closer than a threshold, collapsing those instead\. No singlekkescapes both failures, which is precisely the trade\-off GITA is designed to resolve\. On OGBench, GITA outperforms existing hierarchical baselines, including HIQL\([Park et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib7)\)and OTA\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9)\), as well as a broad range of other GCRL methods across multiple morphologies, datasets, and environment types\. In aggregate, GITA improves over HIQL by 25 percentage points \(p\.p\.\) of success rate \(73% relative\) and over OTA by 7 p\.p\. \(14% relative\)\.

## 2Preliminaries

We consider the offline GCRL setting\([Kaelbling, 1993](https://arxiv.org/html/2610.00849#bib.bib15);[Levine et al\., 2020](https://arxiv.org/html/2610.00849#bib.bib11)\)defined over a Markov Decision Process \(MDP\),ℳ=⟨𝒮,𝒜,P,ρ0,r,γ,𝒢⟩\\mathcal\{M\}=\\langle\\mathcal\{S\},\\mathcal\{A\},P,\\rho\_\{0\},r,\\gamma,\\mathcal\{G\}\\rangle\. Here,𝒮\\mathcal\{S\}and𝒜\\mathcal\{A\}denote the state and action spaces respectively,P⁡\(s′∣s,a\)P\(s^\{\\prime\}\\mid s,a\)is the transition dynamics,ρ0\\rho\_\{0\}is the initial\-state distribution, andγ∈\[0,1\)\\gamma\\in\[0,1\)is the discount factor\. The goal space𝒢⊆𝒮\\mathcal\{G\}\\subseteq\\mathcal\{S\}allows any valid state to serve as a goal\([Schaul et al\., 2015](https://arxiv.org/html/2610.00849#bib.bib1);[Andrychowicz et al\., 2017](https://arxiv.org/html/2610.00849#bib.bib2)\), andr\(s,g\)=−𝟙\[s≠g\]r\(s,g\)=\-\\mathbbm\{1\}\[s\\neq g\]is a goal\-conditioned reward function\. We are given an offline dataset𝒟\\mathcal\{D\}of unlabeled trajectoriesτ=\(s0,a0,s1,…,sT\)\\tau=\(s\_\{0\},a\_\{0\},s\_\{1\},\\ldots,s\_\{T\}\)collected by an unknown behavior policy, and no further environment interaction is permitted\. Writingpπ​\(τ∣g\)=ρ0​\(s0\)​∏tπ⁡\(at∣st,g\)​P​\(st\+1∣st,at\)p^\{\\pi\}\(\\tau\\mid g\)=\\rho\_\{0\}\(s\_\{0\}\)\\prod\_\{t\}\\pi\(a\_\{t\}\\mid s\_\{t\},g\)P\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\}\)for the trajectory distribution induced byπ\\pi, we want to learn a goal\-conditioned policy that maximizesJ\(π\)=𝔼g∼p𝒢,τ∼pπ\(⋅∣g\)\[∑t=0Hγtr\(st,g\)\],J\(\\pi\)=\\mathbb\{E\}\_\{g\\sim p\_\{\\mathcal\{G\}\},\\,\\tau\\sim p^\{\\pi\}\(\\cdot\\mid g\)\}\\left\[\\sum\_\{t=0\}^\{H\}\\gamma^\{t\}r\(s\_\{t\},g\)\\right\],wherep𝒢p\_\{\\mathcal\{G\}\}andHHare the evaluation goal distribution and horizon\.

Hierarchical approaches, the main class of methods we study, facilitate long\-horizon goal\-conditioned control by introducing intermediate subgoals between the current state and the final goal\. HIQL\([Park et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib7)\)uses a high\-level policyπh​\(st\+m∣st,g\)\\pi^\{h\}\(s\_\{t\+m\}\\mid s\_\{t\},g\)to predict a subgoalst\+ms\_\{t\+m\}at offsetmm, and a low\-level policyπℓ​\(at∣st,st\+m\)\\pi^\{\\ell\}\(a\_\{t\}\\mid s\_\{t\},s\_\{t\+m\}\)to select primitive actions toward it\. Both policies rely on a goal\-conditioned value function trained with the expectile temporal\-difference objective:

ℒV=𝔼\(st,st\+1\)∼𝒟,g∼p𝒢​\[L2ν​\(r⁡\(st,g\)\+γ​V¯​\(st\+1,g\)−V⁡\(st,g\)\)\],\\mathcal\{L\}\_\{V\}=\\mathbb\{E\}\_\{\(s\_\{t\},s\_\{t\+1\}\)\\sim\\mathcal\{D\},\\,g\\sim p\_\{\\mathcal\{G\}\}\}\\left\[L\_\{2\}^\{\\nu\}\\left\(r\(s\_\{t\},g\)\+\\gamma\\bar\{V\}\(s\_\{t\+1\},g\)\-V\(s\_\{t\},g\)\\right\)\\right\],\(1\)wherest\+1s\_\{t\+1\}denotes the immediate successor state ofsts\_\{t\}in the dataset,VVis the goal\-conditioned state\-value function,V¯\\bar\{V\}is a target network, andL2ν​\(u\)=\|ν−𝟙​\(u<0\)\|​u2L\_\{2\}^\{\\nu\}\(u\)=\|\\nu\-\\mathbbm\{1\}\(u<0\)\|\\,u^\{2\}, withν\>0\.5\\nu\>0\.5, denotes the expectile loss\.

HIQL extracts both policies with AWR\([Peng et al\., 2019](https://arxiv.org/html/2610.00849#bib.bib12)\)\. The high\- and low\-level advantages are

Ah​\(st,st\+m,g\)=V⁡\(st\+m,g\)−V⁡\(st,g\),Aℓ​\(st,st\+1,st\+m\)=V⁡\(st\+1,st\+m\)−V⁡\(st,st\+m\),\\begin\{gathered\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g\)=V\(s\_\{t\+m\},g\)\-V\(s\_\{t\},g\),\\\\ A^\{\\ell\}\(s\_\{t\},s\_\{t\+1\},s\_\{t\+m\}\)=V\(s\_\{t\+1\},s\_\{t\+m\}\)\-V\(s\_\{t\},s\_\{t\+m\}\),\\end\{gathered\}\(2\)
giving the policy objectives

J⁡\(πh\)=𝔼𝒟,g∼p𝒢​\[exp⁡\(βh​Ah\)​log​πh​\(st\+m∣st,g\)\],J⁡\(πℓ\)=𝔼𝒟​\[exp⁡\(βℓ​Aℓ\)​log​πℓ​\(at∣st,st\+m\)\],\\begin\{gathered\}J\(\\pi^\{h\}\)=\\mathbb\{E\}\_\{\\mathcal\{D\},\\,g\\sim p\_\{\\mathcal\{G\}\}\}\\left\[\\exp\\\!\\left\(\\beta\_\{h\}A^\{h\}\\right\)\\log\\pi^\{h\}\(s\_\{t\+m\}\\mid s\_\{t\},g\)\\right\],\\\\ J\(\\pi^\{\\ell\}\)=\\mathbb\{E\}\_\{\\mathcal\{D\}\}\\left\[\\exp\\\!\\left\(\\beta\_\{\\ell\}A^\{\\ell\}\\right\)\\log\\pi^\{\\ell\}\(a\_\{t\}\\mid s\_\{t\},s\_\{t\+m\}\)\\right\],\\end\{gathered\}\(3\)whereβh,βℓ\>0\\beta\_\{h\},\\beta\_\{\\ell\}\>0are inverse temperatures, andAℓA^\{\\ell\}denotes the low\-level advantage\.

The high\-level policy is trained from the advantageAh=V⁡\(st\+m,g\)−V⁡\(st,g\)A^\{h\}=V\(s\_\{t\+m\},g\)\-V\(s\_\{t\},g\), which should be positive when the candidate subgoalst\+ms\_\{t\+m\}represents progress towardgg\. Recovering this ordering requires the value function to resolve a difference between two states that are onlymmsteps apart\. Asggmoves farther away, one\-step TD learning must propagate information over correspondingly longer horizons, and repeated discounting shrinks that difference until it is dominated by estimation error\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9)\)\. The advantage then carries no reliable signal about which subgoals make progress, and high\-level policy learning degrades\.

Option\-aware temporal abstraction\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9), OTA;\)mitigates this issue by replacing the one\-step value updates in Eq\. \([1](https://arxiv.org/html/2610.00849#S2.E1)\) with temporally\-extended transitions\. For a fixed abstraction factorkk, OTA looks aheadkkprimitive steps along the trajectory in𝒟\\mathcal\{D\}containingsts\_\{t\}and usesst\+ks\_\{t\+k\}as the successor state\. Rather than forming akk\-step TD target with accumulated rewards and discountγk\\gamma^\{k\}, OTA treats this temporally\-extended transition as a single update with TD errorδk=r⁡\(st\+k,g\)\+γ​V¯kh​\(st\+k,g\)−Vkh​\(st,g\)\\delta\_\{k\}=r\(s\_\{t\+k\},g\)\+\\gamma\\bar\{V\}^\{h\}\_\{k\}\(s\_\{t\+k\},g\)\-V^\{h\}\_\{k\}\(s\_\{t\},g\), whereVkhV^\{h\}\_\{k\}denotes the temporally\-extended value function\. Each update therefore propagates value information acrosskkprimitive steps while applying a single discountγ\\gamma, reducing the number of backups needed to connectsts\_\{t\}toggfromd⁡\(st,g\)d\(s\_\{t\},g\)to roughly⌈d⁡\(st,g\)/k⌉\\lceil d\(s\_\{t\},g\)/k\\rceil, whered⁡\(st,g\)d\(s\_\{t\},g\)denotes the number of primitive steps fromsts\_\{t\}togg\. Although a similar contraction can be induced by rescaling the discount asγ~=γ1/k\\widetilde\{\\gamma\}=\\gamma^\{1/k\},[Ahn et al\. \(2025\)](https://arxiv.org/html/2610.00849#bib.bib9)show that this alone does not reproduce the gains of temporally\-extended value updates\.

However, OTA fixes a singlekkfor all state\-goal pairs while learningVkhV^\{h\}\_\{k\}\. This makeskka problem\-dependent hyperparameter that must be tuned per environment, as[Ahn et al\. \(2025\)](https://arxiv.org/html/2610.00849#bib.bib9)do\. Moreover, since the number of backups separatingsts\_\{t\}fromggscales as⌈d⁡\(st,g\)/k⌉\\lceil d\(s\_\{t\},g\)/k\\rceil, thekkthat keeps this quantity small for distant goals is far larger than necessary for nearby ones, and states withinkksteps become indistinguishable underVkhV^\{h\}\_\{k\}\.

We have focused here on the methods GITA builds on most directly\. Appendix[A](https://arxiv.org/html/2610.00849#A1)situates GITA within the broader offline GCRL and temporal abstraction literature\.

## 3No Fixed Abstraction Suits All State\-Goal Distances

We characterize how the temporal abstraction factor affects the high\-level advantage used for subgoal selection\. Letd=d⋆​\(s,g\)d=d^\{\\star\}\(s,g\)denote the minimum number of primitive steps needed to go from a statessto a goalgg, and lets′s^\{\\prime\}be a candidate subgoalΔ\\Deltasteps closer to the goal, such thatd⋆​\(s′,g\)=d⋆​\(s,g\)−Δd^\{\\star\}\(s^\{\\prime\},g\)=d^\{\\star\}\(s,g\)\-\\Delta\. For an abstraction factork\>0k\>0,kkprimitive transitions are represented by one abstract transition, which incurs reward−1\-1and a single discountγ\\gammaunder our analytical convention\. Relaxing the abstract horizon⌈d⋆​\(s,g\)/k⌉\\lceil d^\{\\star\}\(s,g\)/k\\rceiltod⋆​\(s,g\)/kd^\{\\star\}\(s,g\)/kso that it varies smoothly withkk, as a continuous surrogate for the discrete\-horizon problem, yields the temporally\-abstract value function

V~k​\(s,g\)=−1−γd⋆​\(s,g\)/k1−γ\.\\widetilde\{V\}\_\{k\}\(s,g\)=\-\\frac\{1\-\\gamma^\{d^\{\\star\}\(s,g\)/k\}\}\{1\-\\gamma\}\.\(4\)Throughout, we assumed\>Δ\>0d\>\\Delta\>0,k\>0k\>0, and0<γ<10<\\gamma<1\. Specifically, Proposition[5](https://arxiv.org/html/2610.00849#S3.E5)formalizes the relationship between the distance between different states and the difference between their induced value functions across different temporal scaleskk\. Proofs can be found in[AppendixB](https://arxiv.org/html/2610.00849#A2)\.

###### Proposition 1\.

Letssands′s^\{\\prime\}denote two states such thats′s^\{\\prime\}isΔ\\Deltaprimitive steps closer to the goal thanss\. When computed with a fixed stride ofkksteps, the advantage of a subgoals′s^\{\\prime\}in relation tossis

A~k​\(s,s′,g\)=V~k​\(s′,g\)−V~k​\(s,g\)=γ\(d−Δ\)/k⏟reachability⋅1−γΔ/k1−γ⏟resolution\.\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=\\widetilde\{V\}\_\{k\}\(s^\{\\prime\},g\)\-\\widetilde\{V\}\_\{k\}\(s,g\)=\\underbrace\{\\gamma^\{\(d\-\\Delta\)/k\}\}\_\{\\text\{reachability\}\}\\;\\cdot\\;\\underbrace\{\\frac\{1\-\\gamma^\{\\Delta/k\}\}\{1\-\\gamma\}\}\_\{\\text\{resolution\}\}\.\(5\)

This proposition formalizes the intuition we have provided above\. A coarser abstraction allows the advantage to better discriminate distant states when considering distant goals, because the*reachability factor*increases toward11askkgrows\. However, the same coarser abstraction blurs the distinction between nearby states, since the*resolution factor*decreases toward00askkgrows\. We formalize these results in the corollary below\.

###### Corollary 1\.1\.

When computed with a fixed stride ofkksteps, the advantage of a subgoals′s^\{\\prime\}in relation tossvanishes either when the value ofkkis too small or too large\. Formally:

limk→0\+A~k​\(s,s′,g\)=0,limk→∞A~k​\(s,s′,g\)=0\.\\lim\_\{k\\to 0^\{\+\}\}\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=0,\\qquad\\lim\_\{k\\to\\infty\}\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=0\.\(6\)

Becaused⋆​\(s,g\)d^\{\\star\}\(s,g\), the minimum number of primitive steps needed to go from a statessto a goalggvaries depending onss, Proposition[5](https://arxiv.org/html/2610.00849#S3.E5)also allows us to conclude that no singlekkcan be equally effective in distinguishing between different pairs of states\. We formalize this below\.

###### Corollary 1\.2\.

Letd=d⋆​\(s,g\)d=d^\{\\star\}\(s,g\), andΔ=d⋆​\(s,g\)−d⋆​\(s′,g\)\\Delta=d^\{\\star\}\(s,g\)\-d^\{\\star\}\(s^\{\\prime\},g\)\. For a states′s^\{\\prime\}whereΔ≪d\\Delta\\ll d, the optimal abstraction factor in the continuous relaxed model,kopt⋆​\(d,Δ\)k\_\{\\mathrm\{opt\}\}^\{\\star\}\(d,\\Delta\), grows with the state\-goal distance:

kopt⋆​\(d,Δ\)≈−d​log⁡γ\.k\_\{\\mathrm\{opt\}\}^\{\\star\}\(d,\\Delta\)\\approx\-d\\log\\gamma\.\(7\)This choice keeps the remaining state\-goal distance approximately within one effective discount horizon:

d−Δkopt⋆≈−1log⁡γ≈11−γ,γ→1\.\\frac\{d\-\\Delta\}\{k\_\{\\mathrm\{opt\}\}^\{\\star\}\}\\approx\-\\frac\{1\}\{\\log\\gamma\}\\approx\\frac\{1\}\{1\-\\gamma\},\\qquad\\gamma\\to 1\.\(8\)

These predictions hold empirically\. On a subset of OGBench environments, we train independent value functionsVk​\(s,g\)V\_\{k\}\(s,g\), one perkk, and evaluate them on the\(s,s′,g\)\(s,s^\{\\prime\},g\)tuples used to train the high\-level actor\. Since subgoal selection depends only on the ordering induced byAk​\(s,s′,g\)=Vk​\(s′,g\)−Vk​\(s,g\)A\_\{k\}\(s,s^\{\\prime\},g\)=V\_\{k\}\(s^\{\\prime\},g\)\-V\_\{k\}\(s,g\), we measure the Spearman correlation betweenAkA\_\{k\}andd⁡\(s,g\)−d⁡\(s′,g\)d\(s,g\)\-d\(s^\{\\prime\},g\), which we call the Advantage\-Progress Correlation \(APC\)\. We use geodesic shortest paths\([Dijkstra, 1959](https://arxiv.org/html/2610.00849#bib.bib14)\)as a proxy ford⋆d^\{\\star\}\. See[AppendixC](https://arxiv.org/html/2610.00849#A3)for more details\.

Figure 2:No single abstraction factor is best at all distances\.APC for different abstraction factorskkas the goal becomes more distant \(mean±\\pmstd over four seeds\)\. The strip beneath each panel indicates the factor with the highest correlation in each distance regime; the best factor consistently grows with distance, but its range differs across environments\. Inset shows the learned valueVkh​\(s,g\)V^\{h\}\_\{k\}\(s,g\)as a function of geodesic distance\.Figure[2](https://arxiv.org/html/2610.00849#S3.F2)shows a clear distance\-dependent ordering across abstraction factors\. Small factors achieve high APC at short state\-goal distances but degrade rapidly as distance grows, consistent with the collapse of the reachability factor: inpointmaze,k=1k=1falls toAPC≈0\\text\{APC\}\\approx 0byd≈0\.4d\\approx 0\.4, whilek=8k=8still retainsAPC≈0\.5\\text\{APC\}\\approx 0\.5at twice that distance\. Larger factors bring distant pairs back into a value\-sensitive range but attain lower APC at short distances, consistent with the loss of the resolution factor in Proposition[5](https://arxiv.org/html/2610.00849#S3.E5)\. No factor dominates across the distance range; instead, the best\-performing factor increases monotonically with state\-goal distance in every environment, qualitatively consistent with the growth ofkopt⋆k^\{\\star\}\_\{\\mathrm\{opt\}\}withddpredicted by Corollary[8](https://arxiv.org/html/2610.00849#S3.E8)\. That range itself differs across environments—the best factor peaks atk=3k=3inantmazebutk=13k=13inhumanoidmaze—so a single fixed choice cannot serve all environments\.

## 4Generalized Implicit Temporal Abstraction

Section[3](https://arxiv.org/html/2610.00849#S3)established that, for locally separated states, the best abstraction factor grows with state\-goal distance, so any fixedkkis suboptimal for most pairs\.GeneralizedImplicitTemporalAbstraction \(GITA\) removes the need to commit to any onekk\. One value function conditioned onkkrepresents all scales at once, and a single high\-level policy is trained from the advantages this value function induces at every scale, weighted so that the scales best suited to each state\-goal pair dominate its update\. An overview of this mechanism is shown in Figure[1](https://arxiv.org/html/2610.00849#S1.F1)\. The low\-level policy is unchanged from HIQL, predicting primitive actions toward the subgoal the high\-level policy proposes\.

### 4\.1One Network, Multiple Abstraction Scales

GITA introduces the abstraction factorkkas an additional input to the high\-level value network, learning a single parameter\-sharedVθh​\(s,g,k\)V\_\{\\theta\}^\{h\}\(s,g;k\)wherekkis the number of primitive steps represented by each abstract transition\. This makes the abstraction scale a controllable input rather than a fixed architectural choice:kkis sampled during training, andVθhV\_\{\\theta\}^\{h\}can be queried at any scale afterwards\. From here we writekkas an explicit argument rather than a subscript to emphasize the conditioning\.

Given a trajectoryτ=\(s0,a0,s1,…,sT\)\\tau=\(s\_\{0\},a\_\{0\},s\_\{1\},\\ldots,s\_\{T\}\), letsts\_\{t\}denote the current state,g∼p𝒟g\\sim p\_\{\\mathcal\{D\}\}a sampled goal relabeled either from a future state ofτ\\tauor from a random state in𝒟\\mathcal\{D\}, andst\+ks\_\{t\+k\}the state reachedkkprimitive steps aftersts\_\{t\}alongτ\\tau, or the first state at which the goal is reached or the trajectory terminates, whichever comes first\. For each training example, GITA independently samples an abstraction factor from a geometric distribution,

k∼Geom⁡\(1−α\),p⁡\(k\)=\(1−α\)​αk−1,k≥1,k\\sim\\operatorname\{Geom\}\(1\-\\alpha\),\\qquad p\(k\)=\(1\-\\alpha\)\\alpha^\{\\,k\-1\},\\quad k\\geq 1,\(9\)whereα∈\(0,1\)\\alpha\\in\(0,1\)sets the mean scale𝔼⁡\[k\]=1/\(1−α\)\\mathbb\{E\}\[k\]=1/\(1\-\\alpha\)\. Geometric sampling favors shorter scales while retaining a long tail over larger ones, covering both fine and coarse temporal resolutions\.

GITA trains the abstraction\-conditioned value function with the temporal\-difference objective

ℒVθh=𝔼st∼𝒟,g∼p𝒟,k∼p⁡\(k\)​\[L2ν​\(r⁡\(st\+k,g\)\+γ​V¯θ¯h​\(st\+k,g,k\)−Vθh​\(st,g,k\)\)\],\\mathcal\{L\}\_\{V\_\{\\theta\}^\{h\}\}=\\mathbb\{E\}\_\{s\_\{t\}\\sim\\mathcal\{D\},\\,g\\sim p\_\{\\mathcal\{D\}\},\\,k\\sim p\(k\)\}\\left\[L\_\{2\}^\{\\nu\}\\left\(r\(s\_\{t\+k\},g\)\+\\gamma\\bar\{V\}^\{h\}\_\{\\bar\{\\theta\}\}\(s\_\{t\+k\},g;k\)\-V\_\{\\theta\}^\{h\}\(s\_\{t\},g;k\)\\right\)\\right\],\(10\)whereV¯θ¯h\\bar\{V\}^\{h\}\_\{\\bar\{\\theta\}\}is a target network and the successor is evaluated at the samekkthat generated it, so each scale is trained against its own bootstrapped target\. Since all scales are represented by a shared network, they are not learned independently, allowing information to generalize across scales\.

![Refer to caption](https://arxiv.org/html/2610.00849v1/plot_tesselation_value_per_stage_task3.png)Figure 3:Varyingkkdirectly controls the temporal abstraction scale\.Normalized value functionV⁡\(s,g,k\)V\(s,g;k\)onpointmaze\-large\-navigate, evaluated at every navigable cell fork∈\{1,2,3,5,13,21\}k\\in\\\{1,2,3,5,13,21\\\}\. The star marks the fixed reference state\. Askkincreases, the value\-sensitive region expands over longer distances while local resolution progressively degrades\.
### 4\.2High\-Level Policy Training with Multi\-Scale Advantages

The high\-level policy is trained by evaluating each training tuple under several temporal scales at once\. We fix a finite set𝒦\\mathcal\{K\}of abstraction factors as a hyperparameter and evaluate every tuple\(st,st\+m,g\)\(s\_\{t\},s\_\{t\+m\},g\)under allk∈𝒦k\\in\\mathcal\{K\}, wheremmis the subgoal offset in primitive steps and is held fixed independently ofkk\. This is done so the candidate subgoal is the same, only the scale at which its advantage is measured varies\. The policy itself is*not*conditioned on the abstraction factor, giving a singleπϕh​\(st\+m∣st,g\)\\pi\_\{\\phi\}^\{h\}\(s\_\{t\+m\}\\mid s\_\{t\},g\)shared across all scales\.

For every tuple and everyk∈𝒦k\\in\\mathcal\{K\}, we compute the abstraction\-conditioned high\-level advantage

Ah​\(st,st\+m,g,k\)=Vθh​\(st\+m,g,k\)−Vθh​\(st,g,k\)\.A^\{h\}\(s\_\{t\},s\_\{t\+m\},g;k\)=V\_\{\\theta\}^\{h\}\(s\_\{t\+m\},g;k\)\-V\_\{\\theta\}^\{h\}\(s\_\{t\},g;k\)\.\(11\)Each abstraction scale therefore contributes a distinct evaluation of the same observed subgoal\. We then aggregate the corresponding Advantage\-Weighted Regression \(AWR\) weights across scales,

W⁡\(st,st\+m,g\)=1\|𝒦\|​∑k∈𝒦exp⁡\(βh​Ah​\(st,st\+m,g,k\)\),W\(s\_\{t\},s\_\{t\+m\},g\)=\\frac\{1\}\{\|\\mathcal\{K\}\|\}\\sum\_\{k\\in\\mathcal\{K\}\}\\exp\\\!\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g;k\)\\right\),\(12\)and train the high\-level policy according to

ℒπϕh=−𝔼st,st\+m∼𝒟,g∼p𝒟​\[W⁡\(st,st\+m,g\)​log⁡πϕh​\(st\+m∣st,g\)\]\.\\mathcal\{L\}\_\{\\pi\_\{\\phi\}^\{h\}\}=\-\\mathbb\{E\}\_\{s\_\{t\},s\_\{t\+m\}\\sim\\mathcal\{D\},g\\sim p\_\{\\mathcal\{D\}\}\}\\left\[W\(s\_\{t\},s\_\{t\+m\},g\)\\log\\pi\_\{\\phi\}^\{h\}\(s\_\{t\+m\}\\mid s\_\{t\},g\)\\right\]\.\(13\)
![Refer to caption](https://arxiv.org/html/2610.00849v1/implicit.png)Figure 4:The abstraction scale receiving the strongest policy weight changes with state\-goal distance\.Clipped scale\-specific AWR weightmin⁡\{exp⁡\(βh​Ah​\(st,st\+m,g,k\)\),100\}\\min\\\{\\exp\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g;k\)\),100\\\}across geodesic state\-goal distance bins inantmaze\-giant\-navigate, for eachkk\.Applying the exponential before averaging makesWWa soft aggregation of scale\-specific AWR weights\. Each termexp⁡\(βh​Ah\)\\exp\(\\beta\_\{h\}A^\{h\}\)is a weight in its own right, so scales distorted by saturation or excessive abstraction contribute weights nearexp⁡\(0\)\\exp\(0\), while the scale with the clearest positive advantage drives the update\. By Proposition[5](https://arxiv.org/html/2610.00849#S3.E5), that scale is the one nearestkopt⋆k\_\{\\mathrm\{opt\}\}^\{\\star\}, since factors below it are suppressed by the reachability factor and factors above it by the resolution factor\. This is where the*implicit*in GITA comes from:[Section3](https://arxiv.org/html/2610.00849#S3)prescribes an abstraction factor that grows withd⋆​\(s,g\)d^\{\\star\}\(s,g\), butd⋆d^\{\\star\}is unavailable, and GITA never estimates it\. Every scale in𝒦\\mathcal\{K\}is placed on equal footing and the weighting resolves the competition on its own\.

Figure[4](https://arxiv.org/html/2610.00849#S4.F4)shows the clipped scale\-specific AWR weight against state\-goal geodesic distance inantmaze\-giant\-navigate\. At short distances, smaller abstraction factors receive the largest weights, but their advantage signal decays rapidly as the goal moves farther away, causing their weights to approach the neutral valueexp⁡\(0\)=1\\exp\(0\)=1\. Larger factors retain non\-neutral weights over longer distances and become dominant at longer distances\. This mirrors the distance\-dependent ordering in Figure[2](https://arxiv.org/html/2610.00849#S3.F2), which is what makes the implicit assignment work, as the scale that receives the strongest weight in a given distance regime is also the scale that ranks subgoals most reliably there\. See[AppendixI](https://arxiv.org/html/2610.00849#A9)for additional environments\.

At inference, the policy is executed directly asst\+m∼πϕh\(⋅∣st,g\)s\_\{t\+m\}\\sim\\pi\_\{\\phi\}^\{h\}\(\\cdot\\mid s\_\{t\},g\)and needs no abstraction factor as input\. The low\-level policy training objective is unchanged, with the policy predicting a primitive action conditioned on the current state and the generated subgoal,at∼πψℓ\(⋅∣st,st\+m\)a\_\{t\}\\sim\\pi\_\{\\psi\}^\{\\ell\}\(\\cdot\\mid s\_\{t\},s\_\{t\+m\}\)\.

## 5Experiments

We evaluate GITA on the locomotion and manipulation tasks of the Offline Goal Conditioned RL Benchmark\([Park et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib6), OGBench;\)suite, which includes long\-horizon navigation tasks and datasets that require stitching together sub\-trajectories to reach distant goals\. We first describe our evaluation protocol, then compare GITA against representative offline GCRL approaches, including the best fixed\-kkalternative\.

### 5\.1Evaluation on OGBench

#### Experimental Setup\.

We consider three locomotion tasks: PointMaze, AntMaze, and HumanoidMaze, controlling a 2\-DoF ball, an 8\-DoF ant, and a 21\-DoF humanoid, respectively\. Each comes with layouts of increasing geodesic distance—medium,large, andgiant—and thenavigateandstitchdatasets; AntMaze additionally includesexplore\. We also consider two manipulation environments, Cube and Scene, where a robotic arm must reach a target configuration\. The locomotion tasks probe long\-horizon navigation, while the manipulation tasks probe compositional generalization\. We use both the state\- and pixel\-based variants\. An episode terminates when the agent comes within a task\-specific proximity threshold of the goal state, which defines success, or when the environment’s time limit is reached\. Following the OGBench evaluation protocol, we report the average success rate across five goal\-reaching tasks per environment, with5050evaluation episodes per task, averaged over88random seeds for state\-based tasks and44for pixel\-based tasks\.

We compare against the offline GCRL baselines provided by OGBench\([Park et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib6)\): goal\-conditioned behavior cloning\([Ghosh et al\., 2021](https://arxiv.org/html/2610.00849#bib.bib3), GCBC;\), goal\-conditioned implicit V\- and Q\-learning\([Kostrikov et al\., 2022](https://arxiv.org/html/2610.00849#bib.bib13);[Park et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib7), GCIVL and GCIQL;\), quasimetric RL\([Wang et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib5), QRL;\), contrastive RL\([Eysenbach et al\., 2022](https://arxiv.org/html/2610.00849#bib.bib4), CRL;\), hierarchical implicit Q\-learning\([Park et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib7), HIQL;\), and fixed\-scale option\-aware temporally abstracted value learning\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9), OTA;\)\. The comparison with OTA is the most informative, as it directly contrasts GITA’s multi\-scale formulation with fixed\-scale temporal abstraction\. HIQL abstracts only the policy, leaving its value function atk=1k=1, so the three together trace the progression from no value\-level abstraction to fixed\-scale to multi\-scale\. Note also that OTA tuneskkper environment, whereas GITA uses a single candidate set𝒦\\mathcal\{K\}throughout, avoiding per\-environment tuning of the candidate set\. See[AppendixG](https://arxiv.org/html/2610.00849#A7)for further details about the baselines considered\.

Table 1:GITA outperforms fixed\-scale and non\-abstracted baselines across OGBench tasks\.Success rate \(%\), mean±\\pmstandard deviation over 8 seeds \(4 for pixel\-based tasks\)\. The best mean, and any result within one standard deviation of it, are inbold\.Non\-HierarchicalHierarchicalEnvironmentTypeSizeGCBCGCIVLGCIQLQRLCRLHIQLOTAGITAPointMazenavigatelarge2929±6\\pm 64545±5\\pm 53434±3\\pm 3𝟖𝟔\\mathbf\{86\}±𝟗\\mathbf\{\\pm 9\}3939±7\\pm 75858±5\\pm 58585±5\\pm 5𝟗𝟑\\mathbf\{93\}±𝟐\\mathbf\{\\pm 2\}navigategiant11±2\\pm 200±0\\pm 000±0\\pm 06868±7\\pm 72727±10\\pm 104646±9\\pm 97272±6\\pm 6𝟖𝟎\\mathbf\{80\}±𝟑\\mathbf\{\\pm 3\}stitchlarge77±5\\pm 51212±6\\pm 63131±2\\pm 2𝟖𝟒\\mathbf\{84\}±𝟏𝟓\\mathbf\{\\pm 15\}00±0\\pm 01313±6\\pm 64646±7\\pm 7𝟕𝟐\\mathbf\{72\}±𝟏𝟑\\mathbf\{\\pm 13\}stitchgiant00±0\\pm 000±0\\pm 000±0\\pm 05050±8\\pm 800±0\\pm 000±0\\pm 04444±8\\pm 8𝟔𝟎\\mathbf\{60\}±𝟕\\mathbf\{\\pm 7\}AntMazenavigatelarge2424±2\\pm 21616±5\\pm 53434±4\\pm 47575±6\\pm 68383±4\\pm 4𝟗𝟏\\mathbf\{91\}±𝟐\\mathbf\{\\pm 2\}𝟗𝟏\\mathbf\{91\}±𝟏\\mathbf\{\\pm 1\}𝟗𝟏\\mathbf\{91\}±𝟏\\mathbf\{\\pm 1\}navigategiant00±0\\pm 000±0\\pm 000±0\\pm 01414±3\\pm 31616±3\\pm 36565±5\\pm 57070±2\\pm 2𝟕𝟖\\mathbf\{78\}±𝟐\\mathbf\{\\pm 2\}stitchlarge33±3\\pm 31818±2\\pm 277±2\\pm 21818±2\\pm 21111±2\\pm 26767±5\\pm 57979±3\\pm 3𝟖𝟔\\mathbf\{86\}±𝟐\\mathbf\{\\pm 2\}stitchgiant00±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 022±2\\pm 22929±5\\pm 5𝟑𝟖\\mathbf\{38\}±𝟑\\mathbf\{\\pm 3\}explorelarge00±0\\pm 01010±3\\pm 300±0\\pm 000±0\\pm 000±0\\pm 044±5\\pm 5𝟔𝟐\\mathbf\{62\}±𝟏𝟐\\mathbf\{\\pm 12\}𝟕𝟐\\mathbf\{72\}±𝟏𝟑\\mathbf\{\\pm 13\}HumanoidMazenavigatelarge11±0\\pm 022±1\\pm 122±1\\pm 155±1\\pm 12424±4\\pm 44949±4\\pm 4𝟖𝟐\\mathbf\{82\}±𝟐\\mathbf\{\\pm 2\}7373±3\\pm 3navigategiant00±0\\pm 000±0\\pm 000±0\\pm 011±0\\pm 033±2\\pm 21212±4\\pm 4𝟗𝟏\\mathbf\{91\}±𝟏\\mathbf\{\\pm 1\}𝟖𝟗\\mathbf\{89\}±𝟐\\mathbf\{\\pm 2\}stitchlarge66±3\\pm 311±1\\pm 100±0\\pm 033±1\\pm 144±1\\pm 12828±3\\pm 34343±3\\pm 3𝟓𝟏\\mathbf\{51\}±𝟒\\mathbf\{\\pm 4\}stitchgiant00±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 033±2\\pm 2𝟔𝟏\\mathbf\{61\}±𝟑\\mathbf\{\\pm 3\}𝟔𝟑\\mathbf\{63\}±𝟑\\mathbf\{\\pm 3\}Visual\-AntMazenavigatelarge44±0\\pm 055±1\\pm 144±1\\pm 100±0\\pm 0𝟖𝟒\\mathbf\{84\}±𝟏\\mathbf\{\\pm 1\}5353±9\\pm 96363±4\\pm 47575±6\\pm 6navigategiant00±0\\pm 011±1\\pm 100±0\\pm 000±0\\pm 0𝟒𝟕\\mathbf\{47\}±𝟐\\mathbf\{\\pm 2\}66±4\\pm 488±3\\pm 32121±5\\pm 5stitchlarge2424±3\\pm 311±1\\pm 100±0\\pm 011±1\\pm 11111±3\\pm 32828±2\\pm 22121±2\\pm 2𝟒𝟓\\mathbf\{45\}±𝟖\\mathbf\{\\pm 8\}stitchgiant00±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 000±0\\pm 022±2\\pm 2𝟕\\mathbf\{7\}±𝟑\\mathbf\{\\pm 3\}Cubeplaysingle66±2\\pm 25353±4\\pm 4𝟔𝟖\\mathbf\{68\}±𝟔\\mathbf\{\\pm 6\}55±1\\pm 11919±2\\pm 21515±3\\pm 31212±2\\pm 21515±1\\pm 1playdouble11±1\\pm 13636±3\\pm 3𝟒𝟎\\mathbf\{40\}±𝟓\\mathbf\{\\pm 5\}11±0\\pm 01010±2\\pm 266±2\\pm 233±1\\pm 133±1\\pm 1Sceneplay–55±1\\pm 14242±4\\pm 4𝟓𝟏\\mathbf\{51\}±𝟒\\mathbf\{\\pm 4\}55±1\\pm 11919±2\\pm 23838±3\\pm 32222±5\\pm 52626±6\\pm 6Visual\-Cubenoisysingle1414±3\\pm 37575±3\\pm 34848±3\\pm 31010±5\\pm 53939±30\\pm 309999±0\\pm 09999±0\\pm 0𝟏𝟎𝟎\\mathbf\{100\}±𝟎\\mathbf\{\\pm 0\}noisydouble55±1\\pm 11717±4\\pm 42222±2\\pm 266±2\\pm 266±3\\pm 35959±3\\pm 36565±2\\pm 2𝟕𝟓\\mathbf\{75\}±𝟏\\mathbf\{\\pm 1\}Visual\-Scenenoisy–1313±2\\pm 22323±2\\pm 21212±4\\pm 422±0\\pm 01515±2\\pm 25050±1\\pm 15454±2\\pm 2𝟔𝟏\\mathbf\{61\}±𝟏\\mathbf\{\\pm 1\}
#### Performance on OGBench\.

Table[1](https://arxiv.org/html/2610.00849#S5.T1)reports success rates across all tasks considered\. Aggregated over tasks, GITA outperforms every baseline under a two\-sided paired Wilcoxon signed\-rank test with Holm\-Bonferroni correction for multiple comparisons \(all adjustedp<0\.05p<0\.05; see Appendix[F](https://arxiv.org/html/2610.00849#A6)\)\. The locomotion results trace a clear progression fromHIQLtoOTAtoGITA: abstracting the value function at a fixed scale improves over no abstraction, and conditioning on the scale improves further still\. The second gap is the one that supports our central claim, that no single temporal abstraction suits all state\-goal distances\. The gap remains substantial in thegiantlayouts, where the wider range of state\-goal distances within a single environment places strong pressure on methods committed to a single scale\.GITAalso improves on mostlargelayouts, where a fixed scale is already miscalibrated for part of the distance range\.

The exception is state\-based manipulation\. On Cube and Scene, the non\-hierarchicalGCIQLoutperformsGITAby a wide margin, but it also outperformsHIQLandOTA, so the gap reflects a broader difficulty of hierarchical methods, rather than of multi\-scale abstraction specifically\. OGBench attributes this to theplaydatasets, whose open\-loop, non\-Markovian trajectories with temporally correlated noise can make policy and subgoal prediction ambiguous\.GITAstill improves overOTAon every pixel\-based manipulation task, where thenoisydatasets are Markovian and provide broader state coverage, yielding more consistent supervision for hierarchical policy learning\.

Table 2:Every variant degrades sharply somewhere; GITA does not\.Success rate \(%\) on the maze tasks, mean±\\pmstandard deviation over 8 seeds\.

### 5\.2Ablation Studies

Table[2](https://arxiv.org/html/2610.00849#S5.T2)compares four variants of GITA, each isolating one component: whether the value function is conditioned onkk, how scale\-specific advantages are combined, and how the scales are chosen\.

Unconditioned Valuetests whether exposure to multiple temporal scales during value learning suffices on its own, or whether the value function must explicitly represent the abstraction scale\. To test this, we removekkfrom the input to the high\-level value function while still samplingk∼Geom⁡\(1−α\)k\\sim\\operatorname\{Geom\}\(1\-\\alpha\)to construct the abstract transitions in[eq\.10](https://arxiv.org/html/2610.00849#S4.E10), so the network learns one value function averaged across scales\. Only one advantage is then defined per tuple, so the aggregation in[eq\.12](https://arxiv.org/html/2610.00849#S4.E12)reduces to standard AWR; this variant therefore removes both components at once and serves as the cumulative comparison against GITA\.

Highest Advantage Factortests whether policy supervision is better treated as a hard selection problem, with one scale chosen per sample, than as an aggregation across scales\. We test this by retaining the abstraction\-conditioned value function but weighting each tuple using only the scale with the largest advantage,W=exp⁡\(βh​maxk∈𝒦​Akh\)W=\\exp\(\\beta\_\{h\}\\max\_\{k\\in\\mathcal\{K\}\}A^\{h\}\_\{k\}\), discarding the rest\.

Mean Advantage Before Exponentiationtests whether the implicit selection induced by exponentiating before averaging matters, or whether exposing the policy to multiple scales suffices\. We isolate this effect by keeping all scale\-specific advantages but reversing the order, replacing[eq\.12](https://arxiv.org/html/2610.00849#S4.E12)withexp⁡\(βh​1\|𝒦\|​∑kAh\)\\exp\(\\beta\_\{h\}\\frac\{1\}\{\|\\mathcal\{K\}\|\}\\sum\_\{k\}A^\{h\}\)\. Because exponential is convex, this removes the emphasis on large positive scale\-specific advantages and instead averages all advantages uniformly before exponentiation\.

GITA with Geometrick\\bm\{k\}tests whether the high\-level policy needs a pre\-specified candidate set of abstraction factors\. We replace the fixed𝒦\\mathcal\{K\}with\|𝒦\|\|\\mathcal\{K\}\|factors drawn independently fromGeom⁡\(1−ζ\)\\operatorname\{Geom\}\(1\-\\zeta\)at every policy update\. This removes the spacing of𝒦\\mathcal\{K\}as a design choice and ensures the policy is supervised only at scales supported during value learning, though the candidate set is now noisier at each update\.

We observe the largest drop in performance in*Unconditioned Value*\. Exposure to multiple temporal scales during value learning is therefore not sufficient on its own; the scale must be explicitly represented in the value function\.GITAalso significantly outperforms*Highest Advantage Factor*\(p<0\.05p<0\.05, Holm\-Bonferroni\-corrected Wilcoxon signed\-rank test; see[Table7](https://arxiv.org/html/2610.00849#A6.T7)\), so committing to a single scale per sample loses information that soft aggregation retains, even when that scale is the strongest one available\.*Mean Advantage Before Exponentiation*and*GITA with Geometrickk*are not significantly different overall, but each collapses on one of AntMazelarge\-exploreandgiant\-stitchrespectively, the two settings with the widest spread of state\-goal distances, whereGITAdoes not\. Neither the exponentiation order nor the fixed candidate set is thus strictly necessary, but together they are what make the method reliable across regimes\.

## 6Conclusion

In this paper, we introduced Generalized Implicit Temporal Abstraction \(GITA\), an offline goal\-conditioned reinforcement learning \(GCRL\) method that addresses the trade\-off between long\-range value propagation and local value discrimination\. We showed analytically that the advantage used for subgoal selection factors into a reachability term and a resolution term pulling in opposite directions as the abstraction factor grows, so the best factor increases with state\-goal distance, and no fixed choice serves every pair\. GITA resolves this by learning a shared value function across temporal scales and training a single policy from their aggregated advantages, implicitly emphasizing the most informative scales for each state\-goal pair\. Our empirical results establish GITA as a robust and flexible framework for hierarchical offline RL, providing a foundation for future methods that leverage multiple temporal scales to tackle long\-horizon tasks\.

Directions for extending GITA include: \(1\) GITA, like other hierarchical methods, is less effective on Cube and Scene, which are short\-horizon and compositional rather than long\-horizon; incorporating compositionality into the notion of temporal abstraction is a natural next step\. \(2\) In GITA, as in many other approaches, the notion of distance depends on the policy that generated the dataset, and correcting for discrepancies introduced by the off\-policy nature of the data is a promising direction\. \(3\) Finally, the representations GITA learns are only implicitly shaped by the value loss; explicitly learning representations that capture an intrinsic notion of distance\([Shehmar et al\., 2026](https://arxiv.org/html/2610.00849#bib.bib17), e\.g\.,\)would likely benefit the framework\.

### Acknowledgments

This research was supported in part by the Natural Sciences and Engineering Research Council of Canada \(NSERC\), the Canada CIFAR AI Chair Program, and Alberta Innovates\. Computational resources were provided in part by the Digital Research Alliance of Canada\. This work was also partially funded by CNPq, CAPES, FAPEMIG, IAIA\-INCT on AI, INCTTildIAR, and the Brazilian Center on Algorithm Transparency and AI Safety\.

## References

- V\. C\. Adaikkappan, D\. Meger, S\. Rajeswar, and P\. MazzagliaMulti\-scale predictive representations for goal\-conditioned reinforcement learning\.CoRRabs/2605\.09364\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px3.p1.1)\.
- Ahnet al\.\(2025\)H\. Ahn, H\. Choi, J\. Han, and T\. MoonOption\-aware temporally abstracted value for offline goal\-conditioned reinforcement learning\.InNeural Information Processing Systems \(NeurIPS\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px3.p1.1),[Appendix B](https://arxiv.org/html/2610.00849#A2.p1.1),[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p2.1),[§1](https://arxiv.org/html/2610.00849#S1.p4.1),[§2](https://arxiv.org/html/2610.00849#S2.p5.1),[§2](https://arxiv.org/html/2610.00849#S2.p6.1),[§2](https://arxiv.org/html/2610.00849#S2.p7.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1)\.
- Andrychowiczet al\.\(2017\)M\. Andrychowicz, F\. Wolski, A\. Ray, J\. Schneider, R\. Fong, P\. Welinder, B\. McGrew, J\. Tobin, P\. Abbeel, and W\. ZarembaHindsight experience replay\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2610.00849#S2.p1.1)\.
- Baet al\.\(2016\)J\. L\. Ba, J\. R\. Kiros, and G\. E\. HintonLayer normalization\.CoRRabs/1607\.06450\.Cited by:[Appendix E](https://arxiv.org/html/2610.00849#A5.SS0.SSS0.Px1.p1.1)\.
- Choiet al\.\(2026\)J\. Choi, S\. Lee, and S\. SeoChain\-of\-goals hierarchical policy for long\-horizon offline goal\-conditioned RL\.InInternational Conference on Machine Learning \(ICML\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px2.p1.1)\.
- Chunget al\.\(2026\)H\. Chung, J\. Lee, and S\. OhOffline reinforcement learning with universal horizon models\.InInternational Conference on Machine Learning \(ICML\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px3.p1.1)\.
- Dijkstra \(1959\)E\. W\. DijkstraA note on two problems in connexion with graphs\.Numerische Mathematik1\(1\),pp\. 269–271\.Cited by:[§3](https://arxiv.org/html/2610.00849#S3.p4.1)\.
- Espeholtet al\.\(2018\)L\. Espeholt, H\. Soyer, R\. Munos, K\. Simonyan, V\. Mnih, T\. Ward, Y\. Doron, V\. Firoiu, T\. Harley, I\. Dunning, S\. Legg, and K\. KavukcuogluIMPALA: scalable distributed deep\-RL with importance weighted actor\-learner architectures\.InInternational Conference on Machine Learning \(ICML\),Cited by:[Appendix E](https://arxiv.org/html/2610.00849#A5.SS0.SSS0.Px1.p2.1)\.
- Eysenbachet al\.\(2022\)B\. Eysenbach, T\. Zhang, S\. Levine, and R\. SalakhutdinovContrastive learning as goal\-conditioned reinforcement learning\.InNeural Information Processing Systems \(NeurIPS\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px1.p1.1),[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p1.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1)\.
- Fuet al\.\(2020\)J\. Fu, A\. Kumar, O\. Nachum, G\. Tucker, and S\. LevineD4RL: datasets for deep data\-driven reinforcement learning\.External Links:2004\.07219Cited by:[§D\.1](https://arxiv.org/html/2610.00849#A4.SS1.SSS0.Px2.p1.1)\.
- Ghoshet al\.\(2021\)D\. Ghosh, A\. Gupta, A\. Reddy, J\. Fu, C\. M\. Devin, B\. Eysenbach, and S\. LevineLearning to reach goals via iterated supervised learning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p1.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1)\.
- Hendrycks and Gimpel \(2016\)D\. Hendrycks and K\. GimpelGaussian error linear units \(GELUs\)\.CoRRabs/1606\.08415\.Cited by:[Appendix E](https://arxiv.org/html/2610.00849#A5.SS0.SSS0.Px1.p1.1)\.
- Jawaid \(2026\)A\. JawaidOffline RL with hierarchical action chunking\.Reinforcement Learning Journal7\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px2.p1.1)\.
- Jeon and Lee \(2026\)H\. Jeon and Y\. LeeRecursive value learning for long\-horizon offline goal\-conditioned RL\.CoRRabs/2609\.02237\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px1.p1.1)\.
- Kaelbling \(1993\)L\. P\. KaelblingLearning to achieve goals\.InInternational Joint Conference on Artificial Intelligence \(IJCAI\),Cited by:[§2](https://arxiv.org/html/2610.00849#S2.p1.1)\.
- Keet al\.\(2026\)K\. Ke, S\. He, C\. Xu, Y\. Luo, X\. Lan, and C\. YuAdaptive coarse\-to\-fine subgoal refinement for long\-horizon offline goal\-conditioned reinforcement learning\.CoRRabs/2605\.28127\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px2.p1.1)\.
- Kingma and Ba \(2015\)D\. P\. Kingma and J\. L\. BaAdam: a method for stochastic optimization\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix E](https://arxiv.org/html/2610.00849#A5.SS0.SSS0.Px3.p1.1)\.
- Kostrikovet al\.\(2022\)I\. Kostrikov, A\. Nair, and S\. LevineOffline reinforcement learning with implicit Q\-learning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p1.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1)\.
- Levineet al\.\(2020\)S\. Levine, A\. Kumar, G\. Tucker, and J\. FuOffline reinforcement learning: tutorial, review, and perspectives on open problems\.CoRRabs/2005\.01643\.Cited by:[§1](https://arxiv.org/html/2610.00849#S1.p1.1),[§2](https://arxiv.org/html/2610.00849#S2.p1.1)\.
- Opryshkoet al\.\(2026\)E\. Opryshko, J\. Quan, C\. Voelcker, Y\. Du, and I\. GilitschenskiTest\-time graph search for goal\-conditioned reinforcement learning\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2610.00849#S1.p1.1)\.
- Parket al\.\(2025\)S\. Park, K\. Frans, B\. Eysenbach, and S\. LevineOGBench: benchmarking offline goal\-conditioned RL\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§D\.3](https://arxiv.org/html/2610.00849#A4.SS3.p1.1),[Appendix D](https://arxiv.org/html/2610.00849#A4.p1.1),[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p1.1),[§1](https://arxiv.org/html/2610.00849#S1.p4.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1),[§5](https://arxiv.org/html/2610.00849#S5.p1.1)\.
- Parket al\.\(2023\)S\. Park, D\. Ghosh, B\. Eysenbach, and S\. LevineHIQL: offline goal\-conditioned RL with latent states as actions\.InNeural Information Processing Systems \(NeurIPS\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px2.p1.1),[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p1.1),[§1](https://arxiv.org/html/2610.00849#S1.p4.1),[§2](https://arxiv.org/html/2610.00849#S2.p2.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1)\.
- Penget al\.\(2019\)X\. B\. Peng, A\. Kumar, G\. Zhang, and S\. LevineAdvantage\-weighted regression: simple and scalable off\-policy reinforcement learning\.CoRRabs/1910\.00177\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2610.00849#S1.p3.1),[§2](https://arxiv.org/html/2610.00849#S2.p3.1)\.
- Perezet al\.\(2018\)E\. Perez, F\. Strub, H\. de Vries, V\. Dumoulin, and A\. CourvilleFiLM: visual reasoning with a general conditioning layer\.InAAAI Conference on Artificial Intelligence,Cited by:[Appendix E](https://arxiv.org/html/2610.00849#A5.SS0.SSS0.Px2.p2.1)\.
- Schaulet al\.\(2015\)T\. Schaul, D\. Horgan, K\. Gregor, and D\. SilverUniversal value function approximators\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2610.00849#S1.p1.1),[§2](https://arxiv.org/html/2610.00849#S2.p1.1)\.
- Shehmaret al\.\(2026\)D\. Shehmar, M\. Schlegel, M\. E\. Taylor, and M\. C\. MachadoLaplacian representations for decision\-time planning\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§6](https://arxiv.org/html/2610.00849#S6.p2.1)\.
- Sherstanet al\.\(2020\)C\. Sherstan, S\. Dohare, J\. MacGlashan, J\. Günther, and P\. M\. PilarskiGamma\-nets: generalizing value estimation over timescale\.InProceedings of the AAAI Conference on Artificial Intelligence,Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px3.p1.1)\.
- Suttonet al\.\(1999\)R\. S\. Sutton, D\. Precup, and S\. SinghBetween MDPs and semi\-MDPs: a framework for temporal abstraction in reinforcement learning\.Artificial Intelligence112\(1–2\),pp\. 181–211\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2610.00849#S1.p2.1)\.
- Wanget al\.\(2023\)T\. Wang, A\. Torralba, P\. Isola, and A\. ZhangOptimal goal\-reaching reinforcement learning via quasimetric learning\.InInternational Conference on Machine Learning \(ICML\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px1.p1.1),[§G\.1](https://arxiv.org/html/2610.00849#A7.SS1.p1.1),[§5\.1](https://arxiv.org/html/2610.00849#S5.SS1.SSS0.Px1.p2.1)\.
- Wibaultet al\.\(2026\)C\. Wibault, A\. Goldie, A\. Villares, M\. Osborne, and J\. FoersterAbstraction for offline goal\-conditioned reinforcement learning\.CoRRabs/2605\.22711\.Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2026\)B\. Zheng, V\. Myers, B\. Eysenbach, and S\. LevineScaling goal\-conditioned reinforcement learning with multistep quasimetric distances\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Appendix A](https://arxiv.org/html/2610.00849#A1.SS0.SSS0.Px1.p1.1)\.

## Appendix ARelated Work

#### Offline goal\-conditioned value learning\.

Offline goal\-conditioned reinforcement learning \(GCRL\) aims to learn policies for reaching arbitrary goals from fixed datasets\. Existing approaches differ in how they estimate long\-range goal\-conditioned values: contrastive RL learns state–goal representations through contrastive objectives\([Eysenbach et al\., 2022](https://arxiv.org/html/2610.00849#bib.bib4)\), whereas quasimetric RL exploits the structure of temporal distances\([Wang et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib5)\)\. Building on the latter,[Zheng et al\. \(2026\)](https://arxiv.org/html/2610.00849#bib.bib18)fit quasimetric distances with multistep Monte Carlo targets to improve long\-horizon learning\.[Jeon and Lee \(2026\)](https://arxiv.org/html/2610.00849#bib.bib19)instead organize value learning recursively over a balanced decomposition of observed trajectory segments, shortening the chain of bootstrapped dependencies before propagating values across trajectories\. These methods address long\-range estimation through different objectives\.

#### Hierarchical goal\-conditioned reinforcement learning\.

Hierarchical RL addresses long\-horizon decision\-making by decomposing complex tasks into simpler subtasks\. In offline GCRL,[Park et al\. \(2023\)](https://arxiv.org/html/2610.00849#bib.bib7)introduced a two\-level policy that proposes intermediate subgoals and executes primitive actions toward them, training both levels from goal\-conditioned values using advantage\-weighted regression\([Peng et al\., 2019](https://arxiv.org/html/2610.00849#bib.bib12)\)\. More recent methods modify the structure or execution of this hierarchy\.[Choi et al\. \(2026\)](https://arxiv.org/html/2610.00849#bib.bib20)generate sequences of latent subgoals autoregressively before predicting an action, while[Ke et al\. \(2026\)](https://arxiv.org/html/2610.00849#bib.bib21)recursively refine distant goals until a learned reachability criterion indicates that the current target can be executed locally\.[Jawaid \(2026\)](https://arxiv.org/html/2610.00849#bib.bib22)combines high\-level subgoal planning with low\-level action chunks and temporally extended value backups\.[Wibault et al\. \(2026\)](https://arxiv.org/html/2610.00849#bib.bib23)study relativized options and level\-specific representations to reuse experience across similar state–goal contexts\. Rather than changing subgoal generation, hierarchy depth, or low\-level execution, we retain the two\-level policy structure of HIQL and study how temporal abstraction in the high\-level value function affects subgoal evaluation\.

#### Temporal abstraction at multiple scales\.

Temporal abstraction has long been formalized through temporally extended actions\([Sutton et al\., 1999](https://arxiv.org/html/2610.00849#bib.bib8)\)\.[Ahn et al\. \(2025\)](https://arxiv.org/html/2610.00849#bib.bib9)introduced temporally abstracted value learning for offline GCRL, treating a fixed stride ofkkprimitive steps as one abstract transition\. This improves long\-range value propagation, but a fixedkkmust trade off sensitivity to distant goals against the ability to discriminate nearby states\. Multiple temporal horizons have also been studied in other components of offline RL:[Chung et al\. \(2026\)](https://arxiv.org/html/2610.00849#bib.bib24)learn predictive models that can sample future states at arbitrary horizons and use a controlled horizon distribution for model\-based value learning, while[Adaikkappan et al\. \(2026\)](https://arxiv.org/html/2610.00849#bib.bib25)use multiscale predictive supervision to align state and goal representations\. Outside offline RL,[Sherstan et al\. \(2020\)](https://arxiv.org/html/2610.00849#bib.bib31)introducedΓ\\Gamma\-nets, which condition a single value estimator on the discount factorγ\\gamma, enabling value predictions to generalize across multiple timescales\. In contrast, GITA represents multiple temporal abstraction scales directly within a singlekk\-conditioned goal\-value function, wherekkcontrols the number of primitive steps represented by each abstract transition\.

## Appendix BProofs for the Fixed\-Scale Resolution–Horizon Trade\-off

This section provides the proofs for Proposition[5](https://arxiv.org/html/2610.00849#S3.E5)and its corollaries\. Our analysis is an idealized shortest\-path model and is not intended to exactly characterize dataset\-induced transitions\. Letd=d⋆​\(s,g\)d=d^\{\\star\}\(s,g\)denote the minimum number of primitive transitions from statessto goalgg\. We follow the temporally abstracted update convention of OTA\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9)\), which evaluates the sparse reward at the abstract successorst\+ks\_\{t\+k\}, whereas HIQL indexes the one\-step reward atsts\_\{t\}\. This results in a one\-step reward\-indexing shift atk=1k=1, which does not affect the maximizing abstraction factor in the analysis below\. Following the theoretical formulation of OTA\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9)\), the primitive\-step value is

V⋆\(d\)=−∑j=0d−1γj=−1−γd1−γ\.V^\{\\star\}\(d\)=\-\\sum\_\{j=0\}^\{d\-1\}\\gamma^\{j\}=\-\\frac\{1\-\\gamma^\{d\}\}\{1\-\\gamma\}\.\(14\)
With abstraction factorkk, the corresponding abstract horizon isHk​\(d\)=⌈dk⌉H\_\{k\}\(d\)=\\left\\lceil\\frac\{d\}\{k\}\\right\\rceil, which gives

Vk⋆​\(d\)=−1−γ⌈dk⌉1−γ\.V\_\{k\}^\{\\star\}\(d\)=\-\\frac\{1\-\\gamma^\{\\left\\lceil\\frac\{d\}\{k\}\\right\\rceil\}\}\{1\-\\gamma\}\.\(15\)
Relaxing⌈d/k⌉\\lceil d/k\\rceiltod/kd/kso that it varies smoothly withkk, yields

V~k​\(s,g\)=−1−γd⋆​\(s,g\)/k1−γ,k\>0,0<γ<1\.\\widetilde\{V\}\_\{k\}\(s,g\)=\-\\frac\{1\-\\gamma^\{d^\{\\star\}\(s,g\)/k\}\}\{1\-\\gamma\},\\qquad k\>0,\\quad 0<\\gamma<1\.\(16\)
Consider a candidate subgoals′s^\{\\prime\}that makesΔ\\Deltaprimitive steps of progress toward the goal, such that

d⋆​\(s,g\)=d,d⋆​\(s′,g\)=d−Δ,d\>Δ\>0\.d^\{\\star\}\(s,g\)=d,\\qquad d^\{\\star\}\(s^\{\\prime\},g\)=d\-\\Delta,\\qquad d\>\\Delta\>0\.\(17\)
### B\.1Proof of Proposition[5](https://arxiv.org/html/2610.00849#S3.E5)

###### Proof\.

The high\-level advantage of the candidate subgoals′s^\{\\prime\}relative to the current statessis

A~k​\(s,s′,g\)=V~k​\(s′,g\)−V~k​\(s,g\)\.\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=\\widetilde\{V\}\_\{k\}\(s^\{\\prime\},g\)\-\\widetilde\{V\}\_\{k\}\(s,g\)\.\(18\)Applying[Equation16](https://arxiv.org/html/2610.00849#A2.E16)to bothssands′s^\{\\prime\}, with distance given by[Equation17](https://arxiv.org/html/2610.00849#A2.E17), we obtain

A~k​\(s,s′,g\)\\displaystyle\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=−1−γ\(d−Δ\)/k1−γ\+1−γd/k1−γ\\displaystyle=\-\\frac\{1\-\\gamma^\{\(d\-\\Delta\)/k\}\}\{1\-\\gamma\}\+\\frac\{1\-\\gamma^\{d/k\}\}\{1\-\\gamma\}=γ\(d−Δ\)/k−γd/k1−γ\\displaystyle=\\frac\{\\gamma^\{\(d\-\\Delta\)/k\}\-\\gamma^\{d/k\}\}\{1\-\\gamma\}=γ\(d−Δ\)/k​1−γΔ/k1−γ\.\\displaystyle=\\gamma^\{\(d\-\\Delta\)/k\}\\frac\{1\-\\gamma^\{\\Delta/k\}\}\{1\-\\gamma\}\.\(19\)Therefore,

A~k​\(s,s′,g\)=γ\(d−Δ\)/k⏟reachability​1−γΔ/k1−γ⏟resolution\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=\\underbrace\{\\gamma^\{\(d\-\\Delta\)/k\}\}\_\{\\text\{reachability\}\}\\;\\underbrace\{\\frac\{1\-\\gamma^\{\\Delta/k\}\}\{1\-\\gamma\}\}\_\{\\text\{resolution\}\}\(20\)∎

The*reachability factor*measures whether the residual distanced−Δd\-\\Deltafalls within a range where discounting carries signal, while the*resolution factor*measures how sharply the value function separatesssfrom a subgoalΔ\\Deltaprimitive steps closer\.

### B\.2Proof of the Vanishing\-Advantage Corollary

We prove the two limits in[Equation6](https://arxiv.org/html/2610.00849#S3.E6)separately\. Throughout, letλ:=−ln⁡γ\>0\\lambda:=\-\\ln\\gamma\>0, so thatγx=e−λ​x\\gamma^\{x\}=e^\{\-\\lambda x\}for any exponentxx\.

#### Limit ask→0\+k\\to 0^\{\+\}\.

###### Proof\.

Sinced−Δ\>0d\-\\Delta\>0is fixed,\(d−Δ\)/k→∞\(d\-\\Delta\)/k\\to\\inftyask→0\+k\\to 0^\{\+\}, so the reachability factor vanishes:

limk→0\+γ\(d−Δ\)/k=limk→0\+exp⁡\(−λ​d−Δk\)=0\.\\lim\_\{k\\to 0^\{\+\}\}\\gamma^\{\(d\-\\Delta\)/k\}=\\lim\_\{k\\to 0^\{\+\}\}\\exp\\left\(\-\\lambda\\frac\{d\-\\Delta\}\{k\}\\right\)=0\.\(21\)Meanwhile the resolution factor stays bounded for allkk, since0<γΔ/k<10<\\gamma^\{\\Delta/k\}<1implies

0<1−γΔ/k1−γ<11−γ\.0<\\frac\{1\-\\gamma^\{\\Delta/k\}\}\{1\-\\gamma\}<\\frac\{1\}\{1\-\\gamma\}\.\(22\)Combining Equation \([21](https://arxiv.org/html/2610.00849#A2.E21)\) and \([22](https://arxiv.org/html/2610.00849#A2.E22)\) in[Equation19](https://arxiv.org/html/2610.00849#A2.E19)gives

0<A~k​\(s,s′,g\)<11−γ​exp⁡\(−λ​d−Δk\)→k→0\+0,0<\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)<\\frac\{1\}\{1\-\\gamma\}\\exp\\left\(\-\\lambda\\frac\{d\-\\Delta\}\{k\}\\right\)\\xrightarrow\[k\\to 0^\{\+\}\]\{\}0,\(23\)so by the squeeze theorem,

limk→0\+A~k​\(s,s′,g\)=0\.\\lim\_\{k\\to 0^\{\+\}\}\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=0\.\(24\)∎

Althoughkkis discrete in the algorithm, this limit concerns its continuous relaxation\. Equivalently, the same vanishing regime is obtained for any fixed admissiblekkasd→∞d\\to\\infty, since it is governed by the ratio\(d−Δ\)/k\(d\-\\Delta\)/k\.

#### Limit ask→∞k\\to\\infty\.

###### Proof\.

Here the reachability factor tends to11while the resolution factor’s numerator1−γΔ/k1\-\\gamma^\{\\Delta/k\}tends to00\. Usinge−x=1−x\+𝒪⁡\(x2\)e^\{\-x\}=1\-x\+\\mathcal\{O\}\(x^\{2\}\)asx→0x\\to 0withx=λ​Δ/kx=\\lambda\\Delta/kandx=λ⁡\(d−Δ\)/kx=\\lambda\(d\-\\Delta\)/krespectively,

1−γΔ/k=λ​Δk\+𝒪⁡\(k−2\),γ\(d−Δ\)/k=1\+𝒪⁡\(k−1\)\.1\-\\gamma^\{\\Delta/k\}=\\frac\{\\lambda\\Delta\}\{k\}\+\\mathcal\{O\}\(k^\{\-2\}\),\\qquad\\gamma^\{\(d\-\\Delta\)/k\}=1\+\\mathcal\{O\}\(k^\{\-1\}\)\.\(25\)Substituting[Equation25](https://arxiv.org/html/2610.00849#A2.E25)into[Equation19](https://arxiv.org/html/2610.00849#A2.E19), all cross terms are𝒪⁡\(k−2\)\\mathcal\{O\}\(k^\{\-2\}\)or smaller, leaving

A~k​\(s,s′,g\)=λ​Δ\(1−γ\)​k\+𝒪⁡\(k−2\)\.\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=\\frac\{\\lambda\\Delta\}\{\(1\-\\gamma\)k\}\+\\mathcal\{O\}\(k^\{\-2\}\)\.\(26\)Since the leading term isΘ⁡\(1/k\)\\Theta\(1/k\),

limk→∞A~k​\(s,s′,g\)=0\\lim\_\{k\\to\\infty\}\\widetilde\{A\}\_\{k\}\(s,s^\{\\prime\},g\)=0\(27\)∎

Together,[Equations24](https://arxiv.org/html/2610.00849#A2.E24)and[27](https://arxiv.org/html/2610.00849#A2.E27)prove the vanishing\-advantage corollary\. Therefore, a coarser abstraction allows the advantage to better discriminate distant states when considering distant goals, because the*reachability factor*increases toward11askkgrows\. However, the same coarser abstraction blurs the distinction between nearby states, since the*resolution factor*decreases toward00askkgrows\. Nokksimultaneously maximizes both\.

### B\.3Proof of Corollary[8](https://arxiv.org/html/2610.00849#S3.E8)

###### Proof\.

From[Equation19](https://arxiv.org/html/2610.00849#A2.E19),

A~k​\(d,Δ\)=γ\(d−Δ\)/k−γd/k1−γ\.\\widetilde\{A\}\_\{k\}\(d,\\Delta\)=\\frac\{\\gamma^\{\(d\-\\Delta\)/k\}\-\\gamma^\{d/k\}\}\{1\-\\gamma\}\.\(28\)Differentiating with respect tokkgives

∂A~k​\(d,Δ\)∂k=−log⁡γ\(1−γ\)​k2​\[\(d−Δ\)​γ\(d−Δ\)/k−d​γd/k\]\.\\frac\{\\partial\\widetilde\{A\}\_\{k\}\(d,\\Delta\)\}\{\\partial k\}=\\frac\{\-\\log\\gamma\}\{\(1\-\\gamma\)k^\{2\}\}\\left\[\(d\-\\Delta\)\\gamma^\{\(d\-\\Delta\)/k\}\-d\\gamma^\{d/k\}\\right\]\.\(29\)Since−log⁡γ\>0\-\\log\\gamma\>0,1−γ\>01\-\\gamma\>0, andk2\>0k^\{2\}\>0, the sign of the derivative is governed entirely by the terms in brackets, and a stationary point satisfies

\(d−Δ\)​γ\(d−Δ\)/k=d​γd/k\.\(d\-\\Delta\)\\gamma^\{\(d\-\\Delta\)/k\}=d\\gamma^\{d/k\}\.\(30\)Dividing both sides byγ\(d−Δ\)/k\>0\\gamma^\{\(d\-\\Delta\)/k\}\>0and usingγd/k=γ\(d−Δ\)/k​γΔ/k\\gamma^\{d/k\}=\\gamma^\{\(d\-\\Delta\)/k\}\\gamma^\{\\Delta/k\},

d−Δ=d​γΔ/k⟹γΔ/k=d−Δd\.d\-\\Delta=d\\gamma^\{\\Delta/k\}\\qquad\\Longrightarrow\\qquad\\gamma^\{\\Delta/k\}=\\frac\{d\-\\Delta\}\{d\}\.\(31\)Taking logarithms and solving forkkgives the exact continuous optimum

kopt⋆​\(d,Δ\)=−Δ​log⁡γlog⁡\(dd−Δ\)\.k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)=\\frac\{\-\\Delta\\log\\gamma\}\{\\log\\left\(\\frac\{d\}\{d\-\\Delta\}\\right\)\}\.\(32\)
SinceγΔ/k\\gamma^\{\\Delta/k\}increases strictly and continuously from00to11askkranges over\(0,∞\)\(0,\\infty\), and\(d−Δ\)/d∈\(0,1\)\(d\-\\Delta\)/d\\in\(0,1\),[Equation31](https://arxiv.org/html/2610.00849#A2.E31)has a unique solution\. From[Equation29](https://arxiv.org/html/2610.00849#A2.E29), the derivative is positive before this solution and negative after it\. Therefore,kopt⋆​\(d,Δ\)k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)is the unique global maximizer of the relaxed advantage\. ∎

#### Dependence on state\-goal distance\.

For fixedΔ\\Deltaandγ\\gamma, writekopt⋆​\(d,Δ\)=C/L⁡\(d\)k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)=C/L\(d\)withC=−Δ​log⁡γ\>0C=\-\\Delta\\log\\gamma\>0andL⁡\(d\)=log⁡\(dd−Δ\)\.L\(d\)=\\log\\left\(\\frac\{d\}\{d\-\\Delta\}\\right\)\.Since

L′​\(d\)=1d−1d−Δ=−Δd⁡\(d−Δ\)<0,L^\{\\prime\}\(d\)=\\frac\{1\}\{d\}\-\\frac\{1\}\{d\-\\Delta\}=\-\\frac\{\\Delta\}\{d\(d\-\\Delta\)\}<0,\(33\)the quotient rule gives

∂kopt⋆∂d=\(−Δ​log⁡γ\)​Δd⁡\(d−Δ\)​\[log⁡\(dd−Δ\)\]2\>0\.\\frac\{\\partial k^\{\\star\}\_\{\\mathrm\{opt\}\}\}\{\\partial d\}=\\frac\{\(\-\\Delta\\log\\gamma\)\\Delta\}\{d\(d\-\\Delta\)\\left\[\\log\\left\(\\frac\{d\}\{d\-\\Delta\}\\right\)\\right\]^\{2\}\}\>0\.\(34\)Thus, the advantage\-maximizing abstraction factor increases with state\-goal distance\. This monotonicity result holds for fixedΔ\\Deltaandγ\\gamma\.

#### Local\-progress approximation\.

For the regime considered in Corollary[8](https://arxiv.org/html/2610.00849#S3.E8), letϵ=Δd≪1\.\\epsilon=\\frac\{\\Delta\}\{d\}\\ll 1\.Expanding the denominator of[Equation32](https://arxiv.org/html/2610.00849#A2.E32),

log⁡\(dd−Δ\)\\displaystyle\\log\\left\(\\frac\{d\}\{d\-\\Delta\}\\right\)=−log⁡\(1−ϵ\)\\displaystyle=\-\\log\(1\-\\epsilon\)=ϵ\+ϵ22\+𝒪⁡\(ϵ3\)\.\\displaystyle=\\epsilon\+\\frac\{\\epsilon^\{2\}\}\{2\}\+\\mathcal\{O\}\(\\epsilon^\{3\}\)\.\(35\)SubstitutingΔ=d​ϵ\\Delta=d\\epsiloninto[Equation32](https://arxiv.org/html/2610.00849#A2.E32)and dividing through byϵ\\epsilon,

kopt⋆​\(d,Δ\)=−d​ϵ​log⁡γϵ⁡\[1\+ϵ/2\+𝒪⁡\(ϵ2\)\]=−d​log⁡γ1\+ϵ/2\+𝒪⁡\(ϵ2\)\.k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)=\\frac\{\-d\\epsilon\\log\\gamma\}\{\\epsilon\\left\[1\+\\epsilon/2\+\\mathcal\{O\}\(\\epsilon^\{2\}\)\\right\]\}=\\frac\{\-d\\log\\gamma\}\{1\+\\epsilon/2\+\\mathcal\{O\}\(\\epsilon^\{2\}\)\}\.\(36\)Applying the reciprocal expansion1/\(1\+x\)=1−x\+𝒪⁡\(x2\)1/\(1\+x\)=1\-x\+\\mathcal\{O\}\(x^\{2\}\)asx→0x\\to 0, withx=ϵ/2x=\\epsilon/2, gives

kopt⋆​\(d,Δ\)=−d​log⁡γ⁡\[1−ϵ2\+𝒪⁡\(ϵ2\)\]\.k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)=\-d\\log\\gamma\\left\[1\-\\frac\{\\epsilon\}\{2\}\+\\mathcal\{O\}\(\\epsilon^\{2\}\)\\right\]\.\(37\)Therefore,

kopt⋆​\(d,Δ\)≈−d​log⁡γ,Δ≪dk^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)\\approx\-d\\log\\gamma,\\qquad\\Delta\\ll d\(38\)which proves[Equation7](https://arxiv.org/html/2610.00849#S3.E7)\.

#### Effective discount horizon\.

Dividingd−Δ=d⁡\(1−ϵ\)d\-\\Delta=d\(1\-\\epsilon\)by[Equation37](https://arxiv.org/html/2610.00849#A2.E37)and re\-expanding to first order inϵ\\epsilongives

d−Δkopt⋆​\(d,Δ\)=−1log⁡γ​\[1−ϵ2\+𝒪⁡\(ϵ2\)\],\\frac\{d\-\\Delta\}\{k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)\}=\-\\frac\{1\}\{\\log\\gamma\}\\left\[1\-\\frac\{\\epsilon\}\{2\}\+\\mathcal\{O\}\(\\epsilon^\{2\}\)\\right\],\(39\)so that, for allΔ≪d\\Delta\\ll d,

d−Δkopt⋆​\(d,Δ\)≈−1log⁡γ\.\\frac\{d\-\\Delta\}\{k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)\}\\approx\-\\frac\{1\}\{\\log\\gamma\}\.\(40\)Since−1/logγ∼1/\(1−γ\)\-1/\\log\\gamma\\sim 1/\(1\-\\gamma\)asγ→1\\gamma\\to 1, combining the two approximations yields

d−Δkopt⋆​\(d,Δ\)≈−1log⁡γ≈11−γ,Δ≪d,γ→1,\\frac\{d\-\\Delta\}\{k^\{\\star\}\_\{\\mathrm\{opt\}\}\(d,\\Delta\)\}\\approx\-\\frac\{1\}\{\\log\\gamma\}\\approx\\frac\{1\}\{1\-\\gamma\},\\qquad\\Delta\\ll d,\\quad\\gamma\\to 1,\(41\)which proves[Equation8](https://arxiv.org/html/2610.00849#S3.E8)\.

Together, these results show that no fixed abstraction factor can preserve informative advantages across all state\-goal distances: smallkkloses signal for distant goals, while largekksacrifices local resolution, and the best scale therefore grows with distance\.

## Appendix CFixed\-Scale Trade\-off Experiment Details

This section provides additional details on the experiment used to evaluate how the quality of the high\-level advantage varies with the abstraction factor and the state\-goal distance\.

#### Geodesic distance\.

For each environment, we sample\(s,s′,g\)\(s,s^\{\\prime\},g\)triples using the same sampling procedure used to construct training examples for the high\-level policy\. We then compute reference distances using the known geometry of the OGBench mazes\. Each maze is discretized into a fine grid, where navigable cells form the nodes of a four\-connected graph and edges connect adjacent cells without crossing walls\. The geodesic distance between two states is defined as the shortest\-path distance between their corresponding cells and is computed using Dijkstra’s algorithm\.

For each triple, we computedgeo​\(s,g\)d\_\{\\mathrm\{geo\}\}\(s,g\)anddgeo​\(s′,g\)d\_\{\\mathrm\{geo\}\}\(s^\{\\prime\},g\)and define the geodesic progress as

Δ​d=dgeo​\(s,g\)−dgeo​\(s′,g\)\.\\Delta d=d\_\{\\mathrm\{geo\}\}\(s,g\)\-d\_\{\\mathrm\{geo\}\}\(s^\{\\prime\},g\)\.\(42\)The triples are grouped into equal\-width intervals\. Geodesic distance is used only as a proxy for task progress and need not equal expert\-policy temporal distance\.

#### Advantage–Progress Correlation\.

For each abstraction factorkk, we compute the scale\-conditioned high\-level advantage

Akh​\(s,s′,g\)=Vkh​\(s′,g\)−Vkh​\(s,g\)\.A\_\{k\}^\{h\}\(s,s^\{\\prime\},g\)=V\_\{k\}^\{h\}\(s^\{\\prime\},g\)\-V\_\{k\}^\{h\}\(s,g\)\.\(43\)To measure how well this advantage ranks subgoals according to their progress toward the goal, we compute the Spearman rank correlation betweenAkh​\(s,s′,g\)A\_\{k\}^\{h\}\(s,s^\{\\prime\},g\)and the corresponding geodesic progressΔ​d\\Delta d\. We first group training triples\(s,s′,g\)\(s,s^\{\\prime\},g\)according to their geodesic state\-goal distancedgeo​\(s,g\)d\_\{\\mathrm\{geo\}\}\(s,g\)\. Each distance binbbcorresponds to an interval of state\-goal distances, andℐb\\mathcal\{I\}\_\{b\}denotes the set of triples whosedgeo​\(s,g\)d\_\{\\mathrm\{geo\}\}\(s,g\)falls within that interval\. For each abstraction factorkkand distance binbb, we define the Advantage\-Progress Correlation \(APC\) as

APCk,b=ρS​\(\{Ak,ih\}i∈ℐb,\{Δ​di\}i∈ℐb\)\.\\mathrm\{APC\}\_\{k,b\}=\\rho\_\{\\mathrm\{S\}\}\\left\(\\left\\\{A\_\{k,i\}^\{h\}\\right\\\}\_\{i\\in\\mathcal\{I\}\_\{b\}\},\\left\\\{\\Delta d\_\{i\}\\right\\\}\_\{i\\in\\mathcal\{I\}\_\{b\}\}\\right\)\.\(44\)Values close to11indicate that the learned advantage preserves the ordering induced by geodesic progress, values near00indicate little monotonic agreement, and negative values indicate an inverted ordering\. We compute APC independently for each training seed and report the mean and one standard deviation across four seeds\.

## Appendix DEnvironments

This section provides additional details on the environments and offline datasets used in our experiments\. We evaluate GITA on locomotion and manipulation tasks from OGBench\([Park et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib6)\)\. Our locomotion experiments use PointMaze, AntMaze, and HumanoidMaze, which combine maze navigation with increasingly complex control, while our manipulation experiments use Cube and Scene\. We consider state\-based observations for all five environments and pixel\-based variants of AntMaze, Cube, and Scene\.

### D\.1OGBench Environments

#### Maze navigation\.

PointMaze controls a 2\-D point mass directly through planar actions, AntMaze controls an 8\-DoF quadruped, and HumanoidMaze controls a 21\-DoF humanoid\. The task in each case is to reach a target location in the maze, with AntMaze and HumanoidMaze additionally requiring the agent to learn the underlying locomotion dynamics\. State\-based observations contain the full low\-dimensional state, including the agent’sxx\-yyposition\. In Visual\-AntMaze, the low\-dimensional observation is replaced by a64×64×364\\times 64\\times 3RGB image rendered from a third\-person camera, without additional proprioceptive state information\.

#### Maze layouts\.

OGBench provides themedium,large, andgiantlayouts\. Themediumandlargelayouts follow the corresponding D4RL\([Fu et al\., 2020](https://arxiv.org/html/2610.00849#bib.bib26)\)maze structures, whilegiantis approximately twice the size oflargeand contains substantially longer paths\. These layouts provide increasingly long state\-goal distances and therefore directly stress long\-horizon value propagation\. The main benchmark focuses primarily on thelargeandgiantlayouts\.

![Refer to caption](https://arxiv.org/html/2610.00849v1/maze_tessellations.png)Figure 5:Maze layouts\.OGBench provides themedium,large, andgiantlayouts considered in our experiments, with progressively longer state\-goal distances\.
#### Manipulation\.

Cube and Scene use a UR5e robot arm controlled through a 5\-dimensional end\-effector action space\. Cube requires arranging blocks into target configurations and provides variants with different numbers of cubes; our experiments use thesingleanddoublevariants for state observations\. Scene contains a cube, a drawer, a window, and two locking buttons, and its evaluation goals require composing multiple manipulation behaviors into a desired object configuration\. Both environments also support64×64×364\\times 64\\times 3RGB observations; in the visual variants, the arm is rendered transparently to reduce visual occlusion\.

### D\.2Offline Dataset Variants

The maze environments use three offline dataset variants with different trajectory structure and behavioral quality\. Thenavigatedatasets are collected by a noisy expert that repeatedly moves toward randomly sampled goals, producing relatively long and coherent goal\-reaching trajectories\. Thestitchdatasets instead consist of short local trajectory segments, so solving evaluation tasks requires combining information across multiple trajectories\. For AntMaze, we additionally consider theexploredataset, which contains high\-coverage but suboptimal behavior generated from randomly changing movement directions with substantial action noise\. State\- and pixel\-based AntMaze variants share the same underlying trajectories and differ only in their observation modality\. For manipulation, OGBench providesplayandnoisydatasets:playis generated by non\-Markovian scripted policies with temporally correlated noise, whereasnoisyuses Markovian scripted policies with uncorrelated Gaussian noise and higher state coverage\. Our state\-based Cube and Scene experiments useplay, while Visual\-Cube and Visual\-Scene usenoisy\. Figure[6](https://arxiv.org/html/2610.00849#A4.F6)illustrates the characteristic trajectory structure of the maze dataset variants\.

![Refer to caption](https://arxiv.org/html/2610.00849v1/maze_large_trajectories.png)Figure 6:OGBench datasets exhibit distinct trajectory structures\.Representative trajectories from the OGBench offline dataset variants in thelargemaze\. Thenavigatedataset contains long goal\-directed trajectories,stitchconsists of shorter local trajectory segments that must be combined to solve long\-horizon tasks, andexplorecontains high\-coverage but highly suboptimal behavior\.Table[3](https://arxiv.org/html/2610.00849#A4.T3)summarizes the observation and action spaces and evaluation horizons of the environment variants used in our experiments\. Table[4](https://arxiv.org/html/2610.00849#A4.T4)reports the corresponding dataset statistics\. Note that OGBench dataset episode lengths need not coincide with the maximum evaluation episode length\.

Table 3:Environment specifications\.Specifications for the OGBench domains considered in our experiments\.EnvironmentVariantObservationAction Dim\.Max\. Episode LengthPointMazelargestate \(2\)21000giantstate \(2\)21000AntMazelargestate \(29\) / pixels \(64×64×364\\times 64\\times 3\)81000giantstate \(29\) / pixels \(64×64×364\\times 64\\times 3\)81000HumanoidMazelargestate \(69\)212000giantstate \(69\)214000Cubesinglestate \(28\) / pixels \(64×64×364\\times 64\\times 3\)5200doublestate \(37\) / pixels \(64×64×364\\times 64\\times 3\)5500Scene–state \(40\) / pixels \(64×64×364\\times 64\\times 3\)5750Table 4:Offline dataset specifications\.Specifications of the OGBench offline datasets considered in our experiments\. Visual\-AntMaze uses the same trajectory datasets and statistics as the corresponding AntMaze configuration\.EnvironmentDatasetVariant\# Transitions\# EpisodesData Episode LengthPointMazenavigatelarge1M10001000navigategiant1M5002000stitchlarge1M5000200stitchgiant1M5000200AntMazenavigatelarge1M10001000navigategiant1M5002000stitchlarge1M5000200stitchgiant1M5000200explorelarge5M10000500HumanoidMazenavigatelarge2M10002000navigategiant4M10004000stitchlarge2M5000400stitchgiant4M10000400Cubeplaysingle1M10001000playdouble1M10001000Sceneplay–1M10001000Visual\-Cubenoisysingle1M10001000noisydouble1M10001000Visual\-Scenenoisy–1M10001000
### D\.3Evaluation Protocol

We follow the OGBench evaluation protocol\([Park et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib6)\), using five predefined state–goal evaluation tasks for each benchmark environment\. At each evaluation checkpoint, we run5050episodes for each of the five tasks, yielding250250evaluation episodes per checkpoint and seed\. We evaluate state\-based experiments over88random seeds and pixel\-based experiments over44random seeds, and report the mean success rate and standard deviation across seeds\.

The results in our main table are averaged over the final three evaluation checkpoints: for state\-based tasks, models are trained for10610^\{6\}gradient steps and evaluated at800​K800\\text\{K\},900​K900\\text\{K\}, and1​M1\\text\{M\}steps\. For pixel\-based tasks, models are trained for500​K500\\text\{K\}gradient steps and evaluated at300​K300\\text\{K\},400​K400\\text\{K\}, and500​K500\\text\{K\}steps\. Thus, each reported per\-seed score averages performance over3×5×50=7503\\times 5\\times 50=750evaluation episodes before results are aggregated across seeds\.

For the ablation studies, we instead report performance using only the final training checkpoint \(1​M1\\text\{M\}steps for state\-based tasks\)\.

## Appendix EImplementation Details

We implement GITA on top of the HIQL implementation provided by OGBench, which facilitates reproducibility and enables direct comparison with baselines evaluated under the same implementation framework\. Unless explicitly stated, we retain the HIQL network architectures, optimization settings, goal\-sampling distributions, and hyperparameters\. The main architectural change is restricted to the high\-level value function, which is additionally conditioned on the temporal abstraction factorkk\. The high\-level and low\-level policies retain the HIQL architectures\.

#### Network architecture\.

Following the OGBench implementation of HIQL, the value and policy networks use three hidden layers of size\(512,512,512\)\(512,512,512\)with Gaussian Error Linear Unit \(GELU\) activations\([Hendrycks and Gimpel, 2016](https://arxiv.org/html/2610.00849#bib.bib27)\)and Layer Normalization\([Ba et al\., 2016](https://arxiv.org/html/2610.00849#bib.bib28)\)\. For state\-based tasks, the goal representation network maps the concatenated state\-goal input to a1010\-dimensional length\-normalized representation using the same hidden dimensions\. The low\-level policy predicts primitive actions conditioned on the current state and the learned subgoal representation, while the high\-level policy predicts a1010\-dimensional subgoal representation\.

For pixel\-based tasks, we encode images with theimpala\_smallencoder, based on the IMPALA architecture\([Espeholt et al\., 2018](https://arxiv.org/html/2610.00849#bib.bib29)\), before the corresponding value, policy, and goal\-representation networks\. This encoder uses three residual stacks with1616,3232, and3232channels, one residual block per stack, followed by a512512\-dimensional MLP\. We use image augmentation with probability0\.50\.5for visual manipulation tasks and no augmentation for the visual maze tasks\.

The value function consists of two independently parameterized scalar heads,V1V\_\{1\}andV2V\_\{2\}, together with corresponding target networks\. The minimum of the two next\-state target values is used when constructing the conservative bootstrap signal, while policy advantages are computed from the mean of the two online value estimates\.

#### Abstraction\-factor conditioning\.

Only the high\-level value function receives the abstraction factor as input\. We providekkto the network on a logarithmic scale and compute a3232\-dimensional abstraction embedding

ek=fabs​\(log⁡k\),e\_\{k\}=f\_\{\\mathrm\{abs\}\}\(\\log k\),\(45\)wherefabsf\_\{\\mathrm\{abs\}\}is a one\-layer MLP with GELU activation\. We uselog⁡k\\log kbecause the abstraction factors span multiplicative temporal scales; the logarithm converts multiplicative changes inkkinto approximately additive changes in the conditioning input and reduces the dynamic range seen by the embedding network\.

We inject this embedding into the hidden layers of the high\-level value function using Feature\-wise Linear Modulation\([Perez et al\., 2018](https://arxiv.org/html/2610.00849#bib.bib16), FiLM;\)\. For a hidden representationhℓh\_\{\\ell\}at layerℓ\\ell, the abstraction embedding is projected into scale and shift parameters,

\[ηℓ​\(k\),δℓ​\(k\)\]=WℓFiLM​ek\+bℓFiLM,\\left\[\\eta\_\{\\ell\}\(k\),\\,\\delta\_\{\\ell\}\(k\)\\right\]=W\_\{\\ell\}^\{\\mathrm\{FiLM\}\}e\_\{k\}\+b\_\{\\ell\}^\{\\mathrm\{FiLM\}\},\(46\)and the hidden representation is modulated as

hℓ←\(1\+ηℓ​\(k\)\)⊙hℓ\+δℓ​\(k\)\.h\_\{\\ell\}\\leftarrow\\left\(1\+\\eta\_\{\\ell\}\(k\)\\right\)\\odot h\_\{\\ell\}\+\\delta\_\{\\ell\}\(k\)\.\(47\)FiLM is applied after the activation and LayerNorm in the high\-level value network\. The FiLM projection layers are initialized to zero\.

The state\-goal representation is independent of the abstraction factor:kkdoes not modify the state or goal encoder and is introduced only through FiLM in the high\-level value function\. Sincekkis introduced only through FiLM, the same state\-goal representation can be reused across abstraction factors, allowing GITA to support multiple temporal scales with only modest additional training cost\.

#### Optimization and hyperparameters\.

All networks are optimized with Adam\([Kingma and Ba, 2015](https://arxiv.org/html/2610.00849#bib.bib30)\)using a learning rate of3×10−43\\times 10^\{\-4\}\. Following the OGBench protocol, state\-based models use a minibatch size of10241024and are trained for10610^\{6\}gradient steps, while pixel\-based models use a minibatch size of256256and are trained for5×1055\\times 10^\{5\}gradient steps\. The target value networks are updated using Polyak averaging with coefficientτ=0\.005\\tau=0\.005, and the expectile parameter is fixed to0\.70\.7\.

We retain the HIQL subgoal offsetmm, independently of the abstraction factorkk\. We usem=25m=25for PointMaze and AntMaze,m=100m=100for HumanoidMaze, andm=10m=10for Cube and Scene, with the same offsets used for their corresponding visual variants\. The subgoal representation dimension is1010\.

We preserve the environment\-specific HIQL policy\-extraction and goal\-sampling hyperparameters from OGBench\. In particular, the AWR inverse temperature is33for the standardnavigateandstitchdatasets and1010for AntMazeexplore\. For value learning, goals are sampled as the current state, a geometrically sampled future state, or a random dataset state with probabilities0\.20\.2,0\.50\.5, and0\.30\.3, respectively\. For high\-level actor training,navigateuses trajectory goals,stitchuses an equal mixture of trajectory and random goals, andexploreuses random goals\.

The discount factorγ\\gammais the only HIQL hyperparameter that we re\-tune for GITA\. The OGBench reference implementation already uses environment\-dependent discount factors for HIQL, withγ=0\.995\\gamma=0\.995for the longest\-horizon maze environments andγ=0\.99\\gamma=0\.99otherwise\. In GITA, we re\-tuneγ\\gammabecause it plays an additional role beyond TD learning: we use the same value both as the discount factor in the temporal\-difference updates and as the parameter controlling the geometric distribution over abstraction factors\. We therefore sweepγ∈\{0\.98,0\.99,0\.995\}\\gamma\\in\\\{0\.98,0\.99,0\.995\\\}and tune it independently for each task\. For simplicity, we choseα=γ\\alpha=\\gamma\. For the manipulation tasks, we additionally considerγ=0\.97\\gamma=0\.97\. All remaining HIQL hyperparameters are kept unchanged\.

For the candidate abstraction factors, we use values from the Fibonacci sequence, starting from33\. This gives𝒦=\{3,5,8,13,21\}\\mathcal\{K\}=\\\{3,5,8,13,21\\\}\. The range was chosen to cover the abstraction factors used by the fixed\-scale OTA baseline: the smallest candidate,33, is below the smallest task\-specific abstraction factor selected for OTA, while2121extends beyond the largest fixed factor\. For OTA, we use the results reported under the unified\-hyperparameter setting, except for the abstraction factor, which is selected separately for each task\.

Table 5:Hyperparameters used for GITA\.Unless marked as GITA\-specific, the values follow the HIQL reference implementation in OGBench\.

## Appendix FStatistical Testing

We assess differences between GITA and each baseline using a two\-sided paired Wilcoxon signed\-rank test\. Each task constitutes one paired observation, using the mean success rate reported in the corresponding results table; thus, the main comparison uses2323tasks and the ablation comparison uses1313tasks\. Because multiple baselines are compared against GITA, we correct the resultingpp\-values using the Holm–Bonferroni step\-down procedure\. Corrections are applied separately within each family of comparisons: the seven baselines in the main results and the four variants in the ablation study\. We consider a difference statistically significant when the Holm\-adjustedpp\-value is below0\.050\.05\.

Table 6:Two\-sided paired Wilcoxon signed\-rank tests comparing GITA with the baselines in[Table1](https://arxiv.org/html/2610.00849#S5.T1)\.Each test is performed over the2323tasks reported in the main results table\.padjp\_\{\\mathrm\{adj\}\}denotes the Holm–Bonferroni\-adjustedpp\-value across the seven comparisons\.All comparisons remain significant after correction, supporting the aggregate improvement of GITA over each baseline in the main evaluation\. And GITA significantly improves over removing scale conditioning and hard scale selection, while the two alternative aggregation/sampling variants are not significantly different after correction\.

Table 7:Two\-sided paired Wilcoxon signed\-rank tests comparing GITA with the variants in[Table2](https://arxiv.org/html/2610.00849#S5.T2)\.padjp\_\{\\mathrm\{adj\}\}denotes the Holm–Bonferroni\-adjustedpp\-value across the four comparisons\.
## Appendix GBaselines

This section provides additional details on the baseline methods considered in our experiments\.

### G\.1OGBench Baselines

We compare against the six offline GCRL algorithms benchmarked in OGBench\([Park et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib6)\)\. Goal\-conditioned behavioral cloning\([Ghosh et al\., 2021](https://arxiv.org/html/2610.00849#bib.bib3), GCBC;\)trains a goal\-conditioned policy by relabeling transitions with future states from the same trajectory as goals\. Goal\-conditioned implicit value learning and goal\-conditioned implicit Q\-learning\([Kostrikov et al\., 2022](https://arxiv.org/html/2610.00849#bib.bib13);[Park et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib7), GCIVL and GCIQL;\)are goal\-conditioned variants of Implicit Q\-Learning \(IQL\)\. Both use expectile regression to estimate optimal goal\-conditioned value functions\. GCIQL learns both a state\-value functionV⁡\(s,g\)V\(s,g\)and an action\-value functionQ⁡\(s,a,g\)Q\(s,a,g\), whereas GCIVL is the value\-only variant introduced with HIQL and learns onlyV⁡\(s,g\)V\(s,g\)\. Quasimetric reinforcement learning\([Wang et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib5), QRL;\)models goal\-reaching distance with a learned quasimetric and exploits its triangle\-inequality structure to learn goal\-conditioned values\. Contrastive reinforcement learning\([Eysenbach et al\., 2022](https://arxiv.org/html/2610.00849#bib.bib4), CRL;\)estimates a Monte Carlo goal\-conditioned value function through contrastive learning and performs one\-step policy improvement\. Hierarchical implicit Q\-learning\([Park et al\., 2023](https://arxiv.org/html/2610.00849#bib.bib7), HIQL;\)extracts a two\-level hierarchical policy from a single GCIVL value function\. Its high\-level policy predicts a latent representation of an intermediatemm\-step subgoal, while its low\-level policy predicts primitive actions conditioned on that subgoal representation\.

We additionally compare against option\-aware temporally abstracted value learning\([Ahn et al\., 2025](https://arxiv.org/html/2610.00849#bib.bib9), OTA;\), which directly builds on HIQL and is the closest fixed\-scale baseline to GITA\. OTA replaces primitive one\-step backups in the high\-level value objective with temporally extended transitions using a fixed abstraction factorkk, thereby shortening the effective value\-propagation horizon\. Unlike GITA, however, OTA uses a single abstraction factor for all state\-goal pairs within a task and therefore requires task\-specific selection ofkk\. We use the task\-specific abstraction factors and hyperparameters reported for OTA while otherwise following the corresponding HIQL/OGBench implementation\.

## Appendix HAdditional Ablation Details

We provide additional details on the ablations introduced in[Section5\.2](https://arxiv.org/html/2610.00849#S5.SS2)\. Each variant changes a single component of GITA while leaving the remaining architecture, optimization procedure, goal sampling, and low\-level policy unchanged\. Unless stated otherwise, the abstraction\-conditioned high\-level value function is evaluated over the same candidate set𝒦\\mathcal\{K\}used by GITA\. Recall that, for a training tuple\(st,st\+m,g\)\(s\_\{t\},s\_\{t\+m\},g\), GITA computes

Ah​\(st,st\+m,g,k\)=Vθh​\(st\+m,g,k\)−Vθh​\(st,g,k\)A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k\)=V\_\{\\theta\}^\{h\}\(s\_\{t\+m\},g,k\)\-V\_\{\\theta\}^\{h\}\(s\_\{t\},g,k\)\(48\)for everyk∈𝒦k\\in\\mathcal\{K\}and uses the aggregated AWR weight

Wβh=1\|𝒦\|​∑k∈𝒦exp⁡\(βh​Ah​\(st,st\+m,g,k\)\)W\_\{\\beta\_\{h\}\}=\\frac\{1\}\{\|\\mathcal\{K\}\|\}\\sum\_\{k\\in\\mathcal\{K\}\}\\exp\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k\)\\right\)\(49\)to train the shared high\-level policy\. The following ablations separately modify the value conditioning, the scale aggregation rule, or the candidate abstraction factors\.

#### Unconditioned Value\.

This variant removes the abstraction factor from the high\-level value function, replacingVθh​\(s,g,k\)V\_\{\\theta\}^\{h\}\(s,g,k\)with a single unconditioned value functionVθh​\(s,g\)V\_\{\\theta\}^\{h\}\(s,g\)\. Value training still uses temporally abstract transitions spanning different numbers of primitive steps, but the realized abstraction factor is not provided to the value network\. The resulting value function must therefore fit supervision from multiple temporal scales within the same mapping\. Since scale\-specific values are no longer available, the high\-level policy is trained from the single advantage

Ah​\(st,st\+m,g\)=Vθh​\(st\+m,g\)−Vθh​\(st,g\),A^\{h\}\(s\_\{t\},s\_\{t\+m\},g\)=V\_\{\\theta\}^\{h\}\(s\_\{t\+m\},g\)\-V\_\{\\theta\}^\{h\}\(s\_\{t\},g\),\(50\)using the standard AWR weight

exp⁡\(βh​Ah​\(st,st\+m,g\)\)\.\\exp\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g\)\\right\)\.\(51\)This ablation isolates whether training on multiple temporal scales is sufficient by itself, or whether explicitly representing the abstraction factor in the value function is necessary\.

#### Highest Advantage Factor\.

This variant retains the abstraction\-conditioned value function and computesAh​\(st,st\+m,g,k\)A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k\)for everyk∈𝒦k\\in\\mathcal\{K\}as in GITA, but replaces soft aggregation across scales with hard selection\. For each training sample, we select

k⋆=arg⁡maxk∈𝒦​Ah​\(st,st\+m,g,k\)k^\{\\star\}=\\arg\\max\_\{k\\in\\mathcal\{K\}\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k\)\(52\)and train the high\-level policy using only

Wβhmax=exp⁡\(βh​Ah​\(st,st\+m,g,k⋆\)\)\.W\_\{\\beta\_\{h\}\}^\{\\mathrm\{max\}\}=\\exp\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k^\{\\star\}\)\\right\)\.\(53\)All other factors are discarded for that sample\. This ablation tests whether the multi\-scale policy update benefits from combining supervision across several temporal resolutions, or whether selecting only the scale with the largest advantage is sufficient\.

#### Aggregate Advantage Before Exp\.

This variant preserves all scale\-specific advantages but changes the order of aggregation and the nonlinear AWR transformation\. GITA first exponentiates each advantage and then averages the resulting weights,

Wβh=1\|𝒦\|​∑k∈𝒦exp⁡\(βh​Ah​\(st,st\+m,g,k\)\)\.W\_\{\\beta\_\{h\}\}=\\frac\{1\}\{\|\\mathcal\{K\}\|\}\\sum\_\{k\\in\\mathcal\{K\}\}\\exp\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k\)\\right\)\.\(54\)The ablation instead first averages the raw advantages,

A¯h​\(st,st\+m,g\)=1\|𝒦\|​∑k∈𝒦Ah​\(st,st\+m,g,k\),\\bar\{A\}^\{h\}\(s\_\{t\},s\_\{t\+m\},g\)=\\frac\{1\}\{\|\\mathcal\{K\}\|\}\\sum\_\{k\\in\\mathcal\{K\}\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g,k\),\(55\)and then applies the AWR exponential,

Wβhpre=exp⁡\(βh​A¯h​\(st,st\+m,g\)\)\.W\_\{\\beta\_\{h\}\}^\{\\mathrm\{pre\}\}=\\exp\\left\(\\beta\_\{h\}\\bar\{A\}^\{h\}\(s\_\{t\},s\_\{t\+m\},g\)\\right\)\.\(56\)This removes the scale\-specific amplification induced by applying the exponential before aggregation\. The comparison therefore isolates whether preserving differences in the magnitude of the individual scale\-conditioned advantages through the nonlinear AWR weighting is important for policy supervision\.

#### GITA with Geometrickk\.

The default GITA policy update evaluates each training sample on a fixed candidate set𝒦\\mathcal\{K\}\. This ablation replaces that set with freshly sampled abstraction factors at every high\-level policy update\. We sample the same numberK=\|𝒦\|K=\|\\mathcal\{K\}\|of factors independently according to

kj∼Geom\(1−ζ\),p\(kj\)=\(1−ζ\)ζkj−1,j=1,…,K,k\_\{j\}\\sim\\operatorname\{Geom\}\(1\-\\zeta\),\\qquad p\(k\_\{j\}\)=\(1\-\\zeta\)\\zeta^\{k\_\{j\}\-1\},\\qquad j=1,\\ldots,K,\(57\)withζ=0\.9\\zeta=0\.9, corresponding to a mean abstraction factor of1010\. The parameterζ\\zetais specific to this ablation and is distinct from the value\-learning parameterα=γ\\alpha=\\gamma\. The sampled set

𝒦geo=\{k1,…,kK\}\\mathcal\{K\}\_\{\\mathrm\{geo\}\}=\\\{k\_\{1\},\\ldots,k\_\{K\}\\\}\(58\)then replaces the fixed candidate set in the standard GITA aggregation,

Wβhgeo=1K​∑k∈𝒦geoexp⁡\(βh​Ah​\(st,st\+m,g,k\)\)\.W\_\{\\beta\_\{h\}\}^\{\\mathrm\{geo\}\}=\\frac\{1\}\{K\}\\sum\_\{k\\in\\mathcal\{K\}\_\{\\mathrm\{geo\}\}\}\\exp\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g;k\)\\right\)\.\(59\)This change affects only the candidate factors used to construct the high\-level policy supervision; the abstraction\-conditioned value\-learning procedure, including its sampling distribution parameterized byα\\alpha, is unchanged\. The ablation tests whether GITA depends on a manually specified finite set of candidate scales or can obtain similar behavior from stochastic coverage of the temporal\-scale space\.

## Appendix IFurther Results and Visualizations

### I\.1Scale\-Dependent Policy Weights Across State–Goal Distances

To further visualize how different temporal scales contribute to high\-level policy learning, we examine the scale\-specific AWR weights as a function of state–goal distance\. For each abstraction factorkk, we compute

wk=min⁡\{exp⁡\(βh​Ah​\(st,st\+m,g,k\)\),100\},w\_\{k\}=\\min\\left\\\{\\exp\\left\(\\beta\_\{h\}A^\{h\}\(s\_\{t\},s\_\{t\+m\},g;k\)\\right\),100\\right\\\},\(60\)whereAh​\(st,st\+m,g,k\)A^\{h\}\(s\_\{t\},s\_\{t\+m\},g;k\)is the abstraction\-conditioned high\-level advantage from[Equation11](https://arxiv.org/html/2610.00849#S4.E11)\. The clipping threshold of100100is also applied during training, so the plotted quantity corresponds to the final AWR weight used in the high\-level policy update, up to the particular set ofkkvalues shown\. This clipping convention is inherited from the official HIQL and OTA implementations\. We group samples according to their geodesic state–goal distance and report the mean ofwkw\_\{k\}within each distance bin for each abstraction factor\. Curves are averaged across seeds, with shaded regions indicating one standard deviation\.

Recall from[Equation12](https://arxiv.org/html/2610.00849#S4.E12)that GITA averages the exponentiated advantages acrosskk\. Thus, scales assigning larger positive advantages to a training tuple contribute disproportionately to its policy weight, without requiring an explicit scale\-selection rule\. Figure[7](https://arxiv.org/html/2610.00849#A9.F7)shows this behavior across PointMaze, AntMaze, and HumanoidMaze for thenavigateandstitchdatasets and thelargeandgiantmaze layouts\.

Figure 7:Scale\-specific AWR weights across navigation and stitching tasks\.Mean clipped exponentiated high\-level advantage,min⁡\{exp⁡\(βh​Ah\),100\}\\min\\\{\\exp\(\\beta\_\{h\}A^\{h\}\),100\\\}, as a function of geodesic state–goal distance for different abstraction factorskk\. From top to bottom:large\-navigate,giant\-navigate,large\-stitch, andgiant\-stitch\. Curves show the mean across seeds and shaded regions indicate one standard deviation\. Each plot contains the corresponding maze morphologies\.Across the maze environments, the scale\-specific AWR weights exhibit a clear dependence on state–goal distance\. Smaller abstraction factors lose their advantage signal more rapidly as distance increases\. Larger factors retain informative weights over longer ranges, with the transition occurring progressively later askkincreases\. This pattern is especially clear in thegiantlayouts and is consistent with the reachability–resolution trade\-off predicted by our analysis\.

### I\.2Results Across Discount Factors

To characterize the sensitivity of GITA to the discount factor, we report performance for each value considered in our sweep,γ∈\{0\.98,0\.99,0\.995\}\\gamma\\in\\\{0\.98,0\.99,0\.995\\\}, on the maze tasks in[Table8](https://arxiv.org/html/2610.00849#A9.T8)\. Unlike the results reported in the main evaluation, this table uses only the final checkpoint after10610^\{6\}gradient steps\. Results are reported as mean±\\pmstandard deviation over 8 seeds\.

Because we setα=γ\\alpha=\\gamma, this sweep changes both the discount factor and the distribution of abstraction factors\. In particular,γ=α=0\.995\\gamma=\\alpha=0\.995gives𝔼⁡\[k\]=200\\mathbb\{E\}\[k\]=200, which matches the data episode length of PointMaze and AntMazestitch\. As a result, many sampled abstraction factors are truncated by trajectory boundaries, so part of the performance drop at largerγ\\gammamay come from the induced geometric distribution rather than from discounting alone\. This interaction further justifies re\-tuningγ\\gammafor GITA\.

Table 8:Success rate \(%\) of GITA for each discount factor considered in our sweep\.Results use only the final checkpoint after10610^\{6\}gradient steps and are reported as mean±\\pmstandard deviation over 8 seeds\. The*Bestγ\\gammaper Task*column reports the best task\-level result together with the corresponding discount factor in parentheses\. The best global discount factor isγ=0\.99\\gamma=0\.99, selected by the average performance across all tasks; the last column reports the corresponding task\-level performance\.

相似文章

目标条件监督学习用于LLM微调

arXiv cs.LG

本文提出了目标条件监督学习(GCSL)作为LLM的离线微调框架,该方法将反馈作为显式目标,通过一种新颖的目标公式和自然语言目标表示,使用监督学习训练模型。在无毒生成、代码生成和LLM推荐三个任务上的评估显示,该方法优于标准的离线基线方法。

性能驱动的多时间尺度学习环境抽象

arXiv cs.LG

本文提出了一种用于强化学习的性能驱动的状态抽象方法,直接优化决策质量,采用多时间尺度框架共同调整策略和树状结构抽象。该算法基于Q值差异细化或聚合状态空间,相比基线实现了更好的样本效率和更快的重新规划。

自适应多时间视野强化学习

arXiv cs.LG

本文提出一种多时间视野强化学习方法,能够自适应地选择并组合时间视野,无需手动调整折扣因子即可鲁棒地适应变化的奖励结构,并在MiniGrid环境中进行了实验验证。