从随机探索中学习规划
摘要
由UCLA主导的论文表明,可以仅通过随机探索来学习长期规划:采用条件能量模型估计时间对数密度比,无需动作或奖励标签,仅凭状态、图像和第一人称观测即可实现多尺度的认知地图式导航。
arXiv:2609.38383v1 Announce Type: new
Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.
查看缓存全文
缓存时间: 2026/10/02 09:49
# Learning to Plan from Random Exploration
Source: [https://arxiv.org/html/2609.38383](https://arxiv.org/html/2609.38383)
Deqian Kong1,,Guangyan Sun2,11footnotemark:1,Sheng Cheng3,11footnotemark:1,Sirui Xie1,Bo Pang4,Jianwen Xie5,Tony Geng6,Caiwen Ding2,Ying Nian Wu11UCLA2University of Minnesota3Amazon AGI4Salesforce Research5Lambda6Rice University††thanks:Equal contribution\.
###### Abstract
Random exploration reveals how an environment can be traversed before a goal is specified\. Can this experience support long\-range planning without policy\-improvement training? Our random\-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions\. We learn these relations with a conditional energy\-based model that estimates temporal log\-density ratios through horizon\-conditioned embeddings\. The model is trained on observation pairs by noise\-contrastive estimation, without action or reward labels\. The planner queries these learned relations at different horizons as it moves toward the goal\. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons\. The agent executes one action and replans with both models fixed\. Experiments demonstrate long\-range maze planning from random exploration using states and images\. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map\. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data\.
## 1Introduction
In a classic latent\-learning experiment, rats explored a complex maze for ten days with no food waiting at the goal\([Tolman and Honzik, 1930](https://arxiv.org/html/2609.38383#bib.bib30)\)\. After food was introduced, rats that had wandered unrewarded took nearly as few wrong turns on subsequent trials as rats rewarded from the start\. Exploration had taught them far more than their behavior had revealed\. The cognitive\-map hypothesis and studies of hippocampal place cells link navigation to internal spatial representations\([Tolman, 1948](https://arxiv.org/html/2609.38383#bib.bib22);[O’Keefe and Dostrovsky, 1971](https://arxiv.org/html/2609.38383#bib.bib31);[Zhao et al\., 2025b](https://arxiv.org/html/2609.38383#bib.bib14)\), while sequence\-learning accounts explain how these representations can arise from temporal experience\([Raju et al\., 2024](https://arxiv.org/html/2609.38383#bib.bib32)\)\. We ask whether even*random*exploration is enough: can representations learned from trajectories collected without goal\-directed or curiosity\-driven action selection support long\-range planning?
A random walk through an environment leaves temporal traces at multiple scales\. Observations a few steps apart reveal local connections between states, like reading a map at street level\. Observations many steps apart reveal connections between regions, like reading it at district level\. We show that, under random\-walk assumptions, short\-horizon statistics recover geodesic geometry in the diffusion limit while longer horizons capture regional connectivity before mixing erases these distinctions\. Classical results on heat kernels and diffusion processes\([Varadhan, 1967](https://arxiv.org/html/2609.38383#bib.bib18);[Norris, 1997](https://arxiv.org/html/2609.38383#bib.bib19);[Coifman and Lafon, 2006](https://arxiv.org/html/2609.38383#bib.bib11)\)provide the basis for this analysis\. The challenge is to learn these relations from sampled trajectories, without access to the full transition kernel, and let a planner query different scales as it moves toward a goal\.
Existing methods obtain long\-range guidance in different ways\. Goal\-conditioned reinforcement learning learns values and trains policies to act on them\([Eysenbach et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib4);[Park et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib28)\)\. Map\-like representations can emerge through navigation objectives\([Wijmans et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib33)\)\. Latent world models learn action\-conditioned predictions from reward\-free trajectories and evaluate multistep action sequences at test time\([Zhou et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib15);[Wang et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib16)\), but their planners depend on rollout accuracy and a latent geometry that reflects goal progress\. Temporal contrastive representations can encode probability ratios that support planning via interpolation in representation space\([Eysenbach et al\., 2024](https://arxiv.org/html/2609.38383#bib.bib8)\)\. We separate long\-range guidance from local dynamics prediction: a representation learned across horizons evaluates progress toward the goal, while a one\-step dynamics model predicts the outcomes of candidate actions\. Planning requires neither a trained goal\-reaching policy nor multistep dynamics rollouts\.
We model temporal relations with a conditional energy\-based model\([LeCun et al\., 2006](https://arxiv.org/html/2609.38383#bib.bib75);[Xie et al\., 2016](https://arxiv.org/html/2609.38383#bib.bib62)\)whose source and target embeddings estimate temporal log\-density ratios at each horizon\. Noise\-contrastive estimation\([Gutmann and Hyvärinen, 2010](https://arxiv.org/html/2609.38383#bib.bib2)\)learns them from observation pairs, requiring no action or reward labels\. At test time, the dynamics model predicts immediate action outcomes and the temporal representation scores their estimated improvement in goal likelihood\. The temporal representation is reusable across goals, and the agent replans after each action with both models fixed\.
Our contributions are threefold:
1. \(1\)We develop a conditional temporal model that learns multi\-horizon representations from exploration trajectories and supports closed\-loop goal reaching without policy\-improvement training\.
2. \(2\)We characterize the geodesic and connectivity information in the ideal random\-walk temporal score, and unify greedy horizon selection and potential ascent through a temperature\-controlled planning objective\.
3. \(3\)We demonstrate long\-range maze planning from random exploration using states and images\. Score fields, embedding probes, and planned routes reveal properties of a multiscale cognitive map\. Further evaluations cover egocentric navigation and manipulation with suboptimal data\.
## 2Method
Exploration trajectories reveal how an environment can be traversed before a goal is specified\. We learn these temporal relations at multiple horizons, then query them to turn predictions of local action outcomes into goal\-directed decisions\.
### 2\.1Problem Formulation
Given a dataset of trajectories in which each observationxtx\_\{t\}is a state vector or an image, we aim to learn temporal relations that support goal\-directed planning\. Our focus is random exploration for navigation: locally feasible moves are sampled without reference to a task goal\. The learning formulation also accepts suboptimal or expert trajectories, which can be useful in manipulation tasks\. Data collection can therefore balance broad exploration with coverage of task\-relevant transitions\. In each case, we learn the temporal relations induced by the behavior represented in the dataset\.
Let𝒯\\mathcal\{T\}be a finite set of positive prediction horizons\. At horizonτ∈𝒯\\tau\\in\\mathcal\{T\}, we sample a source–target pair\(x,y\)=\(xt,xt\+τ\)\(x,y\)=\(x\_\{t\},x\_\{t\+\\tau\}\)from the same episode, with distributionpdata\(x,y\|τ\)p\_\{\\rm data\}\(x,y\|\\tau\)\. This distribution describes where exploration leads fromxxafterτ\\tausteps, and varyingτ\\taudescribes this at different time scales\. We modelpdata\(y\|x,τ\)p\_\{\\rm data\}\(y\|x,\\tau\)so that, once a goal is given, the planner can ask how likely each candidate next observation is to lead to it withinτ\\tausteps\. Learning this conditional uses only observation pairs, without action or reward labels\. Action\-labeled one\-step transitions\(xt,at,xt\+1\)\(x\_\{t\},a\_\{t\},x\_\{t\+1\}\)are used separately to learn local dynamics\.
### 2\.2Models
A multi\-horizon representation\.The temporal relation between two observations depends on how long the agent explores\. We describe it through the alignment of a source embeddinghθ\(x,τ\)h\_\{\\theta\}\(x,\\tau\)and a target embeddinggθ\(y,τ\)g\_\{\\theta\}\(y,\\tau\)\. The family\{hθ\(⋅,τ\),gθ\(⋅,τ\)\}τ∈𝒯\\\{h\_\{\\theta\}\(\\cdot,\\tau\),g\_\{\\theta\}\(\\cdot,\\tau\)\\\}\_\{\\tau\\in\\mathcal\{T\}\}represents how these relations change with the horizon\. Two observations may have little association over a short horizon but become associated once exploration has had time to traverse a connecting route\.
The planner can query this family at different horizons to evaluate the same candidate move at several temporal scales\. In spatial environments, these relations can exhibit properties of a*multiscale cognitive map*\([Tolman, 1948](https://arxiv.org/html/2609.38383#bib.bib22);[Zhao et al\., 2025b](https://arxiv.org/html/2609.38383#bib.bib14)\)\.[Section3](https://arxiv.org/html/2609.38383#S3)connects them to geodesic geometry and connectivity between regions under random\-walk assumptions\.
Conditional temporal model\.We model the target observationyygiven the sourcexxand horizonτ\\tauwith a conditional energy\-based model \(EBM\)\. Letp0\(y\)p\_\{0\}\(y\)be a reference distribution whose support covers the target observations\. The model is
pθ\(y\|x,τ\)=p0\(y\)Zθ\(x,τ\)exp\(β0\(τ\)⟨hθ\(x,τ\),gθ\(y,τ\)⟩\),p\_\{\\theta\}\(y\|x,\\tau\)=\\frac\{p\_\{0\}\(y\)\}\{Z\_\{\\theta\}\(x,\\tau\)\}\\exp\\\!\\left\(\\beta\_\{0\}\(\\tau\)\\langle h\_\{\\theta\}\(x,\\tau\),g\_\{\\theta\}\(y,\\tau\)\\rangle\\right\),\(1\)whereZθ\(x,τ\)=𝔼p0\(y\)\[exp\(β0\(τ\)⟨hθ\(x,τ\),gθ\(y,τ\)⟩\)\]Z\_\{\\theta\}\(x,\\tau\)=\\mathbb\{E\}\_\{p\_\{0\}\(y\)\}\\\!\\left\[\\exp\\\!\\left\(\\beta\_\{0\}\(\\tau\)\\langle h\_\{\\theta\}\(x,\\tau\),g\_\{\\theta\}\(y,\\tau\)\\rangle\\right\)\\right\]normalizes the conditional distribution\. Both embeddings have unit norm,‖hθ\(x,τ\)‖2=‖gθ\(y,τ\)‖2=1\\\|h\_\{\\theta\}\(x,\\tau\)\\\|\_\{2\}=\\\|g\_\{\\theta\}\(y,\\tau\)\\\|\_\{2\}=1\. For a given source and horizon, targets with larger alignment receive greater weight relative top0p\_\{0\}\. The learned coefficientβ0\(τ\)\\beta\_\{0\}\(\\tau\)controls the strength of this dependence across horizons\.
Parameterization\.hθ\(x,τ\)h\_\{\\theta\}\(x,\\tau\)andgθ\(y,τ\)g\_\{\\theta\}\(y,\\tau\)are encoders that take an observation and a horizon as input\. They receive a learned embedding ofτ\\tau, so one network covers all horizons and an observation can have a different embedding at each\. The architecture depends on whether inputs are states or images\. We writeθ\\thetafor all encoder, horizon\-embedding, and coefficient parameters\.
Local dynamic model\.The temporal model describes where exploration can lead\. To evaluate an action, we also need a prediction of its immediate consequence\. A separately parameterized local model predicts
x^t\+1a=Fϕ\(xt,a\)\.\\widehat\{x\}\_\{t\+1\}^\{\\,a\}=F\_\{\\phi\}\(x\_\{t\},a\)\.\(2\)The hat distinguishes this prediction from the observation obtained after execution\. The temporal representation then evaluates the predicted outcome through its relations to the goal\.
### 2\.3Learning
Noise\-contrastive estimation\.At a given horizon, we distinguish observed temporal pairs from pairs formed by sampling the target observation independently fromp0p\_\{0\}\. These provide positive and negative examples, respectively\. The EBM in[Eq\.1](https://arxiv.org/html/2609.38383#S2.E1)expresses temporal association through the log\-density ratio
logpθ\(y\|x,τ\)p0\(y\)\\displaystyle\\log\\frac\{p\_\{\\theta\}\(y\|x,\\tau\)\}\{p\_\{0\}\(y\)\}=β0\(τ\)⟨hθ\(x,τ\),gθ\(y,τ\)⟩−logZθ\(x,τ\)\\displaystyle=\\beta\_\{0\}\(\\tau\)\\langle h\_\{\\theta\}\(x,\\tau\),g\_\{\\theta\}\(y,\\tau\)\\rangle\-\\log Z\_\{\\theta\}\(x,\\tau\)\(3\)≈β0\(τ\)⟨hθ\(x,τ\),gθ\(y,τ\)⟩\+β1\(τ\)=:Gθ\(x,y,τ\)\.\\displaystyle\\approx\\beta\_\{0\}\(\\tau\)\\langle h\_\{\\theta\}\(x,\\tau\),g\_\{\\theta\}\(y,\\tau\)\\rangle\+\\beta\_\{1\}\(\\tau\)=:G\_\{\\theta\}\(x,y,\\tau\)\.Whenp0p\_\{0\}is the target marginal, the ratio compares how likelyyyis to followxxat horizonτ\\tauwith how oftenyyoccurs in the data\. The learned offsetβ1\(τ\)\\beta\_\{1\}\(\\tau\)approximates−logZθ\(x,τ\)\-\\log Z\_\{\\theta\}\(x,\\tau\), allowing us to fitGθG\_\{\\theta\}by noise\-contrastive estimation \(NCE\) without evaluating the partition function\([Gutmann and Hyvärinen, 2010](https://arxiv.org/html/2609.38383#bib.bib2);[Ma and Collins, 2018](https://arxiv.org/html/2609.38383#bib.bib3)\)\. Because the offset is shared across sources, source\-wise normalization is approximate;[Sec\.B\.1](https://arxiv.org/html/2609.38383#A2.SS1)details this distinction\.
WithNNnegatives per positive and the same source marginal for both classes, the Bayes probability of a positive label is
pdata\(y\|x,τ\)pdata\(y\|x,τ\)\+Np0\(y\)=σ\(logpdata\(y\|x,τ\)p0\(y\)−logN\),\\frac\{p\_\{\\rm data\}\(y\|x,\\tau\)\}\{p\_\{\\rm data\}\(y\|x,\\tau\)\+Np\_\{0\}\(y\)\}=\\sigma\\\!\\left\(\\log\\frac\{p\_\{\\rm data\}\(y\|x,\\tau\)\}\{p\_\{0\}\(y\)\}\-\\log N\\right\),\(4\)whereσ\\sigmais the logistic sigmoid\. Fittingσ\(Gθ−logN\)\\sigma\(G\_\{\\theta\}\-\\log N\)therefore estimates the temporal log\-density ratio\. The reference enters through sampled negatives, so its density need not be evaluated\.
We train across horizons by samplingτ\\taufrom a chosen distributionp\(τ\)p\(\\tau\)over𝒯\\mathcal\{T\}and minimizing
ℒNCE\(θ\)=𝔼p\(τ\)\[\\displaystyle\\mathcal\{L\}\_\{\\rm NCE\}\(\\theta\)=\\mathbb\{E\}\_\{p\(\\tau\)\}\\\!\\Bigl\[−𝔼pdata\(x,y\|τ\)\[logσ\(Gθ\(x,y,τ\)−logN\)\]\\displaystyle\-\\mathbb\{E\}\_\{p\_\{\\rm data\}\(x,y\|\\tau\)\}\\\!\\left\[\\log\\sigma\\\!\\left\(G\_\{\\theta\}\(x,y,\\tau\)\-\\log N\\right\)\\right\]\(5\)−N𝔼pdata\(x\|τ\)p0\(y\)\[log\(1−σ\(Gθ\(x,y,τ\)−logN\)\)\]\]\.\\displaystyle\-N\\,\\mathbb\{E\}\_\{p\_\{\\rm data\}\(x\|\\tau\)p\_\{0\}\(y\)\}\\\!\\left\[\\log\\\!\\left\(1\-\\sigma\\\!\\left\(G\_\{\\theta\}\(x,y,\\tau\)\-\\log N\\right\)\\right\)\\right\]\\Bigr\]\.The two terms distinguish temporal pairs from independent pairs, using the same source marginalpdata\(x\|τ\)p\_\{\\rm data\}\(x\|\\tau\)\. All temporal\-model parameters are learned jointly across horizons\. Our objective shares a similar loss form with SigLIP\([Zhai et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib7)\)\.
At the unrestricted population optimum,G∗\(x,y,τ\)=log\[pdata\(y\|x,τ\)/p0\(y\)\]G^\{\*\}\(x,y,\\tau\)=\\log\[p\_\{\\rm data\}\(y\|x,\\tau\)/p\_\{0\}\(y\)\]where both densities are positive\. The constrained embedding model approximates this score, soeGθe^\{G\_\{\\theta\}\}estimates the conditional target density relative to the reference\.
Learning the local dynamics\.We train the local predictor on action\-labeled transitions at horizonτ=1\\tau=1\. Letpstep\(xt,at,xt\+1\)p\_\{\\rm step\}\(x\_\{t\},a\_\{t\},x\_\{t\+1\}\)denote their empirical distribution\. For state observations, we minimize
ℒdyn\(ϕ\)=𝔼pstep\(xt,at,xt\+1\)\[‖Fϕ\(xt,at\)−xt\+1‖22\]\.\\mathcal\{L\}\_\{\\rm dyn\}\(\\phi\)=\\mathbb\{E\}\_\{p\_\{\\rm step\}\(x\_\{t\},a\_\{t\},x\_\{t\+1\}\)\}\\\!\\left\[\\\|F\_\{\\phi\}\(x\_\{t\},a\_\{t\}\)\-x\_\{t\+1\}\\\|\_\{2\}^\{2\}\\right\]\.\(6\)The framework has a*hierarchical*organization: local dynamics predicts one\-step action outcomes, which the multi\-horizon representation evaluates against the goal across horizons\. Both are learned before planning and remain fixed at test time\.
### 2\.4Planning by Probability Improvement
Given a goal observationygy\_\{g\}, the planner favors actions whose predicted outcomes make the goal more likely\. At the current observationxtx\_\{t\}, the dynamics model predictsx^t\+1a=Fϕ\(xt,a\)\\hat\{x\}^\{a\}\_\{t\+1\}=F\_\{\\phi\}\(x\_\{t\},a\)for eacha∈𝒜\(xt\)a\\in\\mathcal\{A\}\(x\_\{t\}\)\. We define progress at horizonτ\\tauas
Δτ\(a\):=eGθ\(x^t\+1a,yg,τ\)−eGθ\(xt,yg,τ\)\.\\Delta\_\{\\tau\}\(a\):=e^\{G\_\{\\theta\}\(\\hat\{x\}^\{a\}\_\{t\+1\},y\_\{g\},\\tau\)\}\-e^\{G\_\{\\theta\}\(x\_\{t\},y\_\{g\},\\tau\)\}\.\(7\)By[Eq\.3](https://arxiv.org/html/2609.38383#S2.E3), this quantity estimates the increase in goal probability relative top0\(yg\)p\_\{0\}\(y\_\{g\}\),Δτ\(a\)≈\[pθ\(yg\|x^t\+1a,τ\)−pθ\(yg\|xt,τ\)\]/p0\(yg\)\\Delta\_\{\\tau\}\(a\)\\approx\[p\_\{\\theta\}\(y\_\{g\}\|\\hat\{x\}^\{a\}\_\{t\+1\},\\tau\)\-p\_\{\\theta\}\(y\_\{g\}\|x\_\{t\},\\tau\)\]/p\_\{0\}\(y\_\{g\}\)\. At a fixed horizon, the current\-state term is shared by all actions, so ranking actions byΔτ\(a\)\\Delta\_\{\\tau\}\(a\)is equivalent to ranking their predicted outcomes byeGθ\(x^t\+1a,yg,τ\)e^\{G\_\{\\theta\}\(\\hat\{x\}^\{a\}\_\{t\+1\},y\_\{g\},\\tau\)\}\.
Combining horizons\.The same move can make different progress at different horizons\. We combine these improvements with a temperatureβ\>0\\beta\>0, distinct fromβ0\(τ\)\\beta\_\{0\}\(\\tau\)andβ1\(τ\)\\beta\_\{1\}\(\\tau\):
at∗∈argmaxa∈𝒜\(xt\)βlog\[1\|𝒯\|∑τ∈𝒯exp\(Δτ\(a\)β\)\]\.a\_\{t\}^\{\*\}\\in\\argmax\_\{a\\in\\mathcal\{A\}\(x\_\{t\}\)\}\\beta\\log\\\!\\left\[\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\}\\exp\\\!\\left\(\\frac\{\\Delta\_\{\\tau\}\(a\)\}\{\\beta\}\\right\)\\right\]\.\(8\)Asβ→0\+\\beta\\to 0^\{\+\}, the objective becomesmaxτ∈𝒯Δτ\(a\)\\max\_\{\\tau\\in\\mathcal\{T\}\}\\Delta\_\{\\tau\}\(a\)\. The planner greedily picks the action and horizon with the largest improvement, so each decision comes with an explicit horizon selection\.
Asβ→∞\\beta\\to\\infty, the objective becomes\|𝒯\|−1∑τ∈𝒯Δτ\(a\)\|\\mathcal\{T\}\|^\{\-1\}\\sum\_\{\\tau\\in\\mathcal\{T\}\}\\Delta\_\{\\tau\}\(a\)\. Unlike the greedy rule, which uses only the best horizon, averaging lets gains and losses at different horizons offset each other\. The current\-state term cancels, so candidates are ranked by∑τ∈𝒯eGθ\(x^t\+1a,yg,τ\)\\sum\_\{\\tau\\in\\mathcal\{T\}\}e^\{G\_\{\\theta\}\(\\hat\{x\}^\{a\}\_\{t\+1\},y\_\{g\},\\tau\)\}, a single potential over states for a fixed goal\. With exact predictions and positive progress at every step, this potential increases along the trajectory, so no state is revisited\.
We refer to these two limits as*greedy planning*\(GP\) and*potential\-ascent planning*\(PAP\), respectively, throughout the paper\.
Closed\-loop execution\.At each step, it predicts the outcome of every candidate action, scores these outcomes by[Eq\.8](https://arxiv.org/html/2609.38383#S2.E8), and executesat∗a\_\{t\}^\{\*\}\. The agent then observes the actualxt\+1x\_\{t\+1\}and repeats, until it reaches the goal or exhausts its budget\. Because each prediction starts from an actual observation, errors do not compound over a multi\-step rollout, whileGθG\_\{\\theta\}provides long\-range guidance\. Bothθ\\thetaandϕ\\phiremain fixed throughout execution\.
[AppendixC](https://arxiv.org/html/2609.38383#A3)summarizes learningθ\\theta, learningϕ\\phi, and closed\-loop planning\.
## 3Theoretical Understanding
Figure 1:Emergent cognitive maps from random exploration\.\(A\)ScoreGθ\(x,yg,τ\)G\_\{\\theta\}\(x,y\_\{g\},\\tau\)over positionsxxin Large, with goalygy\_\{g\}fixed \(star\)\. Short horizons \(τ=16\\tau=16\) peak near the goal, while long horizons \(τ=512\\tau=512\) spread through connecting passages\.\(B\)t\-SNE ofhθ\(x,τ\)h\_\{\\theta\}\(x,\\tau\)in Giant atτ=1\\tau=1and40964096\. Atτ=1\\tau=1, embeddings preserve local geodesic geometry, while atτ=4096\\tau=4096, they reflect connectivity between distant regions\. 3D views appear in[Fig\.5](https://arxiv.org/html/2609.38383#A4.F5)\.\(C\)GP and PAP routes through Large and Giant\. Colour denotes the selected horizon for GP and the dominant score\-contributing horizon for PAP\. Matching markers pair starts \(∘\\circ\) and goals \(⋆\\star\)\.We study what temporal relations learned from random exploration reveal about an environment\. The results concern the calibrated scoreG∗\(x,y,τ\)=log\[pdata\(y\|x,τ\)/p0\(y\)\]G^\{\*\}\(x,y,\\tau\)=\\log\[p\_\{\\rm data\}\(y\|x,\\tau\)/p\_\{0\}\(y\)\], the ideal log\-density ratio our model approximates, and characterize the environmental structure it contains\.[SectionB\.1](https://arxiv.org/html/2609.38383#A2.SS1)relates this score to the learned model\.
From random walks to the heat equation\.Consider a symmetric local random walk over the free space, where a move into a wall leaves the agent in place\. In free space, its steps are independent with zero mean and covariance2αIn2\\alpha I\_\{n\}\. The constantα\>0\\alpha\>0is set by the one\-step transition kernel and does not depend on position or horizon\. Afterτ\\tausteps, the root\-mean\-square displacement is2nατ\\sqrt\{2n\\alpha\\tau\}, so doubling this diffusion scale takes four times as many steps\. For regular domains and lattice approximations satisfying a reflected\-Brownian invariance principle\([Burdzy and Chen, 2008](https://arxiv.org/html/2609.38383#bib.bib29)\), we study the limit under diffusive rescaling, with time per step proportional to squared step length\. Writingτ\\taualso for this rescaled time, with diffusion coefficientα\\alpha, its transition density satisfies∂τpdata=αΔypdata\\partial\_\{\\tau\}p\_\{\\rm data\}=\\alpha\\,\\Delta\_\{y\}p\_\{\\rm data\}\.
###### Proposition 1\(Short horizons encode geodesic geometry\)\.
Consider the reflected Brownian motion above on a smooth, bounded, connected domain, with uniformp0p\_\{0\}\. Letdgeo\(x,y\)d\_\{\\rm geo\}\(x,y\)be the geodesic distance, the infimum of feasible path lengths betweenxxandyy\. For fixed interiorx,yx,y,
limτ↓0−4ατG∗\(x,y,τ\)=dgeo\(x,y\)2\.\\lim\_\{\\tau\\downarrow 0\}\\,\-4\\alpha\\tau\\,G^\{\*\}\(x,y,\\tau\)=d\_\{\\rm geo\}\(x,y\)^\{2\}\.\(9\)
###### Proof sketch\.
Varadhan’s formula gives−4ατlogpdata\(y\|x,τ\)→dgeo\(x,y\)2\-4\\alpha\\tau\\log p\_\{\\rm data\}\(y\|x,\\tau\)\\to d\_\{\\rm geo\}\(x,y\)^\{2\}\([Varadhan, 1967](https://arxiv.org/html/2609.38383#bib.bib18);[Norris, 1997](https://arxiv.org/html/2609.38383#bib.bib19)\), and4ατlogp0\(y\)→04\\alpha\\tau\\log p\_\{0\}\(y\)\\to 0\. ∎
This limit concerns rescaled diffusion time, rather than integer horizons on a fixed lattice \([Sec\.B\.2](https://arxiv.org/html/2609.38383#A2.SS2)\)\. For the limiting diffusion, targets with shorter geodesic distances have larger scores at sufficiently short horizons, whatever the Euclidean distances\. A nearby point behind a wall may require a long detour\. At short horizons it then scores below a point that is farther in Euclidean distance but reachable along a shorter open route\.
We now return to the discrete walk on a finite connected state space, with transition matrixPPand stationary distributionπ\\pi, which is uniform becausePPis symmetric\.
###### Proposition 2\(Long horizons encode multiscale connectivity\)\.
SupposePPis irreducible, aperiodic, and reversible, with eigenvalues1=λ1≥λ2≥⋯1=\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdotsand eigenfunctionsψj\\psi\_\{j\}orthonormal underπ\\pi\. Setp0=πp\_\{0\}=\\pi\. For every integerτ≥1\\tau\\geq 1and every pair withPτ\(x,y\)\>0P^\{\\tau\}\(x,y\)\>0,
G∗\(x,y,τ\)=log\[1\+∑j≥2λjτψj\(x\)ψj\(y\)\]\.G^\{\*\}\(x,y,\\tau\)=\\log\\Big\[1\+\\sum\_\{j\\geq 2\}\\lambda\_\{j\}^\{\\tau\}\\,\\psi\_\{j\}\(x\)\\psi\_\{j\}\(y\)\\Big\]\.\(10\)
###### Proof sketch\.
DividePτ\(x,y\)=π\(y\)∑jλjτψj\(x\)ψj\(y\)P^\{\\tau\}\(x,y\)=\\pi\(y\)\\sum\_\{j\}\\lambda\_\{j\}^\{\\tau\}\\psi\_\{j\}\(x\)\\psi\_\{j\}\(y\)byπ\(y\)\\pi\(y\), separateψ1=1\\psi\_\{1\}=1, and take the logarithm\([Aldous and Fill, 2002](https://arxiv.org/html/2609.38383#bib.bib20);[Coifman and Lafon, 2006](https://arxiv.org/html/2609.38383#bib.bib11)\)\. ∎
Terms with smaller\|λj\|\|\\lambda\_\{j\}\|decay faster\. Modes near\+1\+1preserve differences between regions that exchange probability slowly\. Picture two rooms joined by a narrow doorway\. When movement between the rooms is slow relative to mixing within each room, a walk forgets which corner it started in long before it forgets which room it started in\. Over these horizons,G∗G^\{\*\}distinguishes the rooms more strongly than positions within each room \([Sec\.B\.3](https://arxiv.org/html/2609.38383#A2.SS3)\)\.
The useful scale changes during planning\.Choosing a horizon is like choosing the zoom level of a map\. Zoomed in, the map shows the next turn but may leave the connecting passage out of view\. Zoomed out, the passage appears, but the turns leading to it blur together\. A navigator zooms out when the destination is far and zooms in as it approaches, and the horizon plays the same role\. Progress compares the true goal probability from each candidate outcome with that from the current state\. At a horizon too short for the walk to reach the goal, these probabilities are all nearly zero, so every candidate makes almost no progress\. As the horizon goes to infinity, they all approach the stationary valueπ\(yg\)\\pi\(y\_\{g\}\), and progress again vanishes\. Only intermediate horizons separate good moves from bad ones\. In free space, the best horizon for a small step is roughly the one at which the walk’s typical displacement reaches the goal, so it grows with the squared goal distance \([Sec\.B\.2](https://arxiv.org/html/2609.38383#A2.SS2)\)\. In a maze, it also depends on the time needed to cross connecting passages\. Empirically, the greedy planner selects shorter horizons when approaching the goal \([Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)C\)\.
## 4Experiments
We evaluate goal reaching, horizon selection, data efficiency, first\-person observations, and manipulation\.[Section4\.1](https://arxiv.org/html/2609.38383#S4.SS1)specifies the data and evaluation conditions and[Sec\.4\.2](https://arxiv.org/html/2609.38383#S4.SS2)reports the results\.
### 4\.1Experimental Setup
Datasets and Environments\.We study three settings\.*\(1\) Maze navigation:*OGBench PointMaze Large and Giant\([Park et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib13)\), with position or×6464\\\!\\times\\\!64overhead images\. Our main models use 4M lattice transitions: eight neighbouring moves or stay sampled uniformly on a 1\-unit grid, with blocked moves replaced by stay and five simulator steps per transition\. Reproduced baselines use 1M transitions from uniformly sampled actions \(20 episodes of 50k steps\)\. Our local dynamics uses 1M transitions on Large and 10M on Giant\.Navigateuses OGBench expert data\.*\(2\) Egocentric navigation:*a first\-person Large maze and Habitat’s Van Gogh Room and Apartment 1\([Savva et al\., 2019](https://arxiv.org/html/2609.38383#bib.bib17)\)\. Maze inputs are images with optional position and/or heading\. Habitat models use one 8M\-step walk per scene; lattice planning compares position alone with images and position, and continuous planning adds heading\.*\(3\) Manipulation:*OGBench cube\-single with expert\-only and mixed data\. The expert\-only setting follows LeWM’s data and evaluation protocol\([Maes et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib25)\)and uses its reported baseline results\. The mixed setting combines expert trajectories with noisy oracle trajectories, generated by adding temporally correlated noise to the oracle’s waypoint plan and replacing each action with a random one with probability0\.10\.1\.
Evaluation\.We report success rate \(SR\) and success weighted by path length \(SPL\)\([Anderson et al\., 2018](https://arxiv.org/html/2609.38383#bib.bib26)\):
SR=1M∑i=1MSi,SPL=1M∑i=1MSiℓimax\(pi,ℓi\),\\mathrm\{SR\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}S\_\{i\},\\qquad\\mathrm\{SPL\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}S\_\{i\}\\frac\{\\ell\_\{i\}\}\{\\max\(p\_\{i\},\\ell\_\{i\}\)\},\(11\)HereMMcounts episodes,SiS\_\{i\}indicates success,pip\_\{i\}is executed path length, andℓi\\ell\_\{i\}is shortest\-path length to the goal region\. Setting \(1\) uses five official tasks \(20 episodes each, 1k\-simulator\-step budget\) and 20 fixed start–goal pairs per maze \(geodesic separation≥20\\geq 20cells, 5k\-simulator\-step budget\)\. Success requires reaching within one unit of the goal\. Candidate outcomes come from learned local dynamics, without simulator queries before action execution\. Setting \(2\) reuses the maze tasks\. Habitat evaluates its two most distant reachable corner pairs in both directions: 25 episodes per task, random initial headings, and a 5k\-lattice\-action budget\. Success requires reaching within one lattice move of the goal or within0\.30\.3m in continuous evaluation\. In setting \(3\), the goal is the observation 25 steps after the initial state in the same dataset trajectory\. Each episode allows 50 steps and succeeds if the cube center comes within0\.040\.04m of its goal position at any step\. We report SR only\. Mixed\-data comparisons with LeWM use 1,000 episodes from the held\-out split\.
Comparisons\.*World models*\(LeWM, DINO\-WM\) plan through learned dynamics using their own planners\.*Offline goal\-conditioned RL*\(GCIQL, HIQL, HILP, QRL\) executes policies learned from offline data\. Random actions and a shortest\-path oracle provide reference controls\.
### 4\.2Results
We first show that random exploration supports long\-range planning from both states and images, and that the learned representation acquires multiscale spatial structure consistent with the random\-walk analysis \([Tab\.1](https://arxiv.org/html/2609.38383#S4.T1),[Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)\)\. We then study how the available horizons and exploration coverage shape planning performance \([Fig\.2](https://arxiv.org/html/2609.38383#S4.F2),[Tab\.3](https://arxiv.org/html/2609.38383#S4.T3)\)\. Finally, we test planning from egocentric observations and manipulation with suboptimal data \([Fig\.3](https://arxiv.org/html/2609.38383#S4.F3),[Tab\.3](https://arxiv.org/html/2609.38383#S4.T3)\)\.
Table 1:PointMaze goal reaching from states \(A\) and images \(B\) on official tasks and random start–goal pairs\. The GCIQL, QRL, and HIQLNavigateresults are quoted from[Park et al\. \(2025\)](https://arxiv.org/html/2609.38383#bib.bib13)\.A State
B Vision
Planning from random exploration\.Temporal relations learned from random exploration support long\-range goal reaching from both states and images \([Tab\.1](https://arxiv.org/html/2609.38383#S4.T1)\)\. Averaged over three training seeds, GP and PAP reach0\.970\.97and0\.920\.92of the random start–goal pairs on Large and0\.850\.85of them on Giant from states\. From images, GP and PAP reach0\.950\.95of the random pairs on Large with SPL0\.850\.85and0\.830\.83, and0\.780\.78and0\.700\.70of them on Giant\. Most reported baselines achieve little or no success on the official tasks in the random\-data setting\. Unlike goal\-conditioned RL, our model learns temporal relations under the exploration behavior without policy\-improvement training\. These relations provide long\-range guidance for local action selection under both planning objectives\.
Emergent cognitive maps\.The learned temporal representation exhibits properties of a multiscale cognitive map\([Tolman, 1948](https://arxiv.org/html/2609.38383#bib.bib22)\)in its scores, embeddings, and planned routes\. For a fixed goal, short\-horizon embedding similarities⟨hθ\(x,τ\),gθ\(yg,τ\)⟩\\langle h\_\{\\theta\}\(x,\\tau\),g\_\{\\theta\}\(y\_\{g\},\\tau\)\\rangleemphasize nearby locations, while longer horizons reveal connections through passages \([Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)A\)\. The embeddingshθ\(x,τ\)h\_\{\\theta\}\(x,\\tau\)represent the same environment at different spatial scales \([Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)B\)\. Distance probes quantify this change: shorter horizons better preserve local geodesic geometry, while longer horizons better reflect connectivity between distant states \([Tab\.4](https://arxiv.org/html/2609.38383#A4.T4)\)\. These patterns are consistent with the random\-walk analysis in[Sec\.3](https://arxiv.org/html/2609.38383#S3)and emerge without distance or route supervision\. In planning, GP and PAP use this structure to turn random exploration into efficient, goal\-directed routes through the maze \([Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)C\)\.
Horizon selection and aggregation\.The larger maze benefits from longer planning horizons \([Fig\.2](https://arxiv.org/html/2609.38383#S4.F2)\): Large reaches about0\.80\.8success atτmax=64\\tau\_\{\\max\}=64, whereas Giant improves mainly as the range grows from256256to about1,0241\{,\}024\. Further extension does not consistently improve performance\. GP selects the largest progress at any horizon, while PAP aggregates progress across horizons\. The two achieve comparable success on random maze pairs \([Tab\.1](https://arxiv.org/html/2609.38383#S4.T1)\), while PAP achieves higher success in the cube\-single comparison \(0\.920\.92versus0\.840\.84;[Tab\.3](https://arxiv.org/html/2609.38383#S4.T3)\)\. Along the routes in[Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)C, GP’s selected horizon and PAP’s dominant score\-contributing horizon generally shorten near the goal, although PAP continues to aggregate all horizons\. This change occurs without a prescribed schedule and is consistent with the spatial\-scale analysis in[Sec\.3](https://arxiv.org/html/2609.38383#S3)\.
Figure 2:Official\-task success versus maximum planning horizon\. Color distinguishes Large and Giant and line style distinguishes GP and PAP\. Lines show three\-seed means and bands one standard deviation\. Planning uses every integer horizon from11toτmax\\tau\_\{\\max\}\.Table 2:Goal reaching vs\. walk length and exploration coveragecc\. Full coverage \(c=1c=1\) repeats the 4M reference\.
Table 3:Cube\-single: mixed data, held\-out episodes \(top\); expert data, LeWM’s protocol and baselines \(bottom\)\.
Exploration coverage\.Every walk in the length rows visits every free cell \(a 0\.1M\-step walk already passes each Large cell about 136 times and each Giant cell about 73 times\), so these rows vary the number of transitions per cell rather than coverage\. The coverage fractionccspecifies the part of the maze available to a 4M\-step walk\. Spreads denote standard deviations\. The 4M andc=1c=1rows are the models of[Tab\.1](https://arxiv.org/html/2609.38383#S4.T1)\. Every row plans with all horizons from 1 to 8192 through the learned local dynamics\. Both data volume and state coverage affect planning performance\. On Large, increasing the explored fraction fromc=0\.25c=0\.25toc=0\.75c=0\.75raises success from0\.090\.09to0\.560\.56for GP and from0\.270\.27to0\.590\.59for PAP, and full coverage reaches0\.920\.92and0\.750\.75\([Tab\.3](https://arxiv.org/html/2609.38383#S4.T3)\)\. With unrestricted exploration, GP on Large rises from0\.450\.45after0\.10\.1M steps to0\.910\.91after11M steps\. Giant needs longer walks: GP reaches0\.010\.01after0\.10\.1M steps,0\.490\.49after11M and0\.760\.76after44M\. Planning therefore improves with both walk length and coverage, and the larger maze requires more exploration\.
Figure 3:An egocentric plan in Habitat Apartment 1 by the model that plans from the image and the position \(horizons capped at 512, no constraint on consecutive headings\)\. Left: the route, coloured by the selected horizon, from the start \(∘\\circ\) to the goal \(⋆\\star\)\. Right: the agent’s view at ten decisions and the horizon selected there\.Egocentric observations\.Planning extends to first\-person observations\. In Habitat, we train on one 8M\-step walk per scene, collected with30∘30^\{\\circ\}turns and0\.50\.5m forward or backward steps on a0\.250\.25m lattice\. On Apartment 1’s four corner tasks, images and position achieve0\.71±0\.120\.71\\pm 0\.12success atτmax=8,192\\tau\_\{\\max\}=8\{,\}192, compared with0\.380\.38for a random policy\. The position\-only planner repeatedly turns in place and reaches no goals\. Under the continuous protocol \(36 headings, walking0\.750\.75m to the scored candidate before re\-planning, no constraint on consecutive headings\), the image\-and\-position model reaches a far pair of the apartment from all three start headings with horizons capped at512512\(SR1\.001\.00, SPL0\.770\.77; position alone: SR1\.001\.00, SPL0\.790\.79\)\. The full horizon range strands the image model, since the longest horizons dominate at headings unseen in training, as on Giant beyondτmax≈1,000\\tau\_\{\\max\}\\approx 1\{,\}000\([Fig\.2](https://arxiv.org/html/2609.38383#S4.F2)\)\. Along a successful route, the selected horizon shortens as the goal comes into view \([Fig\.3](https://arxiv.org/html/2609.38383#S4.F3)\)\. In the egocentric maze, we find position helps more than heading\.
Manipulation with suboptimal data\.Our model supports effective manipulation planning from suboptimal trajectories\. On cube\-single under LeWM’s evaluation protocol, PAP reaches success0\.920\.92while GP reaches0\.840\.84\. We also vary the expert fraction at a fixed training budget and compare PAP with LeWM on matched data \([Tab\.3](https://arxiv.org/html/2609.38383#S4.T3)\)\. Our model achieves higher success at every mixture\. Noisy\-oracle data alone yields success0\.9080\.908, exceeding0\.8760\.876with expert\-only data, and the highest reported success is0\.9330\.933at25%25\\%expert data\. Our temporal model learns how states are connected from the observed trajectories\. Noise may expose a wider range of these connections, making suboptimal trajectories useful for planning even when they are less efficient at reaching goals\.
## 5Related Work
World models and temporal representations\.Latent world models support goal reaching through action\-sequence optimization\. DINO\-WM predicts dynamics in pretrained visual features\([Zhou et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib15)\), while Temporal Straightening jointly learns representations and dynamics with a curvature regularizer\([Wang et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib16)\)\. Temporal representation learning provides another source of structure\. Contrastive predictive coding discriminates future observations\([van den Oord et al\., 2018](https://arxiv.org/html/2609.38383#bib.bib12)\), and Contrastive RL connects temporal associations to goal\-conditioned values\([Eysenbach et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib4)\)\. Successor and forward–backward representations relate future occupancy to rewards\([Dayan, 1993](https://arxiv.org/html/2609.38383#bib.bib1);[Touati and Ollivier, 2021](https://arxiv.org/html/2609.38383#bib.bib5)\), while contrastive representations support planning through interpolation and temporal reasoning\([Eysenbach et al\., 2024](https://arxiv.org/html/2609.38383#bib.bib8);[Ziarko et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib9)\)\. Our model separates long\-range guidance learned from observation pairs from local action\-outcome prediction\.
Random\-walk geometry and planning\.Diffusion maps connect transition statistics to geometry\([Coifman and Lafon, 2006](https://arxiv.org/html/2609.38383#bib.bib11)\)\.[Zhao et al\. \(2025b\)](https://arxiv.org/html/2609.38383#bib.bib14)model place\-cell populations through position embeddings fitted to predefined random\-walk kernels\. We learn conditional temporal log\-density ratios from sampled observation pairs through NCE, using horizon\-conditioned state or image encoders\. These scores guide action selection through horizon\-specific progress in GP or a goal\-dependent potential in PAP\. Quasimetric learning targets directed optimal goal distances\([Wang et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib6)\), while replay\-buffer search composes local connections\([Eysenbach et al\., 2019](https://arxiv.org/html/2609.38383#bib.bib10)\)\. Our temporal scores describe exploration behavior, with geometric and connectivity interpretations under the random\-walk assumptions in[Sec\.3](https://arxiv.org/html/2609.38383#S3)\.
## 6Limitations and Future Work
Random exploration can teach representations that support goal\-directed behavior without policy\-improvement training\. Our current models learn within individual environments\. Planning in unseen environments raises the question of what temporal structure transfers and what must be learned from new experience\. Behavioral timescale synaptic plasticity supports rapid place\-field formation\([Bittner et al\., 2017](https://arxiv.org/html/2609.38383#bib.bib76)\)and may inspire adaptation from limited experience\. Egocentric planning presents a related challenge: a single view may not determine the state\. Observation histories could help recover spatial context and reduce reliance on explicit position inputs\. Scaling training to larger datasets across diverse environments would let us study whether shared temporal representations support transfer and rapid adaptation\.
The connection to animal cognition also remains incomplete\. Although the embeddings are learned, the learning objective, horizon conditioning, and planning rules are designed explicitly\. Our framework does not explain how animals acquire these mechanisms\. A broader account would need to explain how evolution shapes general learning and control mechanisms through which both representations and planning strategies emerge from experience\.
## Acknowledgments
We thank Chenxin Tao for insightful discussions during the spring of 2025, and Minglu Zhao and Dehong Xu for earlier collaborations\. Y\. W\. is partially supported by NSF DMS\-2415226, DARPA W912CG25CA007 and research gift funds from Amazon and Qualcomm\. T\. G\. is supported by NSF under Award No\. 2610649 and by NERSC through DDR\-ERCAP0035256\.
## References
- A\. Ajay, Y\. Du, A\. Gupta, J\. Tenenbaum, T\. Jaakkola, and P\. AgrawalIs conditional generative modeling all you need for decision\-making?\.arXiv preprint arXiv:2211\.15657\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Aldous and Fill \(2002\)D\. Aldous and J\. A\. FillReversible markov chains and random walks on graphs\.Note:Unfinished monograph, recompiled 2014, available at[http://www\.stat\.berkeley\.edu/users/aldous/RWG/book\.html](http://www.stat.berkeley.edu/users/aldous/RWG/book.html)Cited by:[§B\.3](https://arxiv.org/html/2609.38383#A2.SS3.p2.2.1),[§B\.3](https://arxiv.org/html/2609.38383#A2.SS3.p3.4),[§B\.6](https://arxiv.org/html/2609.38383#A2.SS6.p3.2.1),[§3](https://arxiv.org/html/2609.38383#S3.p6.1.1)\.
- Andersonet al\.\(2018\)P\. Anderson, A\. Chang, D\. S\. Chaplot, A\. Dosovitskiy, S\. Gupta, V\. Koltun, J\. Kosecka, J\. Malik, R\. Mottaghi, M\. Savva,et al\.On evaluation of embodied navigation agents\.arXiv preprint arXiv:1807\.06757\.Cited by:[§4\.1](https://arxiv.org/html/2609.38383#S4.SS1.p2.1)\.
- Assranet al\.\(2025\)M\. Assran, A\. Bardes, D\. Fan, Q\. Garrido, R\. Howes, M\. Komeili, M\. Muckley, A\. Rizvi, C\. Roberts, K\. Sinha, A\. Zholus, S\. Arnaud, A\. Gejji, A\. Martin, F\. Robert Hogan, D\. Dugas, P\. Bojanowski, V\. Khalidov, P\. Labatut, F\. Massa, M\. Szafraniec, K\. Krishnakumar, Y\. Li, X\. Ma, S\. Chandar, F\. Meier, Y\. LeCun, M\. Rabbat, and N\. BallasV\-JEPA 2: self\-supervised video models enable understanding, prediction and planning\.arXiv preprint arXiv:2506\.09985\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Bittneret al\.\(2017\)K\. C\. Bittner, A\. D\. Milstein, C\. Grienberger, S\. Romani, and J\. C\. MageeBehavioral time scale synaptic plasticity underlies CA1 place fields\.Science\.Cited by:[§6](https://arxiv.org/html/2609.38383#S6.p1.1)\.
- Boneyet al\.\(2020\)R\. Boney, J\. Kannala, and A\. IlinRegularizing model\-based planning with energy\-based models\.InCoRL,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Burdzy and Chen \(2008\)K\. Burdzy and Z\. ChenDiscrete approximations to reflected Brownian motion\.The Annals of Probability\.Cited by:[§B\.2](https://arxiv.org/html/2609.38383#A2.SS2.p1.1),[§3](https://arxiv.org/html/2609.38383#S3.p2.1)\.
- Caselliet al\.\(2026\)N\. Caselli, F\. Massafra, S\. Punzo, S\. L\. Sardo, I\. Pantelidis, and S\. K\. BhethanabhotlaMind the gap: promises and pitfalls of hierarchical planning in leworldmodel\.arXiv preprint arXiv:2607\.12547\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Chenget al\.\(2025\)S\. Cheng, D\. Kong, J\. Xie, K\. Lee, Y\. N\. Wu, and Y\. YangLatent space energy\-based neural ODEs\.TMLR\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Coifman and Lafon \(2006\)R\. R\. Coifman and S\. LafonDiffusion maps\.Applied and Computational Harmonic Analysis\.Cited by:[§B\.3](https://arxiv.org/html/2609.38383#A2.SS3.p2.2.1),[§B\.4](https://arxiv.org/html/2609.38383#A2.SS4.p5.1),[§1](https://arxiv.org/html/2609.38383#S1.p2.1),[§3](https://arxiv.org/html/2609.38383#S3.p6.1.1),[§5](https://arxiv.org/html/2609.38383#S5.p2.1)\.
- Dayan \(1993\)P\. DayanImproving generalization for temporal difference learning: the successor representation\.Neural Computation\.Cited by:[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Duet al\.\(2020\)Y\. Du, T\. Lin, and I\. MordatchModel\-based planning with energy\-based models\.InCoRL,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Eysenbachet al\.\(2019\)B\. Eysenbach, R\. Salakhutdinov, and S\. LevineSearch on the replay buffer: bridging planning and reinforcement learning\.InNeurIPS,Cited by:[§5](https://arxiv.org/html/2609.38383#S5.p2.1)\.
- Eysenbachet al\.\(2024\)B\. Eysenbach, V\. Myers, R\. Salakhutdinov, and S\. LevineInference via interpolation: contrastive representations provably enable planning and inference\.InNeurIPS,Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p3.1),[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Eysenbachet al\.\(2022\)B\. Eysenbach, T\. Zhang, S\. Levine, and R\. R\. SalakhutdinovContrastive learning as goal\-conditioned reinforcement learning\.InNeurIPS,Cited by:[§B\.6](https://arxiv.org/html/2609.38383#A2.SS6.p5.1),[§1](https://arxiv.org/html/2609.38383#S1.p3.1),[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Gao and Xu \(2026\)Y\. Gao and X\. XuFast LeWorldModel\.arXiv preprint arXiv:2606\.26217\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Gkanatsioset al\.\(2023\)N\. Gkanatsios, A\. Jain, Z\. Xian, Y\. Zhang, C\. G\. Atkeson, and K\. FragkiadakiEnergy\-based models are zero\-shot planners for compositional scene rearrangement\.InRSS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Gutmann and Hyvärinen \(2010\)M\. Gutmann and A\. HyvärinenNoise\-contrastive estimation: a new estimation principle for unnormalized statistical models\.InAISTATS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1),[§B\.1](https://arxiv.org/html/2609.38383#A2.SS1.p2.2),[§1](https://arxiv.org/html/2609.38383#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.38383#S2.SS3.p1.3)\.
- Ha and Schmidhuber \(2018\)D\. Ha and J\. SchmidhuberRecurrent world models facilitate policy evolution\.InNeurIPS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Hafneret al\.\(2020\)D\. Hafner, T\. Lillicrap, J\. Ba, and M\. NorouziDream to control: learning behaviors by latent imagination\.InICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Hafneret al\.\(2019\)D\. Hafner, T\. Lillicrap, I\. Fischer, R\. Villegas, D\. Ha, H\. Lee, and J\. DavidsonLearning latent dynamics for planning from pixels\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Hafneret al\.\(2025\)D\. Hafner, J\. Pasukonis, J\. Ba, and T\. LillicrapMastering diverse control tasks through world models\.Nature\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Hansenet al\.\(2022\)N\. A\. Hansen, H\. Su, and X\. WangTemporal difference learning for model predictive control\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Hansenet al\.\(2024\)N\. Hansen, H\. Su, and X\. WangTD\-MPC2: scalable, robust world models for continuous control\.InICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Huoet al\.\(2026\)Y\. Huo, Z\. Song, and Y\. LuoFlow\-JEPA: flow matching for robust latent dynamics in JEPA world models\.arXiv preprint arXiv:2608\.29029\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Janneret al\.\(2022\)M\. Janner, Y\. Du, J\. Tenenbaum, and S\. LevinePlanning with diffusion for flexible behavior synthesis\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Konget al\.\(2024a\)D\. Kong, Y\. Huang, J\. Xie, E\. Honig, M\. Xu, S\. Xue, P\. Lin, S\. Zhou, S\. Zhong, N\. Zheng, and Y\. N\. WuMolecule design by latent prompt transformer\.InNeurIPS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Konget al\.\(2023\)D\. Kong, B\. Pang, T\. Han, and Y\. N\. WuMolecule design by latent space energy\-based modeling and gradual distribution shifting\.InUAI,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Konget al\.\(2024b\)D\. Kong, D\. Xu, M\. Zhao, B\. Pang, J\. Xie, A\. Lizarraga, Y\. Huang, S\. Xie, and Y\. N\. WuLatent plan transformer for trajectory abstraction: planning as latent space inference\.InNeurIPS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Konget al\.\(2026\)D\. Kong, M\. Zhao, A\. Qin, B\. Pang, C\. Tao, D\. Hartmann, E\. Honig, D\. Xu, A\. H\. Kumar, M\. Sarte, C\. Li, J\. Xie, and Y\. N\. WuInference\-time rethinking with latent thought vectors for math reasoning\.InWorkshop on Latent & Implicit Thinking – Going Beyond CoT Reasoning, ICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Konget al\.\(2025\)D\. Kong, M\. Zhao, D\. Xu, B\. Pang, S\. Wang, E\. Honig, Z\. Si, C\. Li, J\. Xie, S\. Xie, and Y\. N\. WuLatent thought models with variational Bayes inference\-time computation\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Kostrikovet al\.\(2022\)I\. Kostrikov, A\. Nair, and S\. LevineOffline reinforcement learning with implicit Q\-learning\.InICLR,Cited by:[Table 1](https://arxiv.org/html/2609.38383#S4.T1.5.7.1)\.
- LeCunet al\.\(2006\)Y\. LeCun, S\. Chopra, R\. Hadsell, M\. Ranzato, F\. Huang,et al\.A tutorial on energy\-based learning\.Predicting structured data1\(0\)\.Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p4.1)\.
- Liuet al\.\(2026\)M\. Liu, Y\. Huang, Z\. Liang, and X\. GaoToward physically grounded jepa world models for goal\-conditioned robotic planning\.arXiv preprint arXiv:2609\.03565\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Luoet al\.\(2024\)Y\. Luo, C\. Sun, J\. B\. Tenenbaum, and Y\. DuPotential based diffusion motion planning\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Ma and Collins \(2018\)Z\. Ma and M\. CollinsNoise contrastive estimation and negative sampling for conditional models: consistency and statistical efficiency\.InEMNLP,Cited by:[§B\.1](https://arxiv.org/html/2609.38383#A2.SS1.p2.2),[§2\.3](https://arxiv.org/html/2609.38383#S2.SS3.p1.3)\.
- Maeset al\.\(2026\)L\. Maes, Q\. Le Lidec, D\. Scieur, Y\. LeCun, and R\. BalestrieroLeWorldModel: stable end\-to\-end joint\-embedding predictive architecture from pixels\.arXiv preprint arXiv:2603\.19312\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.38383#S4.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.38383#S4.T1.5.11.1),[Table 1](https://arxiv.org/html/2609.38383#S4.T1.7.4.1),[Table 3](https://arxiv.org/html/2609.38383#S4.T3.fig2.1.8.1)\.
- Micheliet al\.\(2023\)V\. Micheli, E\. Alonso, and F\. FleuretTransformers are sample\-efficient world models\.InICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Nohet al\.\(2025\)D\. Noh, D\. Kong, M\. Zhao, A\. Lizarraga, J\. Xie, Y\. N\. Wu, and D\. HongLatent adaptive planner for dynamic manipulation\.InCoRL,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Norris \(1997\)J\. R\. NorrisHeat kernel asymptotics and the distance function in Lipschitz Riemannian manifolds\.Acta Mathematica\.Cited by:[§B\.2](https://arxiv.org/html/2609.38383#A2.SS2.p2.2.1),[§1](https://arxiv.org/html/2609.38383#S1.p2.1),[§3](https://arxiv.org/html/2609.38383#S3.p3.1.1)\.
- O’Keefe and Dostrovsky \(1971\)J\. O’Keefe and J\. DostrovskyThe hippocampus as a spatial map\. preliminary evidence from unit activity in the freely\-moving rat\.Brain Research\.Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p1.1)\.
- Panget al\.\(2020\)B\. Pang, T\. Han, E\. Nijkamp, S\. Zhu, and Y\. N\. WuLearning latent space energy\-based prior model\.InNeurIPS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Parket al\.\(2025\)S\. Park, K\. Frans, B\. Eysenbach, and S\. LevineOGBench: benchmarking offline goal\-conditioned RL\.InICLR,Cited by:[§4\.1](https://arxiv.org/html/2609.38383#S4.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.38383#S4.T1)\.
- Parket al\.\(2023\)S\. Park, D\. Ghosh, B\. Eysenbach, and S\. LevineHIQL: offline goal\-conditioned RL with latent states as actions\.InNeurIPS,Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p3.1),[Table 1](https://arxiv.org/html/2609.38383#S4.T1.5.9.1)\.
- Parket al\.\(2024\)S\. Park, T\. Kreiman, and S\. LevineFoundation policies with Hilbert representations\.InICML,Cited by:[Table 1](https://arxiv.org/html/2609.38383#S4.T1.5.10.1)\.
- Pham and Bera \(2026\)P\. Pham and A\. BeraLatent energy action planning with world models\.arXiv preprint arXiv:2609\.03294\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Qinet al\.\(2025\)A\. Qin, D\. Kong, W\. Wang, Y\. N\. Wu, S\. Zhu, and S\. XieGenerative actor critic\.arXiv preprint arXiv:2512\.21527\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Rajuet al\.\(2024\)R\. V\. Raju, J\. S\. Guntupalli, G\. Zhou, C\. Wendelken, M\. Lázaro\-Gredilla, and D\. GeorgeSpace is a latent sequence: a theory of the hippocampus\.Science Advances\.Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p1.1)\.
- Savvaet al\.\(2019\)M\. Savva, A\. Kadian, O\. Maksymets, Y\. Zhao, E\. Wijmans, B\. Jain, J\. Straub, J\. Liu, V\. Koltun, J\. Malik, D\. Parikh, and D\. BatraHabitat: a platform for embodied AI research\.InICCV,Cited by:[§4\.1](https://arxiv.org/html/2609.38383#S4.SS1.p1.1)\.
- Schrittwieseret al\.\(2020\)J\. Schrittwieser, I\. Antonoglou, T\. Hubert, K\. Simonyan, L\. Sifre, S\. Schmitt, A\. Guez, E\. Lockhart, D\. Hassabis, T\. Graepel, T\. Lillicrap, and D\. SilverMastering Atari, Go, chess and shogi by planning with a learned model\.Nature\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Sobalet al\.\(2025\)V\. Sobal, W\. Zhang, K\. Cho, R\. Balestriero, T\. G\. J\. Rudner, and Y\. LeCunLearning from reward\-free offline data: a case for planning with latent dynamics models\.arXiv preprint arXiv:2502\.14819\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2609.38383#S4.T3.fig2.1.7.1)\.
- Steinerberger \(2018\)S\. SteinerbergerVaradhan asymptotics for the heat kernel on finite graphs\.arXiv preprint arXiv:1801\.02183\.Cited by:[§B\.2](https://arxiv.org/html/2609.38383#A2.SS2.p4.2)\.
- Sunet al\.\(2025\)G\. Sun, M\. Jin, Z\. Wang, C\. Wang, S\. Ma, Q\. Wang, T\. Geng, Y\. N\. Wu, Y\. Zhang, and D\. LiuVisual agents as fast and slow thinkers\.InICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Tolman and Honzik \(1930\)E\. C\. Tolman and C\. H\. HonzikIntroduction and removal of reward, and maze performance in rats\.University of California Publications in Psychology\.Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p1.1)\.
- Tolman \(1948\)E\. C\. TolmanCognitive maps in rats and men\.\.Psychological review\.Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.38383#S2.SS2.p2.1),[§4\.2](https://arxiv.org/html/2609.38383#S4.SS2.p3.1)\.
- Touati and Ollivier \(2021\)A\. Touati and Y\. OllivierLearning one representation to optimize all rewards\.InNeurIPS,Cited by:[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Urainet al\.\(2023\)J\. Urain, N\. Funk, J\. Peters, and G\. ChalvatzakiSE\(3\)\-DiffusionFields: learning smooth cost functions for joint grasp and motion optimization through diffusion\.InICRA,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Urainet al\.\(2022\)J\. Urain, A\. T\. Le, A\. Lambert, G\. Chalvatzaki, B\. Boots, and J\. PetersLearning implicit priors for motion optimization\.InIROS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- van den Oordet al\.\(2018\)A\. van den Oord, Y\. Li, and O\. VinyalsRepresentation learning with contrastive predictive coding\.arXiv preprint arXiv:1807\.03748\.Cited by:[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Varadhan \(1967\)S\. R\. S\. VaradhanOn the behavior of the fundamental solution of the heat equation with variable coefficients\.Communications on Pure and Applied Mathematics\.Cited by:[§B\.2](https://arxiv.org/html/2609.38383#A2.SS2.p2.2.1),[§1](https://arxiv.org/html/2609.38383#S1.p2.1),[§3](https://arxiv.org/html/2609.38383#S3.p3.1.1)\.
- Wanget al\.\(2023\)T\. Wang, A\. Torralba, P\. Isola, and A\. ZhangOptimal goal\-reaching reinforcement learning via quasimetric learning\.InICML,Cited by:[Table 1](https://arxiv.org/html/2609.38383#S4.T1.5.8.1),[§5](https://arxiv.org/html/2609.38383#S5.p2.1)\.
- Wanget al\.\(2026\)Y\. Wang, O\. Bounou, G\. Zhou, R\. Balestriero, T\. G\. J\. Rudner, Y\. LeCun, and M\. RenTemporal straightening for latent planning\.InICML,Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p3.1),[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Wijmanset al\.\(2023\)E\. Wijmans, M\. Savva, I\. Essa, S\. Lee, A\. S\. Morcos, and D\. BatraEmergence of maps in the memories of blind navigation agents\.InICLR,Cited by:[§1](https://arxiv.org/html/2609.38383#S1.p3.1)\.
- Xieet al\.\(2018\)J\. Xie, Y\. Lu, R\. Gao, and Y\. N\. WuCooperative learning of energy\-based model and latent variable model via MCMC teaching\.InAAAI,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Xieet al\.\(2016\)J\. Xie, Y\. Lu, S\. Zhu, and Y\. N\. WuA theory of generative ConvNet\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38383#S1.p4.1)\.
- Xuet al\.\(2025\)D\. Xu, R\. Gao, W\. Zhang, X\. Wei, and Y\. N\. WuOn conformal isometry of grid cells: learning distance\-preserving position embedding\.InICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Xuet al\.\(2023a\)Y\. Xu, D\. Kong, D\. Xu, Z\. Ji, B\. Pang, P\. Fung, and Y\. N\. WuDiverse and faithful knowledge\-grounded dialogue generation via sequential posterior inference\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Xuet al\.\(2023b\)Y\. Xu, J\. Xie, T\. Zhao, C\. Baker, Y\. Zhao, and Y\. N\. WuEnergy\-based continuous inverse optimal control\.IEEE Transactions on Neural Networks and Learning Systems\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2023\)Z\. Yang, J\. Mao, Y\. Du, J\. Wu, J\. B\. Tenenbaum, T\. Lozano\-Pérez, and L\. P\. KaelblingCompositional diffusion\-based continuous constraint solvers\.InCoRL,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Yuet al\.\(2026a\)P\. Yu, D\. Zhang, H\. He, X\. Ma, S\. Xie, R\. Miao, Y\. Lu, Y\. Zhang, D\. Kong, R\. Gao, J\. Xie, G\. Cheng, and Y\. N\. Wu"Noisier" noise contrastive estimation is \(almost\) maximum likelihood\.InICLR,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Yuet al\.\(2026b\)Z\. Yu, X\. Hu, and X\. XuQQWorld: quantile\-quantile matching for world model regularization\.arXiv preprint arXiv:2607\.28415\.Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1)\.
- Zhaiet al\.\(2023\)X\. Zhai, B\. Mustafa, A\. Kolesnikov, and L\. BeyerSigmoid loss for language image pre\-training\.InICCV,Cited by:[§2\.3](https://arxiv.org/html/2609.38383#S2.SS3.p3.3)\.
- Zhaoet al\.\(2025a\)M\. Zhao, D\. Xu, D\. Kong, W\. Zhang, and Y\. N\. WuA minimalistic representation model for head direction system\.InProceedings of the 47th Annual Conference of the Cognitive Science Society,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2025b\)M\. Zhao, D\. Xu, D\. Kong, W\. Zhang, and Y\. N\. WuPlace cells as multi\-scale position embeddings: random walk transition kernels for path planning\.InNeurIPS,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38383#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.38383#S2.SS2.p2.1),[§5](https://arxiv.org/html/2609.38383#S5.p2.1)\.
- Zhouet al\.\(2025\)G\. Zhou, H\. Pan, Y\. LeCun, and L\. PintoDINO\-WM: world models on pre\-trained visual features enable zero\-shot planning\.InICML,Cited by:[Appendix A](https://arxiv.org/html/2609.38383#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38383#S1.p3.1),[Table 1](https://arxiv.org/html/2609.38383#S4.T1.7.5.1),[Table 3](https://arxiv.org/html/2609.38383#S4.T3.fig2.1.9.1),[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
- Ziarkoet al\.\(2025\)A\. Ziarko, M\. Bortkiewicz, M\. Zawalski, B\. Eysenbach, and P\. MiłośContrastive representations for temporal reasoning\.InNeurIPS,Cited by:[§5](https://arxiv.org/html/2609.38383#S5.p1.1)\.
###### Appendix
1. [1Introduction](https://arxiv.org/html/2609.38383#S1)
2. [2Method](https://arxiv.org/html/2609.38383#S2)1. [2\.1Problem Formulation](https://arxiv.org/html/2609.38383#S2.SS1) 2. [2\.2Models](https://arxiv.org/html/2609.38383#S2.SS2) 3. [2\.3Learning](https://arxiv.org/html/2609.38383#S2.SS3) 4. [2\.4Planning by Probability Improvement](https://arxiv.org/html/2609.38383#S2.SS4)
3. [3Theoretical Understanding](https://arxiv.org/html/2609.38383#S3)
4. [4Experiments](https://arxiv.org/html/2609.38383#S4)1. [4\.1Experimental Setup](https://arxiv.org/html/2609.38383#S4.SS1) 2. [4\.2Results](https://arxiv.org/html/2609.38383#S4.SS2)
5. [5Related Work](https://arxiv.org/html/2609.38383#S5)
6. [6Limitations and Future Work](https://arxiv.org/html/2609.38383#S6)
7. [References](https://arxiv.org/html/2609.38383#bib)
8. [AExtended related work](https://arxiv.org/html/2609.38383#A1)
9. [BDerivations and Proofs](https://arxiv.org/html/2609.38383#A2)1. [B\.1Density\-ratio learning and normalization](https://arxiv.org/html/2609.38383#A2.SS1) 2. [B\.2Short horizons: geodesic geometry and spatial scale](https://arxiv.org/html/2609.38383#A2.SS2) 3. [B\.3Long horizons: connectivity and mixing](https://arxiv.org/html/2609.38383#A2.SS3) 4. [B\.4What the embedding parameterization preserves](https://arxiv.org/html/2609.38383#A2.SS4) 5. [B\.5Planning: GP and PAP as temperature limits](https://arxiv.org/html/2609.38383#A2.SS5) 6. [B\.6Geometric weights and goal reaching](https://arxiv.org/html/2609.38383#A2.SS6)
10. [CAlgorithms](https://arxiv.org/html/2609.38383#A3)
11. [DExperiments](https://arxiv.org/html/2609.38383#A4)1. [D\.1Representation Probes](https://arxiv.org/html/2609.38383#A4.SS1) 2. [D\.2Demonstrations](https://arxiv.org/html/2609.38383#A4.SS2)
## Appendix AExtended related work
#### Latent world models\.
Latent world models predict how a latent state evolves under actions, rather than predicting future frames, and plan or learn policies in that latent state space\. They can be grouped by the training signal that shapes the latent state\. Reconstruction\-based models learn the latent through a decoder that reconstructs observations\([Ha and Schmidhuber, 2018](https://arxiv.org/html/2609.38383#bib.bib46);[Hafner et al\., 2019](https://arxiv.org/html/2609.38383#bib.bib47);[Hafner et al\., 2020](https://arxiv.org/html/2609.38383#bib.bib48);[Hafner et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib49);[Micheli et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib50)\)\. Reconstruction provides a dense training signal, but it pushes the latent to encode visual detail regardless of its relevance to control\. Reward\-driven models remove the decoder and shape the latent through task signals such as reward and value prediction, often combined with a latent consistency loss\([Schrittwieser et al\., 2020](https://arxiv.org/html/2609.38383#bib.bib51);[Hansen et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib52);[Hansen et al\., 2024](https://arxiv.org/html/2609.38383#bib.bib53)\)\. These objectives encourage the latent to retain information useful for reward prediction and control, but require reward labels and tie representation learning to the training tasks\. JEPA\-style world models need neither a decoder nor rewards: they are trained to predict future embeddings from reward\-free trajectories, and they plan by reaching a goal embedding\. With no reconstruction or reward signal to anchor the latent, their central challenge is representation collapse\. One line of work avoids it by freezing a pretrained visual encoder during action\-conditioned dynamics training, as in DINO\-WM and V\-JEPA 2\-AC\([Zhou et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib15);[Assran et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib55)\), while another trains the encoder end\-to\-end with explicit anti\-collapse regularization\([Sobal et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib24);[Maes et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib25)\)\. LeWorldModel uses a single Gaussian regularizer\([Maes et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib25)\)\. Subsequent work studies alternative regularization\([Yu et al\., 2026b](https://arxiv.org/html/2609.38383#bib.bib57)\), dynamics models\([Huo et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib58);[Gao and Xu, 2026](https://arxiv.org/html/2609.38383#bib.bib56)\), auxiliary physical supervision\([Liu et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib59)\), hierarchical planning\([Caselli et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib60)\), and action optimization with a terminal state\-space energy over a frozen world model\([Pham and Bera, 2026](https://arxiv.org/html/2609.38383#bib.bib61)\)\.
#### Energy\-based models and inference for planning\.
Energy\-based models \(EBMs\) learn distributions through energy functions, with foundational work connecting generative ConvNets to analysis\-by\-synthesis learning and cooperative training with latent generators\([Xie et al\., 2016](https://arxiv.org/html/2609.38383#bib.bib62);[Xie et al\., 2018](https://arxiv.org/html/2609.38383#bib.bib63)\)\. Noise\-contrastive estimation learns density ratios from data and reference samples\([Gutmann and Hyvärinen, 2010](https://arxiv.org/html/2609.38383#bib.bib2);[Yu et al\., 2026a](https://arxiv.org/html/2609.38383#bib.bib54)\)\. For control, learned energies serve as costs in inverse optimal control\([Xu et al\., 2023b](https://arxiv.org/html/2609.38383#bib.bib64)\), guide state\-sequence planning\([Du et al\., 2020](https://arxiv.org/html/2609.38383#bib.bib34)\), and regularize planning toward likely transitions\([Boney et al\., 2020](https://arxiv.org/html/2609.38383#bib.bib35)\)\. Energy\-based motion planners and diffusion planners further support trajectory optimization and compositional conditioning\([Urain et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib36);[Urain et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib37);[Luo et al\., 2024](https://arxiv.org/html/2609.38383#bib.bib38);[Gkanatsios et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib39);[Yang et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib40);[Janner et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib41);[Ajay et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib42)\)\. A related line moves inference into latent spaces, using latent plans for return\-conditioned generation, online adaptation, and policy improvement\([Kong et al\., 2024b](https://arxiv.org/html/2609.38383#bib.bib43);[Noh et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib44);[Qin et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib65)\), or latent EBMs for task adaptation and continuous\-time dynamics\([Pang et al\., 2020](https://arxiv.org/html/2609.38383#bib.bib74);[Cheng et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib45)\)\. Beyond control, latent inference supports molecular design\([Kong et al\., 2023](https://arxiv.org/html/2609.38383#bib.bib66);[Kong et al\., 2024a](https://arxiv.org/html/2609.38383#bib.bib67)\), knowledge\-grounded dialogue\([Xu et al\., 2023a](https://arxiv.org/html/2609.38383#bib.bib68)\), and reasoning through latent thoughts and iterative refinement\([Kong et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib69);[Kong et al\., 2026](https://arxiv.org/html/2609.38383#bib.bib70)\)\. FaST studies a complementary form of adaptive computation, switching between fast responses and slower inference for visual reasoning\([Sun et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib71)\)\. Spatial representation models connect navigation to head\-direction coding\([Zhao et al\., 2025a](https://arxiv.org/html/2609.38383#bib.bib72)\), grid\-cell embeddings that preserve local distances up to a scale factor\([Xu et al\., 2025](https://arxiv.org/html/2609.38383#bib.bib73)\), and place\-cell embeddings of multiscale random\-walk kernels\([Zhao et al\., 2025b](https://arxiv.org/html/2609.38383#bib.bib14)\)\. Our model learns conditional temporal log\-density ratios from observation pairs through NCE\. These scores provide long\-range goal guidance for local action evaluation\. GP retains the joint move\-and\-horizon maximization of[Zhao et al\. \(2025b\)](https://arxiv.org/html/2609.38383#bib.bib14), applied to estimated density\-ratio progress, while our temperature objective also yields PAP, which aggregates progress across horizons into a single goal\-dependent potential\.
## Appendix BDerivations and Proofs
We first show what NCE learns from temporal pairs, then derive the geometric information in these relations at short and long horizons\. We next examine how the embedding model represents this information and how GP and PAP use it to evaluate progress toward a goal\.
The geometric results concern the ideal scoreG∗=log\[pdata\(y\|x,τ\)/p0\(y\)\]G^\{\*\}=\\log\[p\_\{\\rm data\}\(y\|x,\\tau\)/p\_\{0\}\(y\)\]under the stated diffusion or Markov assumptions\. For the Markov analysis to apply directly to observations, each observation must determine the state\. These results describe the ideal temporal relations; learning at finitely many horizons does not by itself establish the same limiting behavior forGθG\_\{\\theta\}\. The planning identities, however, hold for any finite learned scores\.
### B\.1Density\-ratio learning and normalization
NCE learns temporal association by distinguishing observed pairs from pairs with independently sampled targets\. Assume that the referencep0p\_\{0\}covers the conditional target support\. At a sampled horizon, the two pair densities arepdata\(x,y\|τ\)p\_\{\\rm data\}\(x,y\|\\tau\)andpdata\(x\|τ\)p0\(y\)p\_\{\\rm data\}\(x\|\\tau\)p\_\{0\}\(y\)\. WithNNnegatives per positive, the class prior odds are1:N1:N\. The source marginal cancels from the posterior odds, leavingpdata\(y\|x,τ\)/\(Np0\(y\)\)p\_\{\\rm data\}\(y\|x,\\tau\)/\(Np\_\{0\}\(y\)\)\.
The classifierσ\(G−logN\)\\sigma\(G\-\\log N\)has oddseG/Ne^\{G\}/N\. Matching these odds gives the unrestricted population minimizer of the logistic loss in[Eq\.5](https://arxiv.org/html/2609.38383#S2.E5):
G∗\(x,y,τ\)=logpdata\(y\|x,τ\)p0\(y\)\.G^\{\*\}\(x,y,\\tau\)=\\log\\frac\{p\_\{\\rm data\}\(y\|x,\\tau\)\}\{p\_\{0\}\(y\)\}\.This equality holds on sampled sources and horizons where both densities are positive\([Gutmann and Hyvärinen, 2010](https://arxiv.org/html/2609.38383#bib.bib2);[Ma and Collins, 2018](https://arxiv.org/html/2609.38383#bib.bib3)\)\. Wherep0\(y\)\>0p\_\{0\}\(y\)\>0and the data conditional is zero, the optimum is approached asG→−∞G\\to\-\\infty, and we interpreteG∗=0e^\{G^\{\*\}\}=0\.
The fitted score need not satisfy the same normalization as this ideal ratio\. For the normalized EBM in[Eq\.1](https://arxiv.org/html/2609.38383#S2.E1), the exact relationship is
logpθ\(y\|x,τ\)p0\(y\)=Gθ\(x,y,τ\)−log𝔼p0\(y′\)\[eGθ\(x,y′,τ\)\]\.\\log\\frac\{p\_\{\\theta\}\(y\|x,\\tau\)\}\{p\_\{0\}\(y\)\}=G\_\{\\theta\}\(x,y,\\tau\)\-\\log\\mathbb\{E\}\_\{p\_\{0\}\(y^\{\\prime\}\)\}\\\!\\left\[e^\{G\_\{\\theta\}\(x,y^\{\\prime\},\\tau\)\}\\right\]\.\(12\)The expectation equalseβ1\(τ\)Zθ\(x,τ\)e^\{\\beta\_\{1\}\(\\tau\)\}Z\_\{\\theta\}\(x,\\tau\)and measures how far the fitted ratio is from integrating to one\. The ideal ratio integrates to one at every source\. A shared offsetβ1\(τ\)\\beta\_\{1\}\(\\tau\)cannot enforce this whenZθ\(x,τ\)Z\_\{\\theta\}\(x,\\tau\)varies withxx\. This matters during planning: candidates are different sources paired with the same goal, so a source\-dependent correction can change their ranking\. The planner uses the fitted scoreGθG\_\{\\theta\}\.
### B\.2Short horizons: geodesic geometry and spatial scale
Short temporal relations reveal path geometry through the diffusion limit of the walk\. To obtain this limit, shrink free\-space steps byδ\\deltaand assign each step timeδ2\\delta^\{2\}\. The covariance per unit time remains2αIn2\\alpha I\_\{n\}\. Under the invariance principle assumed in[Sec\.3](https://arxiv.org/html/2609.38383#S3), the walk converges to reflected Brownian motion with generatorαΔ\\alpha\\Delta\([Burdzy and Chen, 2008](https://arxiv.org/html/2609.38383#bib.bib29)\)\. Reflection prevents probability from flowing through walls\. We first take this continuum limit, then let diffusion time tend to zero\.
###### Proof of[Proposition1](https://arxiv.org/html/2609.38383#Thmproposition1)\.
On the smooth, bounded, connected domain of the proposition, Varadhan’s formula for the reflecting heat kernel gives, for fixed interiorx,yx,y,
−4ατG∗\(x,y,τ\)=−4ατlogpdata\(y\|x,τ\)\+4ατlogp0\(y\)⟶dgeo\(x,y\)2,\-4\\alpha\\tau G^\{\*\}\(x,y,\\tau\)=\-4\\alpha\\tau\\log p\_\{\\rm data\}\(y\|x,\\tau\)\+4\\alpha\\tau\\log p\_\{0\}\(y\)\\longrightarrow d\_\{\\rm geo\}\(x,y\)^\{2\},because the uniform reference is positive and independent ofτ\\tau\([Varadhan, 1967](https://arxiv.org/html/2609.38383#bib.bib18);[Norris, 1997](https://arxiv.org/html/2609.38383#bib.bib19)\)\. ∎
This limit also explains how the score ranks candidate moves\. Among a finite set of interior candidates, suppose one has strictly shorter geodesic distance toygy\_\{g\}than all the others\. It maximizesG∗\(⋅,yg,τ\)G^\{\*\}\(\\cdot,y\_\{g\},\\tau\)for sufficiently smallτ\\tau\. Indeed, its score advantage over each competitor, multiplied by4ατ4\\alpha\\tau, tends to the strictly positive difference of squared distances\.
The order of limits matters\. On a fixed graph, letℓ≥1\\ell\\geq 1be the smallest number of edges between distinct reachable statesx,yx,y\. ThenPk\(x,y\)=0P^\{k\}\(x,y\)=0fork<ℓk<\\ell\. Even the continuous\-time walk on this graph satisfies
\[eτ\(P−I\)\]\(x,y\)=τℓPℓ\(x,y\)/ℓ\!\+O\(τℓ\+1\)\.\[e^\{\\tau\(P\-I\)\}\]\(x,y\)=\\tau^\{\\ell\}P^\{\\ell\}\(x,y\)/\\ell\!\+O\(\\tau^\{\\ell\+1\}\)\.Its logarithm therefore scales asℓlogτ\\ell\\log\\tau, rather than−dgeo2/\(4ατ\)\-d\_\{\\rm geo\}^\{2\}/\(4\\alpha\\tau\)\([Steinerberger, 2018](https://arxiv.org/html/2609.38383#bib.bib21)\)\.
The horizon giving the strongest local progress\.The geodesic limit describes which candidate scores higher\. To see which horizon gives the strongest progress signal, consider diffusion in free spaceℝn\\mathbb\{R\}^\{n\}, whose transition density is
pdata\(y\|x,τ\)=\(4πατ\)−n/2exp\(−‖y−x‖224ατ\)\.p\_\{\\rm data\}\(y\|x,\\tau\)=\(4\\pi\\alpha\\tau\)^\{\-n/2\}\\exp\\\!\\left\(\-\\frac\{\\\|y\-x\\\|\_\{2\}^\{2\}\}\{4\\alpha\\tau\}\\right\)\.\(13\)Fixx≠ygx\\neq y\_\{g\}, and write the goal distance asr=‖yg−x‖2r=\\\|y\_\{g\}\-x\\\|\_\{2\}and the direction toward it asu=\(yg−x\)/ru=\(y\_\{g\}\-x\)/r\. For a fixed reference with0<p0\(yg\)<∞0<p\_\{0\}\(y\_\{g\}\)<\\infty, the change in the ratio per unit step toward the goal is
∂∂ϵeG∗\(x\+ϵu,yg,τ\)\|ϵ=0=r2ατpdata\(yg\|x,τ\)p0\(yg\)\.\\left\.\\frac\{\\partial\}\{\\partial\\epsilon\}e^\{G^\{\*\}\(x\+\\epsilon u,y\_\{g\},\\tau\)\}\\right\|\_\{\\epsilon=0\}=\\frac\{r\}\{2\\alpha\\tau\}\\frac\{p\_\{\\rm data\}\(y\_\{g\}\|x,\\tau\)\}\{p\_\{0\}\(y\_\{g\}\)\}\.\(14\)We maximize this local progress over the horizon\. Its logarithmic derivative inτ\\tauis−\(n\+2\)/\(2τ\)\+r2/\(4ατ2\)\-\(n\+2\)/\(2\\tau\)\+r^\{2\}/\(4\\alpha\\tau^\{2\}\), which changes from positive to negative at
τ∗=r22\(n\+2\)α,2nατ∗=rnn\+2\.\\tau^\{\*\}=\\frac\{r^\{2\}\}\{2\(n\+2\)\\alpha\},\\qquad\\sqrt\{2n\\alpha\\tau^\{\*\}\}=r\\sqrt\{\\frac\{n\}\{n\+2\}\}\.\(15\)The strongest local signal therefore occurs when the walk’s root\-mean\-square displacement is a fixed fraction of the goal distance\. At shorter horizons, little probability has reached the goal; at longer horizons, probability has spread over a larger region\. Their balance gives the squared\-distance scaling in[Sec\.3](https://arxiv.org/html/2609.38383#S3)\. This calculation concerns infinitesimal motion in free space; finite moves, walls, and passages can change the maximizing horizon\.
### B\.3Long horizons: connectivity and mixing
At longer horizons, the walk loses information about its starting position\. The spectral expansion describes which differences persist as this happens\. We now return to a finite state space𝒮\\mathcal\{S\}and a discrete walk with irreducible, aperiodic, reversible transition matrixPP\. Letπ\\pibe its stationary distribution, setp0=πp\_\{0\}=\\pi, and order the eigenvalues as1=λ1\>λ2≥⋯1=\\lambda\_\{1\}\>\\lambda\_\{2\}\\geq\\cdots\. At integer horizons,pdata\(y\|x,τ\)=Pτ\(x,y\)p\_\{\\rm data\}\(y\|x,\\tau\)=P^\{\\tau\}\(x,y\)\.
###### Proof of[Proposition2](https://arxiv.org/html/2609.38383#Thmproposition2)\.
PutDπ=diag\(π\)D\_\{\\pi\}=\\operatorname\{diag\}\(\\pi\)\. Detailed balance makesS=Dπ1/2PDπ−1/2S=D\_\{\\pi\}^\{1/2\}PD\_\{\\pi\}^\{\-1/2\}symmetric\. IfSϕj=λjϕjS\\phi\_\{j\}=\\lambda\_\{j\}\\phi\_\{j\}with orthonormalϕj\\phi\_\{j\}, thenψj=Dπ−1/2ϕj\\psi\_\{j\}=D\_\{\\pi\}^\{\-1/2\}\\phi\_\{j\}are orthonormal underπ\\pi\. Irreducibility givesλ1=1\\lambda\_\{1\}=1,ψ1=1\\psi\_\{1\}=1; aperiodicity gives\|λj\|<1\|\\lambda\_\{j\}\|<1forj≥2j\\geq 2\. ExpandingPτ=Dπ−1/2SτDπ1/2P^\{\\tau\}=D\_\{\\pi\}^\{\-1/2\}S^\{\\tau\}D\_\{\\pi\}^\{1/2\}yields
eG∗\(x,y,τ\)=Pτ\(x,y\)π\(y\)=1\+∑j≥2λjτψj\(x\)ψj\(y\)\.e^\{G^\{\*\}\(x,y,\\tau\)\}=\\frac\{P^\{\\tau\}\(x,y\)\}\{\\pi\(y\)\}=1\+\\sum\_\{j\\geq 2\}\\lambda\_\{j\}^\{\\tau\}\\psi\_\{j\}\(x\)\\psi\_\{j\}\(y\)\.\(16\)Taking logarithms wherePτ\(x,y\)\>0P^\{\\tau\}\(x,y\)\>0proves the proposition\([Aldous and Fill, 2002](https://arxiv.org/html/2609.38383#bib.bib20);[Coifman and Lafon, 2006](https://arxiv.org/html/2609.38383#bib.bib11)\); the ratio identity also holds at zero transition probability\. ∎
Why weak connections produce slow modes\.To connect the spectrum to the environment, we measure how much a function changes across transitions of the walk\. Write⟨f,g⟩π=∑xπ\(x\)f\(x\)g\(x\)\\langle f,g\\rangle\_\{\\pi\}=\\sum\_\{x\}\\pi\(x\)f\(x\)g\(x\)\. Expanding the squared change and using stationarity gives
ℰ\(f\)\\displaystyle\\mathcal\{E\}\(f\):=12∑u,vπ\(u\)P\(u,v\)\[f\(u\)−f\(v\)\]2=⟨f,\(I−P\)f⟩π,\\displaystyle:=\\frac\{1\}\{2\}\\sum\_\{u,v\}\\pi\(u\)P\(u,v\)\[f\(u\)\-f\(v\)\]^\{2\}=\\langle f,\(I\-P\)f\\rangle\_\{\\pi\},\(17\)ℰ\(ψj\)\\displaystyle\\mathcal\{E\}\(\\psi\_\{j\}\)=1−λj\.\\displaystyle=1\-\\lambda\_\{j\}\.An eigenfunction with eigenvalue near one therefore varies little across transitions, on average\. Now divide the states into regionsAAandAcA^\{c\}, with0<π\(A\)<10<\\pi\(A\)<1\. The functionf=\(𝟏A−π\(A\)\)/π\(A\)π\(Ac\)f=\(\\mathbf\{1\}\_\{A\}\-\\pi\(A\)\)/\\sqrt\{\\pi\(A\)\\pi\(A^\{c\}\)\}is constant within each region and has zero mean and unit norm underπ\\pi\. Substituting it into the variational formula1−λ2=min⟨f,1⟩π=0,⟨f,f⟩π=1ℰ\(f\)1\-\\lambda\_\{2\}=\\min\_\{\\langle f,1\\rangle\_\{\\pi\}=0,\\,\\langle f,f\\rangle\_\{\\pi\}=1\}\\mathcal\{E\}\(f\)gives
1−λ2≤Q\(A,Ac\)π\(A\)π\(Ac\),Q\(A,Ac\)=∑u∈A,v∉Aπ\(u\)P\(u,v\)\.1\-\\lambda\_\{2\}\\leq\\frac\{Q\(A,A^\{c\}\)\}\{\\pi\(A\)\\pi\(A^\{c\}\)\},\\qquad Q\(A,A^\{c\}\)=\\sum\_\{u\\in A,\\,v\\notin A\}\\pi\(u\)P\(u,v\)\.\(18\)HereQ\(A,Ac\)Q\(A,A^\{c\}\)is the stationary probability flow from one region to the other\. Only transitions between the regions contribute toℰ\(f\)\\mathcal\{E\}\(f\), and reversibility makes the two directions equal\([Aldous and Fill, 2002](https://arxiv.org/html/2609.38383#bib.bib20), Sec\. 3\.6\)\. When this flow is small relative to the regions’ stationary masses, the bound placesλ2\\lambda\_\{2\}near one\.
For two rooms joined by a narrow doorway, this gives a mode that decays slowly\. A large change across the rarely crossed doorway contributes little to its Dirichlet energy, while changes across frequent transitions within a room contribute more\. When each room mixes quickly, this mode describes a contrast between the rooms, as in[Sec\.3](https://arxiv.org/html/2609.38383#S3)\.
What survives before mixing\.Eventually, mixing removes these differences between starting states\. To bound the remaining contrast, letρ=maxj≥2\|λj\|<1\\rho=\\max\_\{j\\geq 2\}\|\\lambda\_\{j\}\|<1\. Completeness gives∑j≥2ψj\(x\)2=π\(x\)−1−1\\sum\_\{j\\geq 2\}\\psi\_\{j\}\(x\)^\{2\}=\\pi\(x\)^\{\-1\}\-1, so Cauchy–Schwarz in[Eq\.16](https://arxiv.org/html/2609.38383#A2.E16)yields
\|eG∗\(x,y,τ\)−1\|≤ρτ\(π\(x\)−1−1\)\(π\(y\)−1−1\)\.\\bigl\|e^\{G^\{\*\}\(x,y,\\tau\)\}\-1\\bigr\|\\leq\\rho^\{\\tau\}\\sqrt\{\\bigl\(\\pi\(x\)^\{\-1\}\-1\\bigr\)\\bigl\(\\pi\(y\)^\{\-1\}\-1\\bigr\)\}\.\(19\)Thus every ratio approaches one,G∗→0G^\{\*\}\\to 0, and progress between any two sources vanishes\. When one positive mode decays more slowly than all the others, it determines the leading contrast\. Specifically, ifλ2\>0\\lambda\_\{2\}\>0andλ2\>\|λj\|\\lambda\_\{2\}\>\|\\lambda\_\{j\}\|for everyj≥3j\\geq 3, then applyinglog\(1\+z\)=z\+O\(z2\)\\log\(1\+z\)=z\+O\(z^\{2\}\)to[Eq\.16](https://arxiv.org/html/2609.38383#A2.E16)gives
G∗\(x,y,τ\)=λ2τψ2\(x\)ψ2\(y\)\+o\(λ2τ\)\.G^\{\*\}\(x,y,\\tau\)=\\lambda\_\{2\}^\{\\tau\}\\psi\_\{2\}\(x\)\\psi\_\{2\}\(y\)\+o\(\\lambda\_\{2\}^\{\\tau\}\)\.\(20\)Without this separation, we must retain all dominant modes in the full expansion\. Negative eigenvalues contribute terms that alternate in sign between odd and even horizons\.
### B\.4What the embedding parameterization preserves
The learned score describes temporal relations through embedding alignment\. When both embeddings have unit norm, we can express the same relation through their distance:
eGθ\(x,y,τ\)=eβ0\(τ\)\+β1\(τ\)exp\[−β0\(τ\)2‖hθ\(x,τ\)−gθ\(y,τ\)‖22\]\.e^\{G\_\{\\theta\}\(x,y,\\tau\)\}=e^\{\\beta\_\{0\}\(\\tau\)\+\\beta\_\{1\}\(\\tau\)\}\\exp\\\!\\left\[\-\\frac\{\\beta\_\{0\}\(\\tau\)\}\{2\}\\\|h\_\{\\theta\}\(x,\\tau\)\-g\_\{\\theta\}\(y,\\tau\)\\\|\_\{2\}^\{2\}\\right\]\.\(21\)This identity holds for any learned coefficientβ0\(τ\)\\beta\_\{0\}\(\\tau\)and offsetβ1\(τ\)\\beta\_\{1\}\(\\tau\)\. The coefficient controls how the score changes with embedding distance; its sign determines the direction of this dependence\. At a fixed horizon, the offset rescales all exponentiated scores equally and leaves their ordering unchanged\. Across horizons, these scales affect how much each horizon contributes to planning\.
For tied encoders, comparison with the goal itself gives the exact identity
Gθ\(yg,yg,τ\)−Gθ\(x,yg,τ\)=β0\(τ\)2‖hθ\(x,τ\)−hθ\(yg,τ\)‖22\.G\_\{\\theta\}\(y\_\{g\},y\_\{g\},\\tau\)\-G\_\{\\theta\}\(x,y\_\{g\},\\tau\)=\\frac\{\\beta\_\{0\}\(\\tau\)\}\{2\}\\\|h\_\{\\theta\}\(x,\\tau\)\-h\_\{\\theta\}\(y\_\{g\},\\tau\)\\\|\_\{2\}^\{2\}\.\(22\)In our evaluated tied\-encoder models, the fitted alignment coefficients are positive and the offsets are negative\. The positive coefficients make closer embeddings score higher and place the goal at a score maximum at each evaluated horizon, consistent with the geometric interpretation\. This is an observation about the fitted models\. A maximum at the goal does not rule out other local maxima or ensure that an improving move is available at every state\.
Tied encoders also preserve a structural property of reversible walks\. At a fixed horizon on a finite state space withp0\(x\)\>0p\_\{0\}\(x\)\>0, suppress the horizon argument and writeKxy=exp\(β0⟨hθ\(x\),hθ\(y\)⟩\)K\_\{xy\}=\\exp\(\\beta\_\{0\}\\langle h\_\{\\theta\}\(x\),h\_\{\\theta\}\(y\)\\rangle\)\. This kernel is symmetric for either sign ofβ0\\beta\_\{0\}\. The normalized conditional matrixp0\(y\)Kxy/Zθ\(x\)p\_\{0\}\(y\)K\_\{xy\}/Z\_\{\\theta\}\(x\)is therefore reversible: with stationary weights proportional top0\(x\)Zθ\(x\)p\_\{0\}\(x\)Z\_\{\\theta\}\(x\), the flow betweenxxandyyis proportional top0\(x\)p0\(y\)Kxyp\_\{0\}\(x\)p\_\{0\}\(y\)K\_\{xy\}, which is symmetric\. This holds at each horizon, but does not require the learned conditionals to be powers of a single transition matrix\.
These properties do not establish exact representation of every reversible walk\. Tied unit embeddings give the same diagonal score at every state, whereas the walk’s return ratio can vary from state to state\.
There is a related feature construction for the ratio itself\. For integersk≥1k\\geq 1, define the nonnegative featuresfk\(x\)z=Pk\(x,z\)/π\(z\)f\_\{k\}\(x\)\_\{z\}=P^\{k\}\(x,z\)/\\sqrt\{\\pi\(z\)\}\. Detailed balance and matrix multiplication give⟨fk\(x\),fk\(y\)⟩=P2k\(x,y\)/π\(y\)\\langle f\_\{k\}\(x\),f\_\{k\}\(y\)\\rangle=P^\{2k\}\(x,y\)/\\pi\(y\)\. This diffusion\-kernel construction\([Coifman and Lafon, 2006](https://arxiv.org/html/2609.38383#bib.bib11)\)represents the ratio as an inner product\. Our affine inner product models its logarithm, so the construction does not establish exact representation by our model\.
### B\.5Planning: GP and PAP as temperature limits
The same action can improve the goal score at one horizon and reduce it at another\. The temperature objective lets each action place more weight on horizons where it makes progress, while penalizing departures from uniform weights\.
Fix a goalygy\_\{g\}, a nonempty finite horizon set𝒯\\mathcal\{T\}, and finite scores\. For a candidate actionaa, writeΔτ\(a\)=eGθ\(x^t\+1a,yg,τ\)−eGθ\(xt,yg,τ\)\\Delta\_\{\\tau\}\(a\)=e^\{G\_\{\\theta\}\(\\widehat\{x\}\_\{t\+1\}^\{\\,a\},y\_\{g\},\\tau\)\}\-e^\{G\_\{\\theta\}\(x\_\{t\},y\_\{g\},\\tau\)\}\. LetUUbe the uniform distribution on𝒯\\mathcal\{T\}, and let𝒫\(𝒯\)\\mathcal\{P\}\(\\mathcal\{T\}\)be the set of distributions over these horizons\.
###### Proposition 3\(GP and PAP as temperature limits\)\.
Forβ\>0\\beta\>0, the planning objective satisfies
Vβ\(a\)\\displaystyle V\_\{\\beta\}\(a\):=βlog\[1\|𝒯\|∑τeΔτ\(a\)/β\]\\displaystyle:=\\beta\\log\\\!\\left\[\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\}e^\{\\Delta\_\{\\tau\}\(a\)/\\beta\}\\right\]\(23\)=maxq∈𝒫\(𝒯\)\{∑τq\(τ\)Δτ\(a\)−βKL\(q∥U\)\},\\displaystyle=\\max\_\{q\\in\\mathcal\{P\}\(\\mathcal\{T\}\)\}\\left\\\{\\sum\_\{\\tau\}q\(\\tau\)\\Delta\_\{\\tau\}\(a\)\-\\beta\\,\\mathrm\{KL\}\(q\\\|U\)\\right\\\},with unique maximizerqβ\(τ\|a\)∝eΔτ\(a\)/βq\_\{\\beta\}\(\\tau\|a\)\\propto e^\{\\Delta\_\{\\tau\}\(a\)/\\beta\}\. Asβ→0\+\\beta\\to 0^\{\+\},Vβ\(a\)V\_\{\\beta\}\(a\)converges tomaxτΔτ\(a\)\\max\_\{\\tau\}\\Delta\_\{\\tau\}\(a\)\(GP\); asβ→∞\\beta\\to\\inftyit converges to\|𝒯\|−1∑τΔτ\(a\)\|\\mathcal\{T\}\|^\{\-1\}\\sum\_\{\\tau\}\\Delta\_\{\\tau\}\(a\)\(PAP\)\.
###### Proof\.
WithKL\(q∥U\)=∑τq\(τ\)log\(\|𝒯\|q\(τ\)\)\\mathrm\{KL\}\(q\\\|U\)=\\sum\_\{\\tau\}q\(\\tau\)\\log\(\|\\mathcal\{T\}\|q\(\\tau\)\)and0log0=00\\log 0=0, direct substitution gives
∑τq\(τ\)Δτ\(a\)−βKL\(q∥U\)=Vβ\(a\)−βKL\(q∥qβ\(⋅\|a\)\)\.\\sum\_\{\\tau\}q\(\\tau\)\\Delta\_\{\\tau\}\(a\)\-\\beta\\,\\mathrm\{KL\}\(q\\\|U\)=V\_\{\\beta\}\(a\)\-\\beta\\,\\mathrm\{KL\}\(q\\\|q\_\{\\beta\}\(\\cdot\|a\)\)\.\(24\)The right\-hand side is largest only whenq=qβq=q\_\{\\beta\}, because KL is nonnegative and vanishes only when the two distributions agree\. This proves the variational identity and uniqueness\. To obtain the limits, writeM=maxτΔτ\(a\)M=\\max\_\{\\tau\}\\Delta\_\{\\tau\}\(a\)\. Evaluating the variational objective at weights concentrated on a maximizing horizon, atUU, and atqβq\_\{\\beta\}gives
M−βlog\|𝒯\|\\displaystyle M\-\\beta\\log\|\\mathcal\{T\}\|≤Vβ\(a\)≤M,\\displaystyle\\leq V\_\{\\beta\}\(a\)\\leq M,\(25\)1\|𝒯\|∑τΔτ\(a\)\\displaystyle\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\}\\Delta\_\{\\tau\}\(a\)≤Vβ\(a\)≤∑τqβ\(τ\|a\)Δτ\(a\)\.\\displaystyle\\leq V\_\{\\beta\}\(a\)\\leq\\sum\_\{\\tau\}q\_\{\\beta\}\(\\tau\|a\)\\Delta\_\{\\tau\}\(a\)\.The first line gives the GP limit; the second gives the PAP limit becauseqβ→Uq\_\{\\beta\}\\to Uasβ→∞\\beta\\to\\infty\. ∎
At low temperature, an action is judged by the horizon where it makes the most progress\. At high temperature, all horizons receive equal weight, so gains and losses offset each other\. Multiplying all progress values by the same positive constant preserves GP and PAP rankings\. At finite temperature, preserving the ranking also requires scalingβ\\betaby that constant\.
Potential ascent and revisits\.PAP evaluates every move through one function of the state\. For the fixed goal, define
Φ\(x\)=1\|𝒯\|∑τ∈𝒯eGθ\(x,yg,τ\)\.\\Phi\(x\)=\\frac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\}e^\{G\_\{\\theta\}\(x,y\_\{g\},\\tau\)\}\.\(26\)PAP progress is exactlyΦ\(x^t\+1a\)−Φ\(xt\)\\Phi\(\\widehat\{x\}\_\{t\+1\}^\{\\,a\}\)\-\\Phi\(x\_\{t\}\)\. Since the current\-state term is shared by all actions, PAP ranks candidates by their next\-state potential\. If the model, goal, and horizon set stay fixed, predictions match deterministic outcomes, and every executed step has strictly positive progress, thenΦ\(xt\+1\)\>Φ\(xt\)\\Phi\(x\_\{t\+1\}\)\>\\Phi\(x\_\{t\}\)\. Returning to a previous state would restore its previous potential, which is impossible\. The argument does not ensure that an improving action exists at every non\-goal state\.
GP can favor different horizons on successive moves and need not increase a common state potential\. Consider two non\-goal states with exponentiated score vectors\(0\.2,0\.8\)\(0\.2,0\.8\)and\(0\.8,0\.2\)\(0\.8,0\.2\)across two horizons\. Each state offers a move to the other and a stay action\. Moving improves one component by0\.60\.6and reduces the other by0\.60\.6, so GP prefers moving in both directions and cycles between the states\. At every finite temperature, moving also givesVβ=βlogcosh\(0\.6/β\)\>0V\_\{\\beta\}=\\beta\\log\\cosh\(0\.6/\\beta\)\>0, while staying gives zero\. PAP assigns both states potential0\.50\.5, so neither move has strictly positive PAP progress\. This cycling example is realizable even with tied, positive unit embeddings,111Choose positive unit vectorsu,vu,vwith⟨u,v⟩=1/2\\langle u,v\\rangle=1/2, useuufor the goal at both horizons, and assign\(v,u\)\(v,u\)and\(u,v\)\(u,v\)to the two states across horizons\. Settingβ0=log16\\beta\_\{0\}=\\log 16andβ1=log0\.05\\beta\_\{1\}=\\log 0\.05gives the stated exponentiated scores\.showing that these embedding constraints do not prevent GP from cycling\. The scores need not be calibrated temporal ratios of a common walk, so the example concerns the learned\-score planner\.
### B\.6Geometric weights and goal reaching
With exact transition ratios, the potential has a direct interpretation under the exploration walk:\|𝒯\|p0\(yg\)Φ\(x\)\|\\mathcal\{T\}\|p\_\{0\}\(y\_\{g\}\)\\Phi\(x\)is the expected number of visits toygy\_\{g\}at horizons in𝒯\\mathcal\{T\}\. Counting visits differs from measuring whether the goal is reached, since one trajectory may visit it several times\.
Geometric weights let us relate these two quantities\. Consider any finite Markov chainPP, take a fixedp0\(yg\)\>0p\_\{0\}\(y\_\{g\}\)\>0, and let0<γ<10<\\gamma<1\. Weighting all horizons, including zero, defines
Φγ\(x\)=\(1−γ\)∑τ=0∞γτPτ\(x,yg\)p0\(yg\)\.\\Phi\_\{\\gamma\}\(x\)=\(1\-\\gamma\)\\sum\_\{\\tau=0\}^\{\\infty\}\\gamma^\{\\tau\}\\frac\{P^\{\\tau\}\(x,y\_\{g\}\)\}\{p\_\{0\}\(y\_\{g\}\)\}\.\(27\)This potential measures discounted occupancy of the goal relative top0\(yg\)p\_\{0\}\(y\_\{g\}\)\. The following result relates it to the time of first arrival\. Neither reversibility nor an absorbing goal is required\.
###### Proposition 4\(Geometric weights connect occupancy to first arrival\)\.
LetTg=inf\{t≥0:Xt=yg\}T\_\{g\}=\\inf\\\{t\\geq 0:X\_\{t\}=y\_\{g\}\\\}be the first arrival time underPP, withγ∞=0\\gamma^\{\\infty\}=0, and let𝔼x\\mathbb\{E\}\_\{x\}denote expectation for the walk started atxx\. Then
Φγ\(x\)=𝔼x\[γTg\]Φγ\(yg\)\.\\Phi\_\{\\gamma\}\(x\)=\\mathbb\{E\}\_\{x\}\[\\gamma^\{T\_\{g\}\}\]\\,\\Phi\_\{\\gamma\}\(y\_\{g\}\)\.\(28\)Every non\-goal state from which the goal is reachable satisfies
maxy:P\(x,y\)\>0Φγ\(y\)≥Φγ\(x\)/γ\>Φγ\(x\)\.\\max\_\{y:P\(x,y\)\>0\}\\Phi\_\{\\gamma\}\(y\)\\geq\\Phi\_\{\\gamma\}\(x\)/\\gamma\>\\Phi\_\{\\gamma\}\(x\)\.\(29\)
###### Proof\.
There are no visits to the goal beforeTgT\_\{g\}\. Once the walk arrives, its subsequent visits have the same law as those of a walk started at the goal\. The strong Markov property therefore gives
𝔼x\[∑t≥0γt𝟏\{Xt=yg\}\]=𝔼x\[γTg\]𝔼yg\[∑j≥0γj𝟏\{Xj=yg\}\]\.\\mathbb\{E\}\_\{x\}\\\!\\left\[\\sum\_\{t\\geq 0\}\\gamma^\{t\}\\mathbf\{1\}\\\{X\_\{t\}=y\_\{g\}\\\}\\right\]=\\mathbb\{E\}\_\{x\}\[\\gamma^\{T\_\{g\}\}\]\\,\\mathbb\{E\}\_\{y\_\{g\}\}\\\!\\left\[\\sum\_\{j\\geq 0\}\\gamma^\{j\}\\mathbf\{1\}\\\{X\_\{j\}=y\_\{g\}\\\}\\right\]\.Multiplying by\(1−γ\)/p0\(yg\)\(1\-\\gamma\)/p\_\{0\}\(y\_\{g\}\)proves[Eq\.28](https://arxiv.org/html/2609.38383#A2.E28)\([Aldous and Fill, 2002](https://arxiv.org/html/2609.38383#bib.bib20), Lemma 2\.25\)\. To show that an improving move exists, split off the horizon\-zero term:Φγ\(x\)=1−γp0\(yg\)𝟏\{x=yg\}\+γ∑yP\(x,y\)Φγ\(y\)\\Phi\_\{\\gamma\}\(x\)=\\frac\{1\-\\gamma\}\{p\_\{0\}\(y\_\{g\}\)\}\\mathbf\{1\}\\\{x=y\_\{g\}\\\}\+\\gamma\\sum\_\{y\}P\(x,y\)\\Phi\_\{\\gamma\}\(y\)\. At a reachable non\-goal state,Φγ\(x\)\>0\\Phi\_\{\\gamma\}\(x\)\>0and the indicator vanishes\. The average next\-state potential is thereforeΦγ\(x\)/γ\>Φγ\(x\)\\Phi\_\{\\gamma\}\(x\)/\\gamma\>\\Phi\_\{\\gamma\}\(x\)\. At least one successor withP\(x,y\)\>0P\(x,y\)\>0has value at least this average, proving[Eq\.29](https://arxiv.org/html/2609.38383#A2.E29)\. ∎
This gives a goal\-reaching guarantee under exact local control\. Suppose all walk\-supported moves are available as deterministic actions, their outcomes are predicted exactly, and planning stops at the goal\. From any start that can reach the goal underPP, choosing a successor with largestΦγ\\Phi\_\{\\gamma\}strictly increases the potential\. Its value remains positive, so the goal remains reachable\. No state can be revisited, and the trajectory cannot stop at a non\-goal state\. It must therefore reach the goal in at most\|𝒮\|−1\|\\mathcal\{S\}\|\-1moves\. The known horizon\-zero term anchors the goal value; omitting it can destroy this guarantee\.
The geometric aggregate is proportional to discounted occupancy, the quantity connected to contrastive goal\-conditioned RL\([Eysenbach et al\., 2022](https://arxiv.org/html/2609.38383#bib.bib4)\)\. Dividing by its value at the goal gives the discounted first\-arrival value\. Under the exact deterministic control assumptions above, choosing the successor with largest value is a policy\-improvement step over the exploration behavior\. Our planner uses a finite set of uniformly weighted positive horizons, which does not generally satisfy this Bellman identity\. The geometric extension therefore explains a connection to goal reaching\.
## Appendix CAlgorithms
[AppendicesC](https://arxiv.org/html/2609.38383#A3)and[C](https://arxiv.org/html/2609.38383#A3)learn the temporal representation and local dynamics separately\.[AppendixC](https://arxiv.org/html/2609.38383#A3)uses both models for closed\-loop planning, with their parameters fixed\.
Algorithm 1: Learningθ\\theta
Input:Exploration trajectories,𝒯\\mathcal\{T\},p\(τ\)p\(\\tau\), reference samplerp0p\_\{0\}, and negative countNN\.
- 1\.Initialize temporal parametersθ\\theta\.
- 2\.whiletemporal updates remaindo
- 3\.Sample a horizonτ∼p\(τ\)\\tau\\sim p\(\\tau\)\.
- 4\.Sample a batch of pairs\(x,y\)\(x,y\)frompdata\(x,y\|τ\)p\_\{\\rm data\}\(x,y\|\\tau\)within episodes\.
- 5\.For each sourcexx, sampleNNindependent negative targets fromp0p\_\{0\}\.
- 6\.ComputeGθ\(x,y,τ\)−logNG\_\{\\theta\}\(x,y,\\tau\)\-\\log Nfor positive and negative pairs\.
- 7\.Updateθ\\thetaby[Eq\.5](https://arxiv.org/html/2609.38383#S2.E5): average the positive loss plus the sum of negative losses over sources\.
- 8\.end while
- 9\.returnlearned parametersθ\\theta\.
Algorithm 2: Learningϕ\\phi
Input:Action\-labeled one\-step transitions with distributionpstepp\_\{\\rm step\}\.
- 1\.Initialize local\-dynamics parametersϕ\\phi\.
- 2\.whiledynamics updates remaindo
- 3\.Sample a batch of transitions\(xt,at,xt\+1\)∼pstep\(x\_\{t\},a\_\{t\},x\_\{t\+1\}\)\\sim p\_\{\\rm step\}\.
- 4\.Updateϕ\\phiby[Eq\.6](https://arxiv.org/html/2609.38383#S2.E6)for state prediction\.
- 5\.end while
- 6\.returnlearned parametersϕ\\phi\.
Algorithm 3: Planning
Input:FrozenGθG\_\{\\theta\}andFϕF\_\{\\phi\}, goalygy\_\{g\}, horizon set𝒯\\mathcal\{T\}, temperatureβ\>0\\beta\>0\(or either limit\), and action budget\.
- 1\.Observe the initialx0x\_\{0\}and sett=0t=0\.
- 2\.whilegoal not reached and budget remainsdo
- 3\.Construct candidate set𝒜\(xt\)\\mathcal\{A\}\(x\_\{t\}\)\.
- 4\.for eacha∈𝒜\(xt\)a\\in\\mathcal\{A\}\(x\_\{t\}\)do
- 5\.Predictx^t\+1a=Fϕ\(xt,a\)\\widehat\{x\}\_\{t\+1\}^\{\\,a\}=F\_\{\\phi\}\(x\_\{t\},a\)\.
- 6\.Compute progress for eachτ∈𝒯\\tau\\in\\mathcal\{T\}:Δτ\(a\)=eGθ\(x^t\+1a,yg,τ\)−eGθ\(xt,yg,τ\)\\Delta\_\{\\tau\}\(a\)=e^\{G\_\{\\theta\}\(\\widehat\{x\}\_\{t\+1\}^\{\\,a\},y\_\{g\},\\tau\)\}\-e^\{G\_\{\\theta\}\(x\_\{t\},y\_\{g\},\\tau\)\}\.
- 7\.end for
- 8\.Selectat∗a\_\{t\}^\{\*\}by[Eq\.8](https://arxiv.org/html/2609.38383#S2.E8), using maximum progress asβ→0\+\\beta\\to 0^\{\+\}or mean progress asβ→∞\\beta\\to\\infty\.
- 9\.Execute onlyat∗a\_\{t\}^\{\*\}in the environment\.
- 10\.Receive the actualxt\+1x\_\{t\+1\}and sett←t\+1t\\leftarrow t\+1\.
- 11\.end while
PyTorch\-Style Pseudocode
The compact version below shows one update for each model and one planning decision\.pair\_scorereturns aB×BB\\times Bmatrix of temporal scores at a shared lag;goal\_scorereturns candidate\-by\-horizon scores\. HereFdenotestorch\.nn\.functional\. Planning uses every integer horizon from 1 toτmax\\tau\_\{\\max\}\.
1deftrain\_temporal\(current,future,tau\):
2logits=temporal\.pair\_score\(current,future,tau\)
3logits=logits\-math\.log\(len\(current\)\-1\)
4positives=logits\.diagonal\(\)
5negatives=F\.softplus\(logits\)\.sum\(1\)
6negatives=negatives\-F\.softplus\(positives\)
7loss=\(F\.softplus\(\-positives\)\+negatives\)\.mean\(\)
8temporal\_opt\.zero\_grad\(\)
9loss\.backward\(\)
10temporal\_opt\.step\(\)
11
12deftrain\_dynamics\(state,action,next\_state\):
13prediction=dynamics\(state,action\)
14error=\(prediction\-next\_state\)\.square\(\)
15loss=error\.flatten\(1\)\.sum\(1\)\.mean\(\)
16dynamics\_opt\.zero\_grad\(\)
17loss\.backward\(\)
18dynamics\_opt\.step\(\)
19
20@torch\.no\_grad\(\)
21defchoose\_action\(state,goal,actions,taus,planner\):
22states=state\.expand\(len\(actions\),\-1\)
23next\_states=dynamics\(states,actions\)
24here=temporal\.goal\_score\(state,goal,taus\)\.exp\(\)
25there=temporal\.goal\_score\(next\_states,goal,taus\)\.exp\(\)
26progress=there\-here
27ifplanner=="GP":
28value=progress\.amax\(1\)
29else:
30value=progress\.mean\(1\)
31returnactions\[value\.argmax\(\)\]
At test time, both models are in evaluation mode\. Execute the returned action, observe the new state, and repeat\. The dynamics call illustrates learned state prediction; oracle experiments instead obtain candidate outcomes from the simulator\.
## Appendix DExperiments
### D\.1Representation Probes
#### Representation Probes and Visualizations\.
[Figures4](https://arxiv.org/html/2609.38383#A4.F4)and[4](https://arxiv.org/html/2609.38383#A4.T4)compare embedding distances with maze geometry at different horizons\. Near pairs are at most eight geodesic steps apart\. Each t\-SNE panel is fitted separately and aligned to the maze by an orthogonal Procrustes transform; projection coordinates are not physical distances\.
Figure 4:Learned embeddingsh\(x,τ\)h\(x,\\tau\)on the Large \(top\) and Giant \(bottom\) state mazes; the left column colours each cell by its position and the other columns keep those colours\. Every panel is rotated and reflected onto the maze by Procrustes, since t\-SNE leaves the orientation undetermined\.Table 4:Spearman correlation between embedding distance and geodesic distance for near \(≤8\\leq 8steps\) and far pairs, Large maze\.
#### Three\-dimensional views\.
[Figure5](https://arxiv.org/html/2609.38383#A4.F5)supplements the Giant\-maze 2D projections in[Fig\.1](https://arxiv.org/html/2609.38383#S3.F1)B with 3D t\-SNE views of both mazes atτ=1\\tau=1,1616, and40964096\. We use perplexity 30, PCA initialization, and random seed 0, fitting each maze and horizon separately\. Rotation/reflection and uniform scaling align each fit to the maze plane without flattening its third dimension\. Within each maze, all three panels share a camera angle and coordinate scale\. The projections illustrate neighbourhood structure, not physical distances or comparable coordinates across horizons\.
Figure 5:3D t\-SNE of state embeddings on Large \(top\) and Giant \(bottom\)\. Atτ=1\\tau=1, 3D views reveal connectivity obscured in 2D\. Colours match the maze regions at left; axes are not physical coordinates\.
### D\.2Demonstrations
#### Planning demonstrations\.
[Figure1](https://arxiv.org/html/2609.38383#S3.F1)C uses all integer horizons from 1 to 512 on Large \(tasks 1 and 3\) and from 1 to 2048 on Giant \(tasks 1 and 2\), with one\-step lookahead and execution \(0\.2 units\)\. These illustrative rollouts use different horizon ranges and a different planning radius from the aggregate evaluations in[Tab\.1](https://arxiv.org/html/2609.38383#S4.T1)\. For GP, path colour marks the horizon with the largest improvement for the chosen action; for PAP, it marks the horizon contributing most to the chosen candidate’s pooled score, not a separately selected planning horizon\. The recorded values are not temporally smoothed\. Matching marker colours pair each start with its recorded episode goal\.
#### Manipulation demonstrations\.
[Figure6](https://arxiv.org/html/2609.38383#A4.F6)shows three planned episodes of different lengths, with evenly spaced frames followed by the goal image\.
Figure 6:Three held\-out cube\-single episodes planned from pixels, one per row, each showing eight evenly spaced executed steps with the goal image last\. The cube travels 0\.455, 0\.277, and 0\.325 m, respectively; in every row the planner reaches the cube, grasps it, and carries it to the goal\.相似文章
基于潜世界模型的强化规划
本文介绍了强化规划,一种利用潜世界模型学习改进多步骤规划的方法,在视觉导航和机器人操作等任务中取得了近乎完美的成功率,并且效率显著高于手工设计的算法。
基于内在好奇心的强化学习中的内生探索
本文提出了一种使用内在好奇心进行内生探索的强化学习框架,适用于非平稳环境,在如LunarLander-v2和BipedalWalker-v3等基准测试上,与PPO和ICM等算法相比,达到了具有竞争力的性能。
学习探索:通过探索感知策略优化扩展代理推理
本文提出一种探索感知的强化学习框架,使LLM代理仅在不确定性高时自适应探索,从而提升在基于文本和基于GUI的基准测试上的性能。
在线规划,离线学习:通过基于模型的控制实现高效学习和探索
OpenAI 提出 POLO(在线规划,离线学习)框架,结合基于模型的控制、价值函数学习和协调探索,能够在人形机器人运动和灵巧手部操纵等复杂控制任务中实现高效学习,同时最小化真实世界经验需求。
部分可观测环境中的生成模型预测规划导航
本文介绍了BeliefDiffusion,一种结合扩散模型表示多模态信念分布和使用模型预测控制在部分可观测环境中进行规划的框架,相比基线方法取得了更好的导航成功率和路径效率。