Learning to Route in Visual Space via Multi-Step Embedding Retrieval
Summary
This paper introduces VHop, a data generation framework and benchmark for multi-step visual retrieval, and VHop-Router, an end-to-end trained autoregressive multi-step retriever that operates in visual latent space, boosting retrieval from under 5% to 76.3% and improving agentic search success by 52.7% while drastically cutting tokens and API payload.
View Cached Full Text
Cached at: 10/01/26, 09:42 AM
# Learning to Route in Visual Space via Multi-Step Embedding Retrieval
Source: [https://arxiv.org/html/2609.38743](https://arxiv.org/html/2609.38743)
\\uselogo
Mingyuan ZhouAffiliation:The University of Texas at AustinJiaxing WuAffiliation:\\thepa
###### Abstract
LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single\-step retrievers\. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results\. We hypothesize that offloading multi\-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck\. To study this systematically, we introduceVHop, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning\. Using this framework, we developVHop\-Router, an end\-to\-end training pipeline—combining supervised fine\-tuning, online imitation learning, and reinforcement learning—that transforms a standard embedding model into an autoregressive multi\-step retriever\. Operating directly in the visual latent space,VHop\-Routerretrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries\. Experiments showVHop\-Routerboosts retrieval performance from under 5% to 76\.3%\. In agentic search, it improves task success rates by 52\.7% and reduces the average token length by 61% from 1886 to 728, whereas upgrading the agent yields only a 3\.7% gain\. Compared to a strong baseline where the agent retrieves the top 50 results per step,VHop\-Routermaintains superior performance while reducing in\-context images by23×23\\timesand cutting the cumulative API payload by35×35\\times\. The models also generalize robustly to unseen difficulty levels and realistic test sets\. Ultimately,VHopandVHop\-Routerprovide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact\.
## 1Introduction
Figure 1:Task performance and upgrade gains\.Agent denotes Gemini 3\.5 Flash\. \(a\) Success rate versus generated tokens plus retrieval hops\. Standalone Qwen3 uses five greedy retrieval steps\. \(b\) Gains over Agent \+ Qwen3 from upgrading the agent model to Gemini 3\.1 Pro or replacing the retrieval tool withVHop\-Router,keeping the other component fixed\.Agentic search equips large language model \(LLM\) agents with access to external knowledge bases, including up\-to\-date information and private data unavailable during training\([Yao et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib9);[Trivedi et al\., 2023](https://arxiv.org/html/2609.38743#bib.bib4)\)\. Recent work fine\-tunes LLMs to improve their reasoning and multi\-step search\([Jin et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib7)\)\. These improvements still depend on a fixed retrieval tool to supply relevant evidence\. Because the agent sees only the few results returned by this tool, it cannot reason over evidence it never receives\.
This bottleneck is especially important in visual search, where each retrieved image may reveal the clue for the next step\. When locating a misplaced item in a photo album, visual clues can be hard to express in a textual query \(Figure[2](https://arxiv.org/html/2609.38743#S1.F2)a\) or be omitted from captions \(Figure[2](https://arxiv.org/html/2609.38743#S1.F2)b\)\. A standard embedding retriever performs one retrieval step per call, leaving the agent to build the search chain from the returned images\. In Figure[2](https://arxiv.org/html/2609.38743#S1.F2)c, the retriever ranks the required next image 39th, outside the top\-five results, leaving the red piggy bank needed for the following hop out of the agent’s context\. In contrast, embedding retrievers can score candidates across the entire indexed corpus\.
Figure 2:Visual retrieval bottlenecks\.\(a\) Exact symbols can be hard to describe\. \(b\) The caption mentions “metal tools” but does not name the boxed pliers\. \(c\) Following the green sewing machine requires the ground\-truth next image\. The retriever ranks it 39th, so the red piggy bank needed for the following hop is not shown\.We hypothesize that the main bottleneck in these scenarios is the retrieval tool itself\. We therefore train this tool to navigate directly in visual embedding space over multiple steps, without agent fine\-tuning\. Shifting multi\-hop search from the agent to the embedding model \(replacing a single\-step model with a multi\-step one\) significantly boosts task success rates while minimizing token length \(Figure[1](https://arxiv.org/html/2609.38743#S1.F1)a\)\. Upgrading to a stronger agent improves performance by only 3\.7%, whereas a multi\-step embedding model yields a 52\.7% increase \(Figure[1](https://arxiv.org/html/2609.38743#S1.F1)b\)\. This validates that enhancing the retrieval tool is the far more effective approach\. The following sections detail our data, evaluation, and training methods\.
Table 1:Comparison with related QA and visual\-retrieval benchmarks\.I\+T denotes image\(s\) and text\. Visual hop clues reveal the next target in an intermediate image; planning involves avoiding or recovering from dead ends; symbolic verification checks answers against structured facts; training recipe indicates an evaluated solver\-training procedure\.BenchmarkQueryOutputVisualhop cluesPlanningabilitySymbolicverificationDifficultycontrolsTrainingrecipe[24](https://arxiv.org/html/2609.38743#bib.bib2)TextText––––✓\\checkmark[18](https://arxiv.org/html/2609.38743#bib.bib3)TextText––––✓\\checkmark[3](https://arxiv.org/html/2609.38743#bib.bib16)TextText––✓\\checkmarkDepth, size–\[2pt/2pt\][2](https://arxiv.org/html/2609.38743#bib.bib12)TextText––––✓\\checkmark[11](https://arxiv.org/html/2609.38743#bib.bib13)I\+TText–––––[22](https://arxiv.org/html/2609.38743#bib.bib21)TextYes/no–––Set size✓\\checkmark[20](https://arxiv.org/html/2609.38743#bib.bib22)I\+TOption–––––[8](https://arxiv.org/html/2609.38743#bib.bib1)I\+TText––✓\\checkmarkScene/questioncomplexity✓\\checkmark\[2pt/2pt\][9](https://arxiv.org/html/2609.38743#bib.bib6)ImageImages––––✓\\checkmark[4](https://arxiv.org/html/2609.38743#bib.bib11)ImageTrack––––✓\\checkmarkVHop\(ours\)I\+TImage✓\\checkmark✓\\checkmark✓\\checkmarkHops, levels,corpus size✓\\checkmark
To train and evaluate multi\-step embedding models, we introduceVHop, a flexible data generation framework for visual multi\-hop image retrieval\. It connects images via shared objects and offers controllable hop counts, object distractors, and varying task difficulty levels \(Table[1](https://arxiv.org/html/2609.38743#S1.T1)\)\. For reliable training rewards and rigorous evaluation, verifiers use structured records of objects and relations to ensure the ground\-truth labels satisfy all detailed query instructions and include all valid answers\.
Building on these data and benchmarks, we developVHop\-Router, an end\-to\-end training framework that transforms a standard single\-step embedding model into an autoregressive multi\-step retriever\. Training combines supervised fine\-tuning \(SFT\), online imitation learning \(Online IL\), and reinforcement learning with verifiable rewards \(RLVR\)\. At inference,VHop\-Routerautoregressively retrieves images and returns a linked image chain in a single tool call based on the initial query and completed hops, without agent text queries for intermediate hops \(Figure[3](https://arxiv.org/html/2609.38743#S2.F3)\)\. Because LLM agents remain completely untouched,VHop\-Routeris compatible with any LLM\-based agent and preserves all native LLM capabilities\.
Through extensive experiments, we demonstrate thatVHop\-Routermodels achieve significant performance gains in both retrieval\-only and agentic search tasks while maintaining superior efficiency\. Furthermore, these models generalize well to unseen tasks across varying difficulty levels and hop counts, as well as to realistic test sets, ensuring the practical utility of the trained multi\-step embedders\.
Our contributions are four\-fold:
- •We formulate*multi\-hop image retrieval in visual latent space*, a novel task where an embedding model navigates independently in the latent embedding space to achieve multi\-hop image retrieval\.
- •We introduceVHop, a flexible and controllable data generation framework and benchmark\. It provides verified ground\-truth trajectories across varying, easily extendable difficulty levels\.
- •We developVHop\-Router, an end\-to\-end framework for training autoregressive multi\-step embedding models\.
- •We demonstrate that offloading multi\-step search from the agent to the retriever offers a highly effective and efficient solution for overcoming the bottleneck in agentic search, improving the success rate while significantly reducing token length and number of images in the LLM context\.
## 2Related Work
Multi\-hop Dense Retrieval\([Xiong et al\., 2020](https://arxiv.org/html/2609.38743#bib.bib10)\)and Q\-RAG\([Sorokin et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib8)\)train retrievers for multi\-step text search\.VHop\-Routershares this focus on learning the retrieval tool, but learns visual matching and recovery from dead ends\. IRCoT\([Trivedi et al\., 2023](https://arxiv.org/html/2609.38743#bib.bib4)\)and Search\-R1\([Jin et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib7)\)use LLM reasoning to guide retrieval; our policy retrieves image chains directly in latent space without generating intermediate text queries or fine\-tuning LLMs\. ILIAS\([Kordopatis\-Zilos et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib6)\)and Ego4D’s visual queries\([Grauman et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib11)\)search for a given object, whereasVHoprequires discovering the next target from intermediate images\. WebQA\([Chang et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib12)\)and Visual Haystacks\([Wu et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib21)\)evaluate question answering with visual evidence; our task focuses on retrieving the answer image through a chain of visual clues\. Finally, CLEVR\([Johnson et al\., 2017](https://arxiv.org/html/2609.38743#bib.bib1)\)and PhantomWiki\([Gong et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib16)\)control reasoning within scenes and text corpora, respectively;VHopcontrols retrieval paths across images, including hop counts and dead ends, and provides verified answers\. Appendix[A](https://arxiv.org/html/2609.38743#A1)provides further discussion of these methods, benchmarks, and training approaches\.
Figure 3:Overview ofVHopandVHop\-Router\.Left:An L4 search example with precise and vague prompts, dead ends, and backtracking, followed by benchmark construction\.Right:A complete L1 retrieval chain and small image comparisons for L2, L3, and L5, with theVHop\-Routertraining stages below\. Boxes, arrows, and detail crops are reader annotations\. Examples of all difficulty levels appear in Figure[9](https://arxiv.org/html/2609.38743#A2.F9)in Appendix[B](https://arxiv.org/html/2609.38743#A2)\.
## 3VHop: Controllable Multi\-Hop Visual Search in Latent Space
Consider that in real life, photo albums capture everyday moments and can help us locate misplaced objects\. A multi\-hop retriever can follow visual clues across these photos to find an image showing where an object was last seen, as illustrated in Figure[7](https://arxiv.org/html/2609.38743#S5.F7)\. We formalize this search process throughVHop, a controlled setting for studying multi\-hop visual retrieval\.
Task definition\.Given a starting image and text instruction, the solver retrieves a sequence of images from a corpus, using each intermediate image to determine what to search for next\. Each retrieval is a*hop*, and the final retrieved image is the answer\. In Figure[3](https://arxiv.org/html/2609.38743#S2.F3), the solver first retrieves an image in which the starting object appears to the right of a yellow objectAA, and then usesAAas the visual clue for the next hop\. We consider two types of prompts\.*Precise prompts*specify each step using color and spatial constraints without directly naming the objects, providing a controlled setting for evaluating visual retrieval under explicit guidance\.*Vague prompts*omit these constraints and step\-by\-step instructions, reflecting settings in which users cannot specify the intermediate retrieval steps\. We train and evaluate both settings; see details in Section[4\.1](https://arxiv.org/html/2609.38743#S4.SS1)\.
##### Controllable difficulty levels\.
VHopprovides five core difficulty levels \(L1–L5; Figure[3](https://arxiv.org/html/2609.38743#S2.F3)\), each retaining earlier requirements while adding constraints or distractors\. L1–L5 link images through the same physical object, with L4 introducing dead ends that require exploration and recovery\. In Figure[3](https://arxiv.org/html/2609.38743#S2.F3), both a yellow tool and a yellow bowling pin appear to the left of the starting object, satisfying the first step, but only the tool permits a valid continuation\. A solver choosing the bowling pin must therefore backtrack and try another candidate\. The hop count, corpus size, and number of distractors are also configurable\. Matching can also be extended to logos and printed text; Appendix[F\.4](https://arxiv.org/html/2609.38743#A6.SS4)describes these extensions and reports their results\.
##### Symbolic generation, rendering, and verification\.
We first construct each image collection symbolically, specifying the object pool, object appearances and attributes, co\-occurrence and spatial relations, the answer sequence, and distractors\. For precise instructions, an independent verifier checks that exactly one sequence satisfies the query across the entire symbolic collection\. If an instruction leaves relations unspecified \(vague prompt\), the verifier records all valid answers, any of which is accepted during evaluation\. We then render the images and check the consistency of object identities, colors, and spatial relations\. Training and evaluation use disjoint sets of object appearances\. Appendices[B](https://arxiv.org/html/2609.38743#A2)and[C](https://arxiv.org/html/2609.38743#A3)provide generation details and formal guarantees, including trajectory uniqueness for precise instructions; Section[4\.1](https://arxiv.org/html/2609.38743#S4.SS1)describes scoring rules and search budgets\.
## 4VHop\-Router: Training Framework
We developedVHop\-Router, an end\-to\-end training framework that enables embedding models to navigate in visual latent spaces autoregressively through a three\-stage process: Supervised fine\-tuning \(SFT\) teaches valid retrieval chains, online imitation learning \(Online IL\) teaches recovery and stopping, and reinforcement learning with verifiable rewards \(RLVR\) optimizes final\-answer success\. The trained policy can search independently and return image chains to VLM agents\.
### 4\.1Experimental Setup
Table 2:Data allocation\.Query counts use precise three\-hop prompts\.##### Training Data\.
All policy\-training data are generated at L4, the first level with locally valid but globally incorrect branches, requiring exploration and backtracking\. Data are organized into*worlds*: a world is one independently generated image corpus together with the queries posed over it, and every query is answered using only the corpus of its own world \(Appendix[B](https://arxiv.org/html/2609.38743#A2)\)\. SFT/Online IL, RLVR, and evaluation use mutually disjoint worlds with distinct objects, backgrounds, and queries\. Gold trajectories supervise SFT and Online IL and are withheld for RLVR \(Table[2](https://arxiv.org/html/2609.38743#S4.T2)\)\.
##### Evaluation setting\.
Table 3:Evaluation sets\.We trainVHop\-Routerexclusively on L4 and evaluate it on L1–L5 without further fine\-tuning\. Each non\-L4 level contains approximately 100 three\-hop queries: L1–L3 test transfer to simpler settings, while L5 tests transfer with visually similar object instances\. For L4, we use three independent evaluation worlds with 100 queries each for a more robust assessment\. All evaluation worlds are disjoint from training, and non\-L4 results measure whether the retrieval, backtracking, and stopping behaviors learned on L4 generalize across difficulty levels, corpora, and visual conditions\. Table[3](https://arxiv.org/html/2609.38743#S4.T3)summarizes the statistics\.
##### Baselines\.
We consider two types of baselines\.*\(1\) Retriever\-only methods*search without an LLM agent\. We evaluate the open\-source Qwen3\-VL\-Embedding\-2B \(Qwen3\) and the proprietary Gemini Embedding 2 \(GE2\)\([Shanbhogue et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib26)\)under three strategies\.*Single\-shot*retrieval encodes the original multimodal query and returns the top\-5 images\.*Greedy*retrieval combines the original instruction with the latest retrieved image to select one new image at each of five steps\.*History\-based*retrieval instead encodes the complete retrieval history\. Caption\-based variants use cached Gemini\-3\.1\-Flash\-Lite captions to test whether textual descriptions preserve the visual clues needed for multi\-hop retrieval\.
*\(2\) Agentic\-search methods*use an LLM to set up tool calls with observations\([Yao et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib9)\)\. We use Gemini 3\.5 Flash with GE2 or Qwen3: the agent inspects the returned images and formulates its next query\. To isolate the effect of the retrieval tool, we keep the agent fixed and replace its retriever withVHop\-Router, which returns an image chain for the agent to inspect before answering or continuing the search\. We also evaluate stronger agentic baselines by replacing Flash with Gemini 3\.1 Pro \(Section[5\.2](https://arxiv.org/html/2609.38743#S5.SS2)\) or returning up to 50 images per retrieval call \(Section[5\.3](https://arxiv.org/html/2609.38743#S5.SS3)\)\.
##### Metrics\.
We evaluate both*task performance*and*search efficiency*\. Task performance is measured by success rate\. Embedding and caption\-based baselines succeed if a valid answer appears among the five retrieved images\. StandaloneVHop\-Routersucceeds if it stops with a valid answer at the top of the stack, while agentic systems succeed if their declared final answer is valid\. For efficiency, we report the total number of generated tokens, image inputs, and API calls per query\. Agentic search allows up to1616tool calls and8,0008\{,\}000output tokens per model call\. EachVHop\-Routerrollout allows up to1616actions; retrieval, backtracking, and stopping each count as one action\.
### 4\.2Autoregressive Retrieval Policy
##### State and actions\.
Letxxbe the text instruction,q0q\_\{0\}the query image, and𝒞\\mathcal\{C\}the image corpus\. At decisiontt, the state consists ofxxand the active stackσt=\(q0,q1,…,qdt\)\\sigma\_\{t\}=\(q\_\{0\},q\_\{1\},\\ldots,q\_\{d\_\{t\}\}\), whereq1,…,qdtq\_\{1\},\\ldots,q\_\{d\_\{t\}\}are the images retained after any backtracking anddtd\_\{t\}is the current hop count\. The action space is\{Select\(a\):a∈𝒞\}∪\{Backtrack,Stop\}\\\{\\textsc\{Select\}\(a\):a\\in\\mathcal\{C\}\\\}\\cup\\\{\\textsc\{Backtrack\},\\textsc\{Stop\}\\\}\.Select\(a\)\(a\)appends imageaato the stack;Backtrackremoves its last retrieved image;Stopends the search and returns its current image stack\. The query imageq0q\_\{0\}always remains in the active stack\.
##### Select
A state encoderEsE\_\{s\}jointly encodes the instructionxxand active stackσt\\sigma\_\{t\}assts\_\{t\}, while an action encoderEaE\_\{a\}independently encodes each candidate imageaa:
st=Es\(x,σt\),ea=Ea\(a\),∥st∥2=∥ea∥2=1\.s\_\{t\}=E\_\{s\}\(x,\\sigma\_\{t\}\),\\qquad e\_\{a\}=E\_\{a\}\(a\),\\qquad\\lVert s\_\{t\}\\rVert\_\{2\}=\\lVert e\_\{a\}\\rVert\_\{2\}=1\.\(1\)Both encoders start from Qwen3\-VL\-Embedding\-2B and use LoRA\([Hu et al\., 2021](https://arxiv.org/html/2609.38743#bib.bib24)\)\. We freezeEaE\_\{a\}after SFT and reuse its cached corpus embeddings during Online IL, RLVR, and inference\. Candidates are scored by cosine similarity,zt\(a\)=st⊤eaz\_\{t\}\(a\)=s\_\{t\}^\{\\top\}e\_\{a\}\. After masking ineligible images, we select from the top\-KKcandidates𝒞t\\mathcal\{C\}\_\{t\}using a softmax with temperatureτ\\tau:
πθsel\(a∣st,𝒞t\)=exp\(zt\(a\)/τ\)∑a′∈𝒞texp\(zt\(a′\)/τ\)\.\\pi\_\{\\theta\}^\{\\mathrm\{sel\}\}\(a\\mid s\_\{t\},\\mathcal\{C\}\_\{t\}\)=\\frac\{\\exp\(z\_\{t\}\(a\)/\\tau\)\}\{\\sum\_\{a^\{\\prime\}\\in\\mathcal\{C\}\_\{t\}\}\\exp\(z\_\{t\}\(a^\{\\prime\}\)/\\tau\)\}\.\(2\)
##### BacktrackandStop\.
Two learned gates control stopping and backtracking\. Each is a two\-layer MLP over\[st;ct\]\[s\_\{t\};c\_\{t\}\], wherect∈ℝ4c\_\{t\}\\in\\mathbb\{R\}^\{4\}contains four retrieval statistics, such as the maximum candidate score \(Appendix[D\.3](https://arxiv.org/html/2609.38743#A4.SS3)\)\. During stochastic rollouts, the gate probabilities are
ptstop=sigmoid\(gstop\(\[st;ct\]\)/τg\),ptbt=sigmoid\(gbt\(\[st;ct\]\)/τg\),p\_\{t\}^\{\\mathrm\{stop\}\}=\\operatorname\{sigmoid\}\\\!\\left\(g\_\{\\mathrm\{stop\}\}\(\[s\_\{t\};c\_\{t\}\]\)/\\tau\_\{g\}\\right\),\\qquad p\_\{t\}^\{\\mathrm\{bt\}\}=\\operatorname\{sigmoid\}\\\!\\left\(g\_\{\\mathrm\{bt\}\}\(\[s\_\{t\};c\_\{t\}\]\)/\\tau\_\{g\}\\right\),\(3\)Here,τg\\tau\_\{g\}is the gate temperature\. Writeztstop=𝟏\[at=Stop\]z\_\{t\}^\{\\mathrm\{stop\}\}=\\mathbf\{1\}\[a\_\{t\}=\\textsc\{Stop\}\]andztbt=𝟏\[at=Backtrack\]z\_\{t\}^\{\\mathrm\{bt\}\}=\\mathbf\{1\}\[a\_\{t\}=\\textsc\{Backtrack\}\]for the two action indicators\. The full policy distribution is
πθ\(at∣st\)=\(ptstop\)ztstop\[\(1−ptstop\)\(ptbt\)ztbt\(\(1−ptbt\)πθsel\(a∣st,𝒞t\)\)1−ztbt\]1−ztstop\.\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)=\\bigl\(p\_\{t\}^\{\\mathrm\{stop\}\}\\bigr\)^\{z\_\{t\}^\{\\mathrm\{stop\}\}\}\\Bigl\[\\bigl\(1\-p\_\{t\}^\{\\mathrm\{stop\}\}\\bigr\)\\bigl\(p\_\{t\}^\{\\mathrm\{bt\}\}\\bigr\)^\{z\_\{t\}^\{\\mathrm\{bt\}\}\}\\Bigl\(\\bigl\(1\-p\_\{t\}^\{\\mathrm\{bt\}\}\\bigr\)\\,\\pi\_\{\\theta\}^\{\\mathrm\{sel\}\}\(a\\mid s\_\{t\},\\mathcal\{C\}\_\{t\}\)\\Bigr\)^\{1\-z\_\{t\}^\{\\mathrm\{bt\}\}\}\\Bigr\]^\{1\-z\_\{t\}^\{\\mathrm\{stop\}\}\}\.\(4\)At evaluation time, the policy acts deterministically: it stops ifptstop≥0\.5p\_\{t\}^\{\\mathrm\{stop\}\}\\geq 0\.5, otherwise backtracks ifptbt≥0\.5p\_\{t\}^\{\\mathrm\{bt\}\}\\geq 0\.5, and otherwise selects the highest\-scoring image\.
A rolloutρ=\(a0,…,aT\)\\rho=\(a\_\{0\},\\ldots,a\_\{T\}\)records all decisions, including selections later undone by backtracking\. Its prefixρ<t\\rho\_\{<t\}determines the active stackσt\\sigma\_\{t\}before decisiontt\. Its log\-probability underπθ\\pi\_\{\\theta\}is
logπθ\(ρ∣x,q0\)=∑t=0Tlogπθ\(at∣st\)\.\\log\\pi\_\{\\theta\}\(\\rho\\mid x,q\_\{0\}\)=\\sum\_\{t=0\}^\{T\}\\log\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)\.\(5\)
### 4\.3Training Pipeline
The three stages address successive requirements of sequential retrieval: ranking the correct continuation, recovering from mistakes, and optimizing final\-answer success\.
#### 4\.3\.1Supervised Fine\-Tuning
SFT teaches next\-hop retrieval from gold trajectory prefixes\. For a gold trajectoryg1:Hg\_\{1:H\}and hopk∈\{1,…,H\}k\\in\\\{1,\\ldots,H\\\}, the state encoder reads the instructionxx, query imageq0q\_\{0\}, and prefixg1:k−1g\_\{1:k\-1\}\. The InfoNCE objective\([Oord et al\., 2018](https://arxiv.org/html/2609.38743#bib.bib25)\)aligns the resulting state with the next gold imagegkg\_\{k\}:
sk−1\\displaystyle s\_\{k\-1\}=Es\(x,\(q0,g1:k−1\)\),\\displaystyle=E\_\{s\}\\\!\\left\(x,\(q\_\{0\},g\_\{1:k\-1\}\)\\right\),\\qquadℒk\\displaystyle\\mathcal\{L\}\_\{k\}=−logexp\(⟨sk−1,egk⟩/τ\)∑a∈𝒞exp\(⟨sk−1,ea⟩/τ\)\.\\displaystyle=\-\\log\\frac\{\\exp\\\!\\left\(\\langle s\_\{k\-1\},e\_\{g\_\{k\}\}\\rangle/\\tau\\right\)\}\{\\sum\_\{a\\in\\mathcal\{C\}\}\\exp\\\!\\left\(\\langle s\_\{k\-1\},e\_\{a\}\\rangle/\\tau\\right\)\}\.\(6\)Training proceeds in two phases: we first freezeEaE\_\{a\}and update onlyEsE\_\{s\}, then jointly optimize both encoders\. This warm\-up stabilizes training and prevents collapse \(Figure[10](https://arxiv.org/html/2609.38743#A4.F10)\)\.
#### 4\.3\.2Online Imitation Learning for Recovery and Stopping
Online IL trains on states visited during search to teach recovery and stopping\([Ross et al\., 2011](https://arxiv.org/html/2609.38743#bib.bib17)\)\. Each iteration collects fresh rollouts using the current policy, oracle guidance, and dead\-end exploration\. For each prefixρ<t\\rho\_\{<t\}, we recover the active stackσt\\sigma\_\{t\}, encode it assts\_\{t\}, and obtain the oracle actionat⋆a\_\{t\}^\{\\star\}\. We freezeEaE\_\{a\}and updateEsE\_\{s\}and the gates using the per\-state loss in Eq\.[7](https://arxiv.org/html/2609.38743#S4.E7)\. Only states collected in the current iteration are used for the update\. Algorithm[1](https://arxiv.org/html/2609.38743#alg1)summarizes this procedure\.
For oracle labelsytbt=𝟏\[at⋆=Backtrack\]y\_\{t\}^\{\\mathrm\{bt\}\}=\\mathbf\{1\}\[a\_\{t\}^\{\\star\}=\\textsc\{Backtrack\}\]andytstop=𝟏\[at⋆=Stop\]y\_\{t\}^\{\\mathrm\{stop\}\}=\\mathbf\{1\}\[a\_\{t\}^\{\\star\}=\\textsc\{Stop\}\], the loss is
ℓt=−\(1−ytbt\)\(1−ytstop\)logπθsel\(gt∣st,𝒞t\)\+Bt\(ytbt\)\+St\(ytstop\)\.\\ell\_\{t\}=\-\(1\-y\_\{t\}^\{\\mathrm\{bt\}\}\)\(1\-y\_\{t\}^\{\\mathrm\{stop\}\}\)\\log\\pi\_\{\\theta\}^\{\\mathrm\{sel\}\}\(g\_\{t\}\\mid s\_\{t\},\\mathcal\{C\}\_\{t\}\)\+B\_\{t\}\(y\_\{t\}^\{\\mathrm\{bt\}\}\)\+S\_\{t\}\(y\_\{t\}^\{\\mathrm\{stop\}\}\)\.\(7\)The selection term applies only when the oracle selects imagegtg\_\{t\}, which is added to𝒞t\\mathcal\{C\}\_\{t\}if absent\. The termsBt\(y\)B\_\{t\}\(y\)andSt\(y\)S\_\{t\}\(y\)are binary cross\-entropy losses for the backtrack and stop gates \(Eq\.[3](https://arxiv.org/html/2609.38743#S4.E3)\), with targety∈\{0,1\}y\\in\\\{0,1\\\}\. Appendix[D](https://arxiv.org/html/2609.38743#A4)specifies their weights and action masks\.
Algorithm 1Online IL training1:SFT checkpoint; oracle
π⋆\\pi^\{\\star\}
2:Freeze action encoder
EaE\_\{a\}
3:foreach training iterationdo
4:Collect fresh rollouts
ρ\\rho
5:foreach rollout prefix
ρ<t\\rho\_\{<t\}do
6:Replay
ρ<t\\rho\_\{<t\}to recover
σt\\sigma\_\{t\}
7:Encode
st=Es\(x,σt\)s\_\{t\}=E\_\{s\}\(x,\\sigma\_\{t\}\)
8:Obtain
at⋆=π⋆\(x,σt\)a\_\{t\}^\{\\star\}=\\pi^\{\\star\}\(x,\\sigma\_\{t\}\)
9:endfor
10:Update
EsE\_\{s\}and gates using
ℓt\\ell\_\{t\}
11:endfor
#### 4\.3\.3Outcome\-Based Reinforcement Learning
Starting from the Online IL checkpoint, RLVR trainsEsE\_\{s\}and the two gate MLPs on the separate RLVR data in Table[2](https://arxiv.org/html/2609.38743#S4.T2), withEaE\_\{a\}fixed\. For each query, we sampleGGrolloutsρg\\rho\_\{g\}and assignrg=1r\_\{g\}=1only when the policy stops with a valid answer, and00otherwise\. Each rollout receives a leave\-one\-out advantageAgA\_\{g\}\([Ahmadian et al\., 2024](https://arxiv.org/html/2609.38743#bib.bib19)\)\. We use a PPO\-style clipped surrogate\([Schulman et al\., 2017](https://arxiv.org/html/2609.38743#bib.bib18)\)with DAPO\-style asymmetric clipping\([Yu et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib20)\)\. Algorithm[3](https://arxiv.org/html/2609.38743#alg3)summarizes training, and Figure[11](https://arxiv.org/html/2609.38743#A4.F11)shows its dynamics\. Appendix[D\.4](https://arxiv.org/html/2609.38743#A4.SS4)details the loss terms, clipping, and KL penalties\.
## 5Results
We evaluate combinations of four independent configurations\. First,task difficultytests capability across query complexities\. Second, we vary thenumber of hops\. Third, we compareprecise prompts—which give detailed instructions and constrain valid answers—againstvague promptsthat omit step\-by\-step guidance to mimic real\-world searches\. Finally, we train withfixedormixedhop counts\. By default, training uses L4 \(requiring planning and navigation around dead ends\), 1–3 mixed hops, and precise prompts\. We evaluate L1–L5 and ablate settings to test generalization, alongside a test set reflecting realistic user scenarios\.
### 5\.1Significant Improvements Across Difficulty Levels
Table[4](https://arxiv.org/html/2609.38743#S5.T4)reports three\-hop results across L1–L5 for standalone retrieval and agentic search\. We trainVHop\-Routerwith precise or vague prompts, using either three\-hop\-only or mixed\-hop data\. On all these four training configurations,VHop\-Routerimproved retrieval and agentic search across 5 difficulty levels significantly over baseline retrievers\.
With our default mixed\-hop training with precise prompts, standaloneVHop\-Routerreaches76\.3%success on L4, compared with3\.7%for the strongest standard single\-step retriever\. It also improves agentic search: replacing Qwen3 withVHop\-Routerwhile keeping the agent fixed as Gemini 3\.5 Flash raises success from32\.7% to 85\.4% at in\-domain L4 difficultyand from28\.0% to 84\.0% on L5 with zero\-shot transfer\.
Table 4:Three\-hop success rate \(%\) across L1–L5\.Left: standalone retrieval\. Right: Gemini\-3\.5\-Flash with different tools\. Blocks \(a\) and \(b\) use precise and vague prompts, respectively\. Shaded columns showVHop\-Routerwith three\-hop\-only and mixed\-hop training; other baselines are shared across both regimes\. Within each row,boldandunderliningmark the best and second\-best scores on each side of the dashed rule; ties share a rank\.Retriever onlyAgentic searchEmbedding retrievalCaption→\\rightarrowembeddingVHop\-RouterGE2Qwen3\+\+VHop\-RouterGE2Qwen3GE2Qwen333\-hopmixed33\-hopmixedLevel1\-shotGreedyHist\.1\-shotGreedyHist\.1\-shotGreedyHist\.1\-shotGreedyHist\.\(a\) Precise promptL1 \(re\-ID\)2\.06\.12\.05\.19\.16\.11\.00\.01\.06\.14\.03\.066\.761\.660\.673\.785\.980\.8L2 \(spatial\)4\.15\.14\.14\.18\.24\.11\.02\.01\.06\.12\.03\.143\.959\.258\.269\.462\.280\.6L3 \(color\)3\.15\.22\.13\.17\.22\.11\.00\.00\.00\.01\.00\.044\.385\.656\.777\.358\.889\.7L4 \(dead ends\)1\.33\.71\.71\.03\.31\.00\.00\.00\.00\.71\.00\.348\.076\.319\.332\.759\.785\.4L5 \(twins\)1\.02\.02\.03\.04\.02\.00\.00\.00\.01\.01\.01\.042\.072\.015\.028\.057\.084\.0\(b\) Vague promptL1 \(re\-ID\)1\.01\.01\.03\.05\.12\.00\.00\.01\.03\.02\.02\.046\.541\.411\.19\.167\.745\.5L2 \(spatial\)2\.04\.16\.10\.08\.22\.02\.02\.02\.00\.02\.02\.055\.136\.76\.10\.077\.651\.0L3 \(color\)0\.06\.24\.24\.26\.24\.22\.14\.20\.00\.00\.00\.056\.229\.26\.26\.272\.933\.3L4 \(dead ends\)2\.01\.31\.30\.70\.00\.70\.02\.00\.00\.00\.00\.038\.042\.06\.02\.055\.350\.7L5 \(twins\)4\.02\.02\.02\.00\.00\.00\.00\.00\.00\.00\.00\.042\.024\.00\.02\.050\.032\.0
### 5\.2Retriever Upgrades Yield 14x the Gain of Agent Upgrades
Table 5:Upgrading agent vs\. upgrading toolOn L4 and L5, upgrading the tool yields larger gains than upgrading the agent\. Replacing Flash with Gemini 3\.1 Pro while keeping Qwen3 adds only 3\.7 percentage points on L4, compared with 52\.7 points upgrading tool from Qwen3 toVHop\-Router\(Figure[1](https://arxiv.org/html/2609.38743#S1.F1)b, Table[5](https://arxiv.org/html/2609.38743#S5.T5)\)\. On L5, the corresponding gains are 6\.0 and 56\.0 points\. Appendix Figure[12](https://arxiv.org/html/2609.38743#A5.F12)reports both tool training configurations across all levels\.
### 5\.3Can Retrieving More Images Match a Learned Multi\-Step Retriever?
Agentic baselines in Table[4](https://arxiv.org/html/2609.38743#S5.T4)return only the top\-ranked image \(k=1k=1\) per call\. Could a single\-step retriever close the gap by returning more candidates for the agent to reason over? We evaluate GE2 and Qwen3 atk∈\{1,3,5,10,50\}k\\in\\\{1,3,5,10,50\\\}, keeping Gemini 3\.5 Flash fixed\. The L4 test corpus contains450450images, sok=50k=50exposes more than one ninth per call\. Increasingkkimproves success, but neither retriever matchesVHop\-Router\(85\.4%85\.4\\%\) even atk=50k=50\(Figure[4](https://arxiv.org/html/2609.38743#S5.F4)\)\. Visual context also grows rapidly\. With Qwen3 atk=50k=50, the agent averages275275retrieved images in context per question and a cumulative API payload of1,1631\{,\}163image inputs across turns\. In contrast, Flash withVHop\-Routeraverages2\.12\.1tool calls,1212returned images, and a cumulative API payload of3333images\. It thusoutperforms while requiring23×\\penalty\\ 23\\timesfewer images in context and a35×35\\timessmaller cumulative API payloadthan thek=50k=50baselines\. All results use precise three\-hop test queries; Appendix Table[13](https://arxiv.org/html/2609.38743#A5.T13)includes the three\-hop\-trained tool and other levels\.
\(a\)Success rate\.\(b\)Images in context\.Retrieval toolkkCallsImgsSentGE2114\.515120311\.83524959\.648300106\.969344505\.42701,084Qwen3113\.013104311\.93625559\.347295107\.373392505\.52751,163VHop\-Routerchain2\.11233
Table 6:Tool call usage\.
Figure 4:Wider retrieval does not outperformVHop\-Router\.Gemini 3\.5 Flash uses precise three\-hop L4 test queries;VHop\-Routeruses precise mixed\-hop training\.Left:Success rate; dashed line: Flash withVHop\-Router\.Middle:Returned images per question\.Right:Mean calls, returned images \(*Imgs*\), and image occurrences across agent requests \(*Sent*\), per question\.
### 5\.4Generalization on One\- and Two\-Hop Queries
We evaluateVHop\-Routeron one\- and two\-hop queries across L1–L5, using the same checkpoints and held\-out worlds as in the three\-hop evaluation\. Under the default precise mixed\-hop configuration, the standaloneVHop\-Routerachieves a99\.7%success rate on one\-hop L4 queries, showing strong performance on single\-step retrieval\. Furthermore, it reaches90\.0%success on two\-hop L4 queries\. Together, these results show that one mixed\-hop policy handles both shorter query types without further fine\-tuning\. Comprehensive results for all four training configurations and evaluation levels are detailed in Appendix Tables[10](https://arxiv.org/html/2609.38743#A5.T10)and[11](https://arxiv.org/html/2609.38743#A5.T11)\.
### 5\.5Robustness to Additional Distractor Objects
In the standard evaluation \(Table[4](https://arxiv.org/html/2609.38743#S5.T4)\), each hop image contains two foreground objects\. To test robustness, we add a third object to every hop image in the L4 test set\. These objects introduce competing visual cues without changing the answers\. This modification substantially increases retrieval difficulty \(Figure[6](https://arxiv.org/html/2609.38743#S5.F6)\): success for the strongest standalone policy,VHop\-Routerwith precise mixed\-hop training, drops from76\.3%76\.3\\%to52\.2%52\.2\\%\. However, it still exceeds the strongest agentic baseline with a single\-step retriever \(27\.4%27\.4\\%\) by24\.824\.8percentage points, retaining a substantial advantage under this unseen visual shift\.
\(a\)Standard two\-object hop image\.
\(b\)The same setting with an added distractor\.
Figure 5:L4 answer accuracy \(%\)\.VHop\-Routerpolicy alone\.
Figure 6:Robustness to a foreground distractor\.Left:A standardVHopimage\.Middle:A variant with a third, unseen distractor\.Right:Retrieval success with distractor\.
### 5\.6Generalization to Realistic Queries and Settings
Table 7:Task success on realistic queries\.Answer found among returned images \(%\)\.Agentic Search MethodTraining configurationSuccess rateFlash\+\+VHop\-RouterPrecise,33\-hop36\.0Precise, mixed\-hop34\.0Flash\+\+GE2\-22\.0Flash\+\+Qwen3\-26\.0
All preceding experiments use procedurally generated images and prompts\. To test the generalization capability in real\-world scenarios, we curate another test set with realistic queries and settings including5050questions and386386images\. Annotators write the questions and specify everyday scenes\. The images are re\-rendered to remove privacy\-sensitive details while preserving theintendedcontent\. Each query requires following visual clues to locate a target item, with an image of the corresponding box serving as the final answer \(Figure[7](https://arxiv.org/html/2609.38743#S5.F7)\)\.
Without additional training, Flash withVHop\-Routerretrieves the correct answer among its returned images on36\.0%of queries with precise three\-hop training and34\.0%with precise mixed\-hop training, compared with22\.0%for Flash with GE2 and26\.0%with Qwen3 \(Table[7](https://arxiv.org/html/2609.38743#S5.T7)\)\. Full results, including standaloneVHop\-Router, appear in Appendix[F\.3](https://arxiv.org/html/2609.38743#A6.SS3)\.
Figure 7:A realistic retrieval task\.Dashed insets show corpus distractors depicting the same object types but different physical instances\. Zoom in for details\.
## 6Analysis
### 6\.1Controlling Rollout Length in RLVR
Figure 8:RLVR length\-penalty dynamics\.\(a\)Mean\|ρ\|\|\\rho\|\.\(b\)L4 test success rate\.The default RLVR objective rewards only answer correctness, without a search cost \(Section[4\.3\.3](https://arxiv.org/html/2609.38743#S4.SS3.SSS3)\)\. Training thus lengthens rollouts and increases budget exhaustion \(Figure[11](https://arxiv.org/html/2609.38743#A4.F11)\)\. To test whether a length penalty shortens rollouts, we anneal the terminal reward asrλ=r−λ\|ρ\|r\_\{\\lambda\}=r\-\\lambda\|\\rho\|,λ∈\{0,0\.02,0\.05\}\\lambda\\in\\\{0,0\.02,0\.05\\\}\.
Increasingλ\\lambdaconsistently shortens rollouts \(Figure[8](https://arxiv.org/html/2609.38743#S6.F8)\)\. On the L4 test set,λ=0\.02\\lambda=0\.02achieves accuracy comparable to the unpenalized objective, whereasλ=0\.05\\lambda=0\.05degrades performance\. A modest penalty thus removes unnecessary actions without sacrificing accuracy; a stronger penalty affects the accuracy reward\. Further results appear in Appendix[E\.4](https://arxiv.org/html/2609.38743#A5.SS4)and Table[14](https://arxiv.org/html/2609.38743#A5.T14)\.
### 6\.2Trajectory Pattern Analysis
Representative trajectories in Appendix[G](https://arxiv.org/html/2609.38743#A7)illustrate the contributions of each training stage and the agent\.SFTimproves visual matching, retrieving the correct first hop where the untrained model fails, but can still follow an incorrect spatial relation\.Online ILlearns backtracking and stopping but may stop on a wrong branch\.RLVRrefines these decisions through final\-answer rewards, recovering from the same initial mistake as Online IL, completing the correct chain, and stopping at the answer \(Figure[14](https://arxiv.org/html/2609.38743#A7.F14)\)\. The agent further complementsVHop\-Routerthroughanswer selectionandre\-querying: it selects the correct answer from the returned chain whenVHop\-Routercontinues searching until its budget runs out \(Figure[20](https://arxiv.org/html/2609.38743#A7.F20)\), or revises the query and calls the same policy again when the answer is missing, recovering the complete correct image chain \(Figure[19](https://arxiv.org/html/2609.38743#A7.F19)\)\.
## 7Conclusion and Limitations
We address visual agentic search bottlenecks by offloading multi\-hop search directly to the retrieval tool\. We introduceVHop, a benchmark and data generation framework, alongsideVHop\-Router, an end\-to\-end training pipeline for autoregressive multi\-step embedders\.VHop\-Routersignificantly boosts retrieval and agentic search success rates while drastically reducing LLM context sizes and API payloads and preserving all native LLM capabilities\. Scaling the data generation framework to more objects remains challenging because it relies on explicitly defined relationships\. The trained policy also struggles to generalize beyond the hop counts seen during training\. Future work will focus on scaling bothVHopandVHop\-Routerand incorporating real\-world datasets\.
#### Reproducibility statement
The stdlib\-only generator and verifier are deterministic from stored seeds andθ\\theta\. Appendix[B](https://arxiv.org/html/2609.38743#A2)specifies rendering checks, data splits, and prompts; Appendix[D](https://arxiv.org/html/2609.38743#A4)gives training procedures and hyperparameters\. Section[4\.1](https://arxiv.org/html/2609.38743#S4.SS1)defines baselines and metrics\.
#### AI use statement
We used AI tools to assist with manuscript drafting and editing, figure preparation, and reference checking\. Generative models were also used to render benchmark images \(Appendix[B](https://arxiv.org/html/2609.38743#A2)\) and produce captions for retrieval baselines \(Section[4\.1](https://arxiv.org/html/2609.38743#S4.SS1)\)\. The authors take responsibility for the final content, including the text, figures, references, and experimental claims\.
## References
- Ahmadianet al\.\(2024\)A\. Ahmadian, C\. Cremer, M\. Gallé, M\. Fadaee, J\. Kreutzer, O\. Pietquin, A\. Üstün, and S\. HookerBack to basics: revisiting reinforce\-style optimization for learning from human feedback in llms\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12248–12267\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px6.p1.1),[§4\.3\.3](https://arxiv.org/html/2609.38743#S4.SS3.SSS3.p1.1)\.
- Changet al\.\(2022\)Y\. Chang, M\. Narang, H\. Suzuki, G\. Cao, J\. Gao, and Y\. BiskWebqa: multihop and multimodal qa\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 16495–16504\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.5.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Gonget al\.\(2025\)A\. Gong, K\. Stankevičiūtė, C\. Wan, A\. Kabra, R\. Thesmar, J\. Lee, J\. Klenke, C\. P\. Gomes, and K\. Q\. WeinbergerPhantomwiki: on\-demand datasets for reasoning and retrieval evaluation\.arXiv preprint arXiv:2502\.20377\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.4.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Graumanet al\.\(2022\)K\. Grauman, A\. Westbury, E\. Byrne, Z\. Chavis, A\. Furnari, R\. Girdhar, J\. Hamburger, H\. Jiang, M\. Liu, X\. Liu,et al\.Ego4d: around the world in 3,000 hours of egocentric video\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 18995–19012\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.11.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Huet al\.\(2021\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLora: low\-rank adaptation of large language models\.arXiv preprint arXiv:2106\.09685\.Cited by:[§4\.2](https://arxiv.org/html/2609.38743#S4.SS2.SSS0.Px2.p1.2)\.
- Hudson and Manning \(2019\)D\. A\. Hudson and C\. D\. ManningGqa: a new dataset for real\-world visual reasoning and compositional question answering\.In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 6693–6702\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px4.p1.1)\.
- Jinet al\.\(2025\)B\. Jin, H\. Zeng, Z\. Yue, J\. Yoon, S\. Arik, D\. Wang, H\. Zamani, and J\. HanSearch\-r1: training llms to reason and leverage search engines with reinforcement learning\.arXiv preprint arXiv:2503\.09516\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.38743#S1.p1.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Johnsonet al\.\(2017\)J\. Johnson, B\. Hariharan, L\. Van Der Maaten, L\. Fei\-Fei, C\. Lawrence Zitnick, and R\. GirshickClevr: a diagnostic dataset for compositional language and elementary visual reasoning\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 2901–2910\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.9.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Kordopatis\-Ziloset al\.\(2025\)G\. Kordopatis\-Zilos, V\. Stojnić, A\. Manko, P\. Suma, N\. Ypsilantis, N\. Efthymiadis, Z\. Laskar, J\. Matas, O\. Chum, and G\. ToliasIlias: instance\-level image retrieval at scale\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 14777–14787\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.10.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Liet al\.\(2026\)M\. Li, Y\. Zhang, D\. Long, K\. Chen, S\. Song, S\. Bai, Z\. Yang, P\. Xie, A\. Yang, D\. Liu,et al\.Qwen3\-vl\-embedding and qwen3\-vl\-reranker: a unified framework for state\-of\-the\-art multimodal retrieval and ranking\.arXiv preprint arXiv:2601\.04720\.Cited by:[§D\.1](https://arxiv.org/html/2609.38743#A4.SS1.p1.1)\.
- Mensinket al\.\(2023\)T\. Mensink, J\. Uijlings, L\. Castrejon, A\. Goel, F\. Cadar, H\. Zhou, F\. Sha, A\. Araujo, and V\. FerrariEncyclopedic vqa: visual questions about detailed properties of fine\-grained categories\.In2023 IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 3090–3101\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.6.1)\.
- Oordet al\.\(2018\)A\. v\. d\. Oord, Y\. Li, and O\. VinyalsRepresentation learning with contrastive predictive coding\.arXiv preprint arXiv:1807\.03748\.Cited by:[§4\.3\.1](https://arxiv.org/html/2609.38743#S4.SS3.SSS1.p1.2)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px2.p1.1)\.
- Rosset al\.\(2011\)S\. Ross, G\. Gordon, and D\. BagnellA reduction of imitation learning and structured prediction to no\-regret online learning\.InProceedings of the fourteenth international conference on artificial intelligence and statistics,pp\. 627–635\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px6.p1.1),[§4\.3\.2](https://arxiv.org/html/2609.38743#S4.SS3.SSS2.p1.1)\.
- Schulmanet al\.\(2017\)J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. KlimovProximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px6.p1.1),[§D\.4](https://arxiv.org/html/2609.38743#A4.SS4.SSS0.Px3.p1.1),[§4\.3\.3](https://arxiv.org/html/2609.38743#S4.SS3.SSS3.p1.1)\.
- Shanbhogueet al\.\(2026\)M\. Shanbhogue, Z\. Li, S\. Zhang, G\. H\. Ábrego, S\. Huang, A\. Jain, D\. Salz, S\. Goenka, C\. Hegde, J\. Ma,et al\.Gemini embedding 2: a native multimodal embedding model from gemini\.arXiv preprint arXiv:2605\.27295\.Cited by:[§4\.1](https://arxiv.org/html/2609.38743#S4.SS1.SSS0.Px3.p1.1)\.
- Sorokinet al\.\(2026\)A\. Sorokin, N\. Buzun, A\. Anokhin, E\. Vedernikov, P\. Anokhin, M\. Burtsev, and E\. BurnaevQ\-rag: long context multi\-step retrieval via value\-based embedder training\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 32481–32506\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Trivediet al\.\(2022\)H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. SabharwalMusique: multihop questions via single\-hop question composition\.Transactions of the Association for Computational Linguistics10,pp\. 539–554\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.3.1)\.
- Trivediet al\.\(2023\)H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. SabharwalInterleaving retrieval with chain\-of\-thought reasoning for knowledge\-intensive multi\-step questions\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 10014–10037\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38743#S1.p1.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Wanget al\.\(2025\)F\. Wang, X\. Fu, J\. Y\. Huang, Z\. Li, Q\. Liu, X\. Liu, M\. D\. Ma, N\. Xu, W\. Zhou, K\. Zhang,et al\.Muirbench: a comprehensive benchmark for robust multi\-image understanding\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 62624–62650\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.8.1)\.
- Welleret al\.\(2026\)O\. Weller, M\. Boratko, I\. Naim, and J\. LeeOn the theoretical limitations of embedding\-based retrieval\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 11202–11216\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px4.p1.1)\.
- Wuet al\.\(2025\)T\. Wu, G\. Biamby, J\. Quenum, R\. Gupta, J\. E\. Gonzalez, D\. Chan,et al\.Visual haystacks: a vision\-centric needle\-in\-a\-haystack benchmark\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 64592–64615\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.7.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Xionget al\.\(2020\)W\. Xiong, X\. L\. Li, S\. Iyer, J\. Du, P\. Lewis, W\. Y\. Wang, Y\. Mehdad, W\. Yih, S\. Riedel, D\. Kiela,et al\.Answering complex open\-domain questions with multi\-hop dense retrieval\.arXiv preprint arXiv:2009\.12756\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.38743#S2.p1.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 2369–2380\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.38743#S1.T1.6.1.1.1.2.1)\.
- Yaoet al\.\(2022\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReact: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.38743#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.38743#S4.SS1.SSS0.Px3.p2.1)\.
- Yuet al\.\(2026\)Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.Dapo: an open\-source llm reinforcement learning system at scale\.Advances in Neural Information Processing Systems38,pp\. 113222–113244\.Cited by:[Appendix A](https://arxiv.org/html/2609.38743#A1.SS0.SSS0.Px6.p1.1),[§D\.4](https://arxiv.org/html/2609.38743#A4.SS4.SSS0.Px3.p1.1),[§4\.3\.3](https://arxiv.org/html/2609.38743#S4.SS3.SSS3.p1.1)\.
## Appendix AAdditional Related Work
##### Multi\-hop retrieval and question answering\.
HotpotQA\([Yang et al\., 2018](https://arxiv.org/html/2609.38743#bib.bib2)\)requires combining evidence from multiple Wikipedia articles, while MuSiQue\([Trivedi et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib3)\)composes single\-hop questions into connected reasoning chains and filters out shortcuts\. Multi\-hop Dense Retrieval\([Xiong et al\., 2020](https://arxiv.org/html/2609.38743#bib.bib10)\)learns to retrieve a passage conditioned on earlier evidence\. IRCoT\([Trivedi et al\., 2023](https://arxiv.org/html/2609.38743#bib.bib4)\)interleaves retrieval with reasoning so that each informs the next step\. VHOP studies a visual form of this dependency: a retrieved image supplies the object or mark needed to find the next image, and the final answer is itself an image\.
##### Visual retrieval and memory\.
CLIP\([Radford et al\., 2021](https://arxiv.org/html/2609.38743#bib.bib5)\)learns a shared image–text embedding space that supports retrieval by similarity\. Instance\-level benchmarks such as ILIAS\([Kordopatis\-Zilos et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib6)\)test whether a system can find the same particular object among many distractors\. Ego4D’s visual\-query task\([Grauman et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib11)\)uses an object image to locate its last appearance in an egocentric video\. VHOP extends such visual search to a chain of linked objects: the starting image may not show the final target, so directly matching that image to the answer is insufficient\. The solver must use intermediate images to determine what to retrieve next\.
##### Multimodal and multi\-image reasoning\.
WebQA\([Chang et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib12)\)combines evidence retrieval from image or text sources with natural\-language answer generation\. Encyclopedic\-VQA\([Mensink et al\., 2023](https://arxiv.org/html/2609.38743#bib.bib13)\)asks detailed questions about visually identified entities and provides a Wikipedia knowledge base\. Visual Haystacks\([Wu et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib21)\)tests question answering over large image collections, including questions that require evidence from multiple images\. MuirBench\([Wang et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib22)\)covers a broad range of multi\-image tasks, including visual retrieval\. VHOP focuses on a specific question: can a system repeatedly recover a new visual clue, follow it through a corpus, and recover when a locally plausible match does not lead to an answer? Its generator makes these dependencies and alternative branches explicit for controlled evaluation\.
##### Controlled evaluation\.
CLEVR\([Johnson et al\., 2017](https://arxiv.org/html/2609.38743#bib.bib1)\)uses generated scenes and executable question programs to diagnose compositional visual reasoning\. GQA\([Hudson and Manning, 2019](https://arxiv.org/html/2609.38743#bib.bib14)\)generates compositional questions from scene graphs of real images\. PhantomWiki\([Gong et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib16)\)generates consistent document collections and questions with adjustable reasoning difficulty and corpus size\. VHOP follows this approach of making task structure controllable, while placing the evidence across separate images\. Its verifier checks satisfying trajectories over the pooled corpus before rendering\. The guarantees are conditional on faithful rendering; the additional retrieval bounds also require the solver restrictions and sampling assumptions stated in Appendix[C](https://arxiv.org/html/2609.38743#A3)\. Separately,[Weller et al\. \(2026\)](https://arxiv.org/html/2609.38743#bib.bib15)establish limitations imposed by embedding dimension on the relevance sets a single\-vector retriever can represent\. Our analysis concerns missing intermediate evidence and locally ambiguous branches, rather than embedding dimension\.
##### Agents and learned search policies\.
ReAct\([Yao et al\., 2022](https://arxiv.org/html/2609.38743#bib.bib9)\)interleaves language reasoning with actions and external observations\. Search\-R1\([Jin et al\., 2025](https://arxiv.org/html/2609.38743#bib.bib7)\)trains language models to issue search queries using outcome\-based reinforcement learning\. Q\-RAG\([Sorokin et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib8)\)trains embeddings for multi\-step text retrieval through a value\-based objective\. Our method also adapts retrieval representations\. It operates on image sequences and learns selection, backtracking, and stopping from imitation and final\-answer rewards\. The trained policy can also serve as a tool for an agent that revises queries or selects an answer from returned images\.
##### Imitation learning and outcome supervision\.
Collecting expert labels on states visited by the learner addresses the distribution shift that motivates DAgger\([Ross et al\., 2011](https://arxiv.org/html/2609.38743#bib.bib17)\)\. Our Online IL stage uses this idea but updates only on the current iteration’s states, without accumulating a dataset across iterations; we therefore do not call the algorithm DAgger or invoke its guarantees\. The subsequent RLVR stage uses a terminal answer check, a leave\-one\-out baseline\([Ahmadian et al\., 2024](https://arxiv.org/html/2609.38743#bib.bib19)\), and a clipped policy objective\([Schulman et al\., 2017](https://arxiv.org/html/2609.38743#bib.bib18);[Yu et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib20)\)\. We adapt these components to a visual retrieval policy with explicit selection, backtracking, and stopping decisions\.
## Appendix BBenchmark Generation and Verification
VHopseparates task construction from image rendering\. The generator first specifies objects, their links across images, and the constraints in each question\. A symbolic verifier checks the resulting search problem over the entire corpus\. Rendering then turns this specification into images\. This separation provides exact task labels while allowing the visual appearance and search difficulty to vary\.
### B\.1Difficulty Levels and World Construction
Table 8:Cumulative difficulty levels\.Each level retains the requirements of the preceding levels\.Figure 9:Visual examples of the difficulty levels\.L1 shows a complete retrieval chain\. L2 and L3 add spatial and color constraints; L4 introduces dead ends; and L5 requires distinguishing look\-alike objects\. L6 and L7 extend the matching rule to logos and printed text\. The examples come from separate questions\. Boxes, arrows, and enlarged details are reader annotations\.A*world*is one image corpus and its associated questions\. A*chain family*contains a gold chain and its related distractors\. The generator controls hop count, matching rules, relations, colors, family count, and the numbers of distractors and forks\. It assigns fresh matching keys to each family and records every image containing each key\. The final portrait is included in the searchable corpus\. Query images are fresh views of the starting objects\.
Distractors target different errors\. Relation and color decoys violate a stated constraint\. A dead end satisfies the local match but lacks a valid continuation\. An appearance twin resembles the correct object while representing a different instance\. Extra portraits share query\-specified properties with the answer, preventing those properties alone from identifying it\. Images and filenames are shuffled without exposing their roles\. The corpus is available from the start; the task requires discovering the links between images, rather than revealing hidden corpus entries\.
Algorithm 2Generating and verifying aVHopworld1:Difficulty level, hop count, family count, and object/mark pools
2:foreach chain familydo
3:sample matching rules, spatial relations, and color constraints
4:assign fresh keys and construct the gold scenes and final portrait
5:add the level’s decoys, dead ends, twins, and matching distractor portraits
6:construct a precise query without naming the intermediate key values
7:endfor
8:pool the images and shuffle their identifiers
9:foreach querydo
10:enumerate all satisfying trajectories over the pooled corpus
11:check the unique precise path and the intended distractor structure
12:enumerate acceptable endpoints for the vague variant, if supported
13:reject and regenerate any family that fails verification
14:endfor
15:render the verified world and check its images against the specification
16:returnimage corpus, query images, instructions, and verified labels
### B\.2Prompt Variants and Data Separation
Precise instructions state the sequence of matching rules, relations, and filters\. They refer to an intermediate object through the preceding image, without giving its identity, logo value, or printed string in advance\. Vague instructions omit a relation and may admit more than one answer\. The verifier enumerates the acceptable endpoints after that omission; evaluation accepts any member of this set\. Vague variants are reported for L1–L5\. One\- and two\-hop questions are verified suffixes of the three\-hop chains and use the same image corpora\.
Training and evaluation use separate worlds and separate pools of object appearances and marks\. SFT and Online IL use worlds with full trajectory labels\. RLVR uses a separate practice pool and obtains its learning signal from the final\-answer verifier\. Development worlds support model selection; the reported L4 test results use three held\-out worlds\. World identifiers such asw0w\_\{0\}are local to their split, so a trainingw0w\_\{0\}and a developmentw0w\_\{0\}do not denote the same corpus\.
### B\.3Rendering Conditions
Each rendering prompt specifies the objects, their colors and marks, and their spatial arrangement\. Reference images maintain an object’s identity or a mark across views\. A portrait prompt requests the relevant object alone\. The question text is generated separately and does not reveal the intermediate matching values\.
The symbolic labels describe the intended scene\. For these labels to remain valid in the images, rendering must preserve object identity across views, distinguish different instances where required, reproduce colors and marks, and show the stated relations unambiguously\. An image audit checks these properties against the manifest; images with visible violations require regeneration\. These conditions are also the assumptions needed to transfer the symbolic guarantees in Appendix[C](https://arxiv.org/html/2609.38743#A3)to pixels\. The verifier does not, by itself, certify rendering fidelity or rule out every possible captioning shortcut\.
## Appendix CSymbolic Guarantees
This section states what the generator guarantees about its symbolic worlds, and the assumptions needed for the retrieval bounds\. A unique symbolic answer does not by itself guarantee that a rendered image is unambiguous\. Applying these statements to pixels also requires the rendering conditions in Appendix[B](https://arxiv.org/html/2609.38743#A2)\.
##### Definitions\.
An object’s*link key*is its identity, logo, or printed string, depending on the matching rule in the instruction\. A scene contains an*anchor*that matches the preceding object and a*witness*that supplies the next key\. The witness must satisfy the stated spatial relation and any color constraint\. A portrait contains the final object alone\. LetHHbe the total number of retrieval hops, including the final portrait; the three\-hop task therefore has two scene hops and one portrait hop\.
For a queryQQand pooled corpusDD, letSat\(Q,D\)\\mathrm\{Sat\}\(Q,D\)be the set of image sequences that satisfy every matching rule, relation, and filter, and end in an eligible portrait\. The acceptable answer set𝒜\(Q,D\)\\mathcal\{A\}\(Q,D\)contains the last image of each such sequence\. Precise queries are verified to have one satisfying sequence\. For vague queries, the verifier enumerates the sequences allowed by the omitted relation and accepts all their endpoints\.
###### Theorem 1\(Whole\-corpus uniqueness\)\.
Suppose each chain family has fresh link keys, every matched scene has a unique anchor and witness, each non\-forked gold prefix has one valid continuation, and each alternative branch at a fork has no satisfying completion\. If the terminal key has exactly one eligible portrait, then a precise query has exactly one satisfying trajectory in the pooled corpus\.
###### Proof\.
The constructed gold trajectory satisfies the query\. Consider a different satisfying trajectory and its first departure from the gold trajectory\. Fresh keys prevent a continuation through another family\. At a non\-forked prefix, continuation uniqueness rules out the departure\. At a fork, the alternative has no satisfying completion\. A change in the final image is excluded by terminal\-portrait uniqueness\. Thus no different satisfying trajectory exists\. The verifier checks these conditions over the pooled corpus, including decoys and appearance twins\. ∎
###### Theorem 2\(Reusing appearance templates\)\.
Reusing an object’s appearance template across families preserves symbolic uniqueness if link keys remain fresh and the conditions of Theorem[1](https://arxiv.org/html/2609.38743#Thmtheorem1)remain satisfied\.
###### Proof\.
Symbolic matching depends on link keys and scene structure, not on the appearance template used to render an object\. Reusing a template therefore does not add an object to a key’s matching set\. For an identity link, different instances must still be distinguishable in the rendered images\. ∎
##### Assumptions for direct retrieval\.
A pairwise scorer assignsf\(Q,z\)f\(Q,z\)to each candidate imagezzusing only the query and that image\. It has no retrieved intermediate evidence, corpus\-level features, or metadata that reveals an image’s role\. Considernncandidate portraits that agree on all answer properties supplied by the query, such as color and matching rule\. The required symmetry assumption is that, conditioned on the evidence available to this scorer, each of these portraits is equally likely to be the answer\. This requires role\-independent sampling of keys and appearances and symmetric handling of collisions\. Omitting a key from the query alone is insufficient to establish this symmetry\.
###### Theorem 3\(Direct pairwise answer bound\)\.
Under the conditional symmetry above, ifn≥m\+1n\\geq m\+1portraits share the answer’s query\-specified properties and at least one scene hop precedes the answer, any pairwise scorer has top\-1 answer accuracy at most1/\(m\+1\)1/\(m\+1\), averaged over the sampled worlds\.
###### Proof\.
Conditioned on the scorer’s available evidence, the answer index is uniform over thennportraits\. Any selected index has probability1/n1/nof being correct\. Selecting an image outside this set cannot increase success\. Averaging over the evidence gives the bound\. A solver that retrieves an intermediate image and reads its next key has additional evidence, so this bound does not apply to it\. ∎
###### Proposition 1\(Completing a chain with a tight retrieval budget\)\.
Consider a policy that follows the query’s local transitions and must retrieve a completeHH\-hop chain using at mostHHimage retrievals\. At each forktt, suppose there aredtd\_\{t\}incorrect branches and, conditioned on a correct prefix and the policy’s available evidence, the correct branch is uniform among the1\+dt1\+d\_\{t\}choices\. If an incorrect branch cannot complete the query, the probability of completing the chain is at most∏t:dt\>0\(1\+dt\)−1\\prod\_\{t:\\,d\_\{t\}\>0\}\(1\+d\_\{t\}\)^\{\-1\}\.
###### Proof\.
The budget leaves no retrievals for testing a wrong branch and returning to the correct one\. The conditional probability of choosing correctly at forkttis at most\(1\+dt\)−1\(1\+d\_\{t\}\)^\{\-1\}\. Applying the chain rule over successive forks gives the product\. Unconditional independence between forks is not required\. ∎
This proposition concerns complete\-chain retrieval under the stated local information restriction\. It does not bound finding the answer anywhere in a returned set, guessing the final portrait directly, or using a larger budget for recovery\. Its budget counts image retrievals; the experimental decision budget also countsBacktrackandStop\. The benefit of additional search is measured empirically\.
## Appendix DTraining and Evaluation Details
### D\.1Data, Models, and Evaluation
All four training configurations in Table[4](https://arxiv.org/html/2609.38743#S5.T4)use L4 data and Qwen3\-VL\-Embedding\-2B\([Li et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib23)\)\. They vary the prompt type \(precise or vague\) and training hop counts \(three hops only or a mixture of one, two, and three\)\. The fully labeled seed world contains978978questions and4,4924\{,\}492corpus images\. SFT and Online IL use its gold trajectories\. RLVR uses separate practice worlds and a final\-answer verifier; it does not train on their intermediate gold actions\. The detailed optimization settings below describe the precise three\-hop reference run, which producesrloo10w10\_200\.
For the three\-hop results, the precise\-prompt sample sizes at L1–L7 are\(99,98,97,300,100,94,95\)\(99,98,97,300,100,94,95\); the vague\-prompt sizes at L1–L5 are\(99,49,48,150,50\)\(99,49,48,150,50\)\. L4 pools three held\-out test worlds, while the other levels use development worldw0w\_\{0\}\. Development scores are used for model selection and should not be read as independently held\-out estimates\. Embedding and caption baselines measure whether an acceptable answer is among their five returned images\. For a standalone policy, correctness requiresStopwith an acceptable image at the top of the active stack\. Agent accuracy checks the returned answer, including the evaluator’s fallback if there is no valid final declaration\. Finding an answer among retrieved images is a separate metric, defined for the human\-authored evaluation in Appendix[F\.3](https://arxiv.org/html/2609.38743#A6.SS3)\.
The active stackσt\\sigma\_\{t\}contains the query image and the retrieved images that have not been removed by backtracking\. Its encoding isst=Es\(x,σt\)s\_\{t\}=E\_\{s\}\(x,\\sigma\_\{t\}\)\. The rolloutρ\\rhorecords every decision, including selections later undone\. Candidate masks and stored parent scores are maintained by the search procedure\. Thus a popped image leaves subsequent state\-encoder inputs but its earlier decision and training record remain in the rollout\. This distinction applies in both Online IL and RLVR\.
### D\.2Supervised Fine\-Tuning
Figure 10:State\-encoder warm\-up stabilizes SFT\.The state and action encoders use separate LoRA adapters over a frozen backbone, with rank3232and scaling parameter3232\. The adapters modify the query, key, and value projections and the MLP projections\. Embeddings use last\-token pooling and are normalized to unit length, with dimension20482048\. The reference SFT recipe uses learning rate2×10−52\\times 10^\{\-5\}, batch size88per GPU, and a linear schedule over four epochs\.
The first stage trains only the state encoder on gold prefixes with full\-corpus InfoNCE at temperature0\.050\.05\. The next stage introduces hop\-level hard negatives, including designed dead ends and constraint violations, and also updates the action encoder\. Hard\-negative images are re\-encoded with gradients and their scores replace the corresponding entries in the cached corpus logits\. These re\-encoded images provide the action encoder’s gradient; the rest of the cached index is treated as constant during an update\. Later stages freeze the action encoder and its corpus index\.
### D\.3Online Imitation Learning
##### Rollout collection\.
Online IL runs for120120iterations with a1414\-decision budget per rollout\. The oracle selects the next gold image from a valid prefix, backtracks from an incorrect prefix, and stops at the completed chain\. At iterationii, the expert\-control and forced\-excursion probabilities are
βi=max\(0,1−i60\),pi=0\.6max\(0,1−i80\)\.\\beta\_\{i\}=\\max\\\!\\left\(0,1\-\\frac\{i\}\{60\}\\right\),\\qquad p\_\{i\}=0\.6\\max\\\!\\left\(0,1\-\\frac\{i\}\{80\}\\right\)\.\(8\)At a valid prefix with an available dead end, the rollout enters a dead end with probabilitypip\_\{i\}\. Otherwise, it follows the oracle with probabilityβi\\beta\_\{i\}and the current policy with probability1−βi1\-\\beta\_\{i\}\. Every visited state receives an oracle label, regardless of how its action was chosen\. Each iteration trains only on the states collected in that iteration; states from earlier iterations are not retained for training\. Each record stores the stack, candidate mask, parent maximum score, and oracle action\. During the update, we re\-encode the recorded stack and recompute the current retrieval scores andctc\_\{t\}, retaining the recorded mask and parent score\. A later backtrack does not remove earlier records\.
##### Online IL losses and gradients\.
Selection uses the top\-88eligible candidates, with the expert’s next gold image appended if absent, and a softmax temperature of0\.050\.05\. The gate losses in Eq\.[7](https://arxiv.org/html/2609.38743#S4.E7)are
Bt\(y\)\\displaystyle B\_\{t\}\(y\)=−w\+ylogptbt−\(1−y\)log\(1−ptbt\),\\displaystyle=\-w\_\{\+\}y\\log p\_\{t\}^\{\\mathrm\{bt\}\}\-\(1\-y\)\\log\(1\-p\_\{t\}^\{\\mathrm\{bt\}\}\),St\(y\)\\displaystyle S\_\{t\}\(y\)=−ylogptstop−\(1−y\)log\(1−ptstop\),\\displaystyle=\-y\\log p\_\{t\}^\{\\mathrm\{stop\}\}\-\(1\-y\)\\log\(1\-p\_\{t\}^\{\\mathrm\{stop\}\}\),\(9\)wherew\+=4w\_\{\+\}=4compensates for the relative scarcity of backtrack targets\. The gate probabilities use Eq\. \([3](https://arxiv.org/html/2609.38743#S4.E3)\) withτg=1\\tau\_\{g\}=1\. We setBt\(y\)=0B\_\{t\}\(y\)=0at the root andSt\(y\)=0S\_\{t\}\(y\)=0wherever stopping is disabled\. For the procedural evaluations, three\-hop\-only checkpoints permit stopping once the stack depth reaches the question’s hop count; mixed\-hop checkpoints permit it after any retrieval\. This depth floor is an evaluation constraint and should be kept distinct from the learned stop decision\. The per\-state losses are averaged over the collected dataset:
ℒOnlineIL=1\|𝒟i\|∑t=1\|𝒟i\|ℓt\.\\mathcal\{L\}\_\{\\mathrm\{Online\\,IL\}\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{i\}\|\}\\sum\_\{t=1\}^\{\|\\mathcal\{D\}\_\{i\}\|\}\\ell\_\{t\}\.\(10\)The action encoder and its corpus index remain frozen\. Selection gradients reach the state encoder through the retrieval scores, while gate gradients reach both the corresponding MLP and the state encoder throughsts\_\{t\}\. The state embedding is not detached on backtrack or stop examples\. The four retrieval\-context features are treated as constants: the current maximum score, the top\-1/top\-2 gap, the parent maximum, and their difference\.
### D\.4Outcome\-Based RL Details
\(a\)Reward and test accuracy
\(b\)Length ofρ\\rhoand backtracks
\(c\)Truncation rate
Figure 11:Training dynamics of RLVR\.\(a\) Training reward and test accuracy\. \(b\) Rollout length and number of backtracks\. \(c\) The fraction of rollouts that exceed budget\.##### Rewards and advantages\.
For each input\(x,q0\)\(x,q\_\{0\}\), we sampleGGrolloutsρg\\rho\_\{g\}from the behavior policyμ\\mu, a snapshot of the current policy at the start of the iteration\. Let𝒜\\mathcal\{A\}be the input’s certified answer set: a singleton for a precise question, or the set of valid answers for a vague question\. WithTgT\_\{g\}denoting the final decision index, the terminal verifier returns
rg=\[ag,Tg=Stop∧top\(σg,Tg\)∈𝒜\]\.r\_\{g\}=\\mathbf\{1\}\\\!\\left\[a\_\{g,T\_\{g\}\}=\\textsc\{Stop\}\\ \\land\\ \\operatorname\{top\}\(\\sigma\_\{g,T\_\{g\}\}\)\\in\\mathcal\{A\}\\right\]\.\(11\)Each decision inρg\\rho\_\{g\}receives the same leave\-one\-out advantage
Ag=rg−1G−1∑g′≠grg′,A\_\{g\}=r\_\{g\}\-\\frac\{1\}\{G\-1\}\\sum\_\{g^\{\\prime\}\\neq g\}r\_\{g^\{\\prime\}\},\(12\)where the sum is over the other rollouts for the same input\. Groups with identical rewards have zero advantage and contribute no policy\-gradient signal\.
Algorithm 3RLVR training1:Online IL policy
πθ\\pi\_\{\\theta\}; terminal answer verifier; frozen action encoder
2:freeze a reference copy
π0\\pi\_\{0\}of the initial policy
3:foreach iterationdo
4:snapshot the current policy as the behavior policy
μ\\mu
5:foreach sampled querydo
6:generate
GGcomplete rollouts
ρg\\rho\_\{g\}using
μ\\mu
7:for every prefix
ρg,<t\\rho\_\{g,<t\}, record its active stack
σg,t\\sigma\_\{g,t\}, masks, context, and behavior probabilities
8:compute terminal rewards
rgr\_\{g\}and leave\-one\-out advantages
AgA\_\{g\}
9:endfor
10:foreach inner update epochdo
11:re\-encode retained stacks as
sg,t=Es\(x,σg,t\)s\_\{g,t\}=E\_\{s\}\(x,\\sigma\_\{g,t\}\); recompute candidate scores and gate probabilities
12:update
EsE\_\{s\}and the gates with clipped branch terms and KL penalties
13:endfor
14:endfor
Table 9:Basic RL policy\-gradient form\.
##### Branch\-local credit assignment\.
For decisionttin rolloutρg\\rho\_\{g\}, letpg,tp\_\{g,t\}denote the sampled\-branch probability factor in Table[9](https://arxiv.org/html/2609.38743#A4.T9), andp¯g,t\\bar\{p\}\_\{g,t\}its stored value underμ\\mu\. The SELECT and gate probabilities follow Eqs\. \([2](https://arxiv.org/html/2609.38743#S4.E2)\) and \([3](https://arxiv.org/html/2609.38743#S4.E3)\)\. We omit the complementary gate factors in Eq\. \([4](https://arxiv.org/html/2609.38743#S4.E4)\):\(1−ptstop\)\(1−ptbt\)\(1\-p\_\{t\}^\{\\mathrm\{stop\}\}\)\(1\-p\_\{t\}^\{\\mathrm\{bt\}\}\)for SELECT and\(1−ptstop\)\(1\-p\_\{t\}^\{\\mathrm\{stop\}\}\)for BACKTRACK\. This generally yields a biased surrogate for the full policy gradient obtained from Eq\. \([5](https://arxiv.org/html/2609.38743#S4.E5)\)\. This choice prevents a SELECT policy\-loss term from directly changing the gates through their rejection probabilities\.
The branch terms share one objective and optimizer\. SELECT gradients reachEsE\_\{s\}; BACKTRACK and STOP gradients reachEsE\_\{s\}and the respective gate MLP\. The policy loss for SELECT has no direct gate\-MLP gradient\. Shared\-encoder updates and KL penalties can still change other branch probabilities;EaE\_\{a\}remains fixed\.
##### DAPO\-style clipping and negative\-advantage weighting\.
For batch reuse, we use a PPO\-style ratio surrogate\([Schulman et al\., 2017](https://arxiv.org/html/2609.38743#bib.bib18)\)with DAPO’s asymmetric clipping bounds\([Yu et al\., 2026](https://arxiv.org/html/2609.38743#bib.bib20)\):
ωg,t\(p\)\\displaystyle\\omega\_\{g,t\}\(p\)=pp¯g,t,\\displaystyle=\\frac\{p\}\{\\bar\{p\}\_\{g,t\}\},\(13\)𝒥g,t\(p\)\\displaystyle\\mathcal\{J\}\_\{g,t\}\(p\)=−w\(Ag\)min\{ωg,t\(p\)Ag,clip\(ωg,t\(p\),1−ϵlo,1\+ϵhi\)Ag\},\\displaystyle=\-w\(A\_\{g\}\)\\min\\\!\\left\\\{\\omega\_\{g,t\}\(p\)A\_\{g\},\\,\\operatorname\{clip\}\\\!\\left\(\\omega\_\{g,t\}\(p\),1\-\\epsilon\_\{\\mathrm\{lo\}\},1\+\\epsilon\_\{\\mathrm\{hi\}\}\\right\)A\_\{g\}\\right\\\},\(14\)with
w\(A\)=\{1,A≥0,λneg,A<0\.w\(A\)=\\begin\{cases\}1,&A\\geq 0,\\\\ \\lambda\_\{\\mathrm\{neg\}\},&A<0\.\\end\{cases\}\(15\)Settingϵhi\>ϵlo\\epsilon\_\{\\mathrm\{hi\}\}\>\\epsilon\_\{\\mathrm\{lo\}\}permits larger probability increases for positive\-advantage decisions before clipping\. The additional weight0<λneg<10<\\lambda\_\{\\mathrm\{neg\}\}<1reduces the magnitude of negative\-advantage terms without reversing their sign\. Atp=p¯g,tp=\\bar\{p\}\_\{g,t\}, withw=1w=1and fixed probability support, the surrogate’s gradient matches the−Aglogp\-A\_\{g\}\\log pform in Table[9](https://arxiv.org/html/2609.38743#A4.T9)\.
The implemented policy objective sums the clipped terms:
ℒPG=1Nroll∑ρg∈ℬ∑t=0Tg𝒥g,t\(pg,t\)\.\\mathcal\{L\}\_\{\\mathrm\{PG\}\}=\\frac\{1\}\{N\_\{\\mathrm\{roll\}\}\}\\sum\_\{\\rho\_\{g\}\\in\\mathcal\{B\}\}\\sum\_\{t=0\}^\{T\_\{g\}\}\\mathcal\{J\}\_\{g,t\}\(p\_\{g,t\}\)\.\(16\)Here,ℬ\\mathcal\{B\}contains the nonzero\-advantage rollouts, whileNrollN\_\{\\mathrm\{roll\}\}counts all sampled rollouts, including zero\-advantage groups\. States excluded by the training filters below contribute zero policy loss\. During batch reuse, a SELECT term is evaluated only while the sampled image remains in the current top\-KKset𝒞t\\mathcal\{C\}\_\{t\}\.
##### RL regularization\.
The full objective adds behavior and reference KL penalties:
ℒRL=ℒPG\+βbeh𝒦beh\+βref\(i\)𝒦ref\.\\mathcal\{L\}\_\{\\mathrm\{RL\}\}=\\mathcal\{L\}\_\{\\mathrm\{PG\}\}\+\\beta\_\{\\mathrm\{beh\}\}\\mathcal\{K\}\_\{\\mathrm\{beh\}\}\+\\beta\_\{\\mathrm\{ref\}\}\(i\)\\mathcal\{K\}\_\{\\mathrm\{ref\}\}\.\(17\)The behavior policyμ\\muis the policy that generated the current rollout batch\. We record its selection probabilities before updating the model\. The behavior penalty sumsKL\(πθsel∥μsel\)\\mathrm\{KL\}\(\\pi\_\{\\theta\}^\{\\mathrm\{sel\}\}\\,\\\|\\,\\mu^\{\\mathrm\{sel\}\}\)over retained SELECT states and divides byNrollN\_\{\\mathrm\{roll\}\}\. Both distributions use the candidate set stored during rollout\. The stored probabilities remain fixed through all updates on this batch; they are refreshed only when the updated policy collects the next batch\. Thus, behavior KL limits changes from the policy at the start of the current iteration, even when the batch is reused for several gradient updates\. The reference policyπ0\\pi\_\{0\}is a frozen copy of the initial Online IL policy\. On a subsample of retained states,𝒦ref\\mathcal\{K\}\_\{\\mathrm\{ref\}\}adds the SELECT KL on the stored candidate set when selection was taken, and Bernoulli KLs for each permitted gate\. All KLs use the current policy as their first argument\. The reference sum is normalized by the number of reference states, floored at1616\. Its weight decreases during training to stabilize early updates while allowing later improvement without gold\-trajectory supervision\.
##### RL configuration\.
The precise three\-hop reference run uses eight data\-parallel GPUs\. At each iteration, each rank samples44practice questions andG=8G=8stochastic rollouts per question, each limited to1616decisions\. The selection temperature isτ=0\.05\\tau=0\.05and the gate temperature isτg=0\.5\\tau\_\{g\}=0\.5\. Each group stays on one rank, so leave\-one\-out baselines require no communication\. The state\-encoder LoRA learning rate is5×10−65\\times 10^\{\-6\}and the gate\-MLP rate is3×10−43\\times 10^\{\-4\}\. The action encoder and corpus index remain frozen throughout\. We reuse each batch for two inner epochs with DAPO\-style asymmetric clipping \(ϵlo=0\.2\\epsilon\_\{\\mathrm\{lo\}\}=0\.2,ϵhi=0\.3\\epsilon\_\{\\mathrm\{hi\}\}=0\.3\) andλneg=0\.5\\lambda\_\{\\mathrm\{neg\}\}=0\.5\.
Gradient\-bearing states are restricted to\|σt\|≤9\|\\sigma\_\{t\}\|\\leq 9, including the query image \(equivalently, chain depthdt≤8d\_\{t\}\\leq 8\), to bound activation memory\. If more than128128eligible states remain per rank, we sample128128uniformly and scale the policy loss by the original count divided by128128\. The behavior\-KL weight isβbeh=0\.05\\beta\_\{\\mathrm\{beh\}\}=0\.05\. Reference KL is evaluated on at most4848retained states per rank, withβref\(i\)=0\.3\\beta\_\{\\mathrm\{ref\}\}\(i\)=0\.3for iterations11–7575and0\.10\.1thereafter\. The reference run lasts200200iterations \(≈14\\approx 14hours\) and yields checkpointrloo10w10\_200\.
##### Length\-penalty variant\.
This variant keeps the terminal verifier unchanged but substitutesr~g=rg−λndecisions\(ρg\)\\tilde\{r\}\_\{g\}=r\_\{g\}\-\\lambda n\_\{\\mathrm\{decisions\}\}\(\\rho\_\{g\}\)forrgr\_\{g\}in Eq\. \([12](https://arxiv.org/html/2609.38743#A4.E12)\)\. The decision count includes SELECT, BACKTRACK, and STOP; strict reward remains the evaluation and checkpoint\-selection criterion\. The reference run usesλ=0\\lambda=0\. The recordedλ=0\.02\\lambda=0\.02run uses one\-state gradient microbatches and raises the gradient\-bearing stack limit to1212images to accommodate the additional failed\-rollout states with nonzero advantages\. This implementation difference should be considered alongside the reward change\. Appendix[E\.4](https://arxiv.org/html/2609.38743#A5.SS4)examines the effect of charging for each decision\.
##### Roles of the optimization components\.
The leave\-one\-out baseline compares outcomes for the same query without normalizing by a small within\-group reward standard deviation\. Asymmetric clipping limits changes during batch reuse, while the negative\-advantage weight reduces the contribution of failed rollouts\. Behavior KL constrains updates relative to the policy that collected the batch\. Reference KL limits drift from the initial Online IL policy\.
## Appendix EAdditional Results and Ablations
### E\.1One\- and Two\-Hop Queries
Tables[10](https://arxiv.org/html/2609.38743#A5.T10)and[11](https://arxiv.org/html/2609.38743#A5.T11)report the complete results on verified suffixes of the three\-hop queries\. The corpus and checkpoints are unchanged\. The three\-hop\-only policies retain the question\-specific minimum stop depth; mixed\-hop policies can stop after any retrieval\. As in the main table, embedding baselines receive credit if the answer is among five results, whereas the standalone policy must stop on one correct answer\. This distinction is especially relevant on one\-hop queries, where a top\-5 list contains five possible answers\.
Table 10:One\-hop success rate \(%\)\.Same corpora, checkpoints, and scoring conventions as Table[4](https://arxiv.org/html/2609.38743#S5.T4)\. L4 uses test worldsw1w\_\{1\}–w3w\_\{3\}; other levels use development worldw0w\_\{0\}\. Caption→\\toQwen3 retains the query image as pixels\. In block \(b\), the agent uses the learned per\-step retriever; the other trained\-agent blocks use the chain tool\. Vague prompts are evaluated on L1–L5\. Bold marks the best result within each training block on each side of the dashed rule\.Retrieval methodsAgentic searchEmbedding retrievalCaption→\\rightarrowembeddingVHop\-RouterGE2Qwen3\+\+VHop\-RouterGE2Qwen3GE2Qwen3Levelnn1\-shotGreedyHist\.1\-shotGreedyHist\.1\-shotGreedyHist\.1\-shotGreedyHist\.\(a\) Precise prompt,33\-hop trainingL1 \(re\-ID\)9990\.977\.878\.898\.092\.996\.027\.316\.218\.292\.973\.773\.769\.780\.874\.796\.0L2 \(spatial\)9892\.977\.681\.693\.989\.890\.827\.610\.214\.386\.764\.370\.467\.381\.678\.695\.9L3 \(color\)9793\.879\.483\.596\.989\.793\.814\.44\.17\.284\.566\.070\.174\.285\.682\.595\.9L4 \(dead ends\)30090\.770\.073\.797\.388\.093\.714\.78\.09\.786\.764\.065\.367\.377\.775\.796\.7L5 \(twins\)10089\.067\.076\.098\.088\.095\.014\.06\.08\.086\.066\.073\.059\.078\.066\.094\.0L6 \(logo\)9457\.442\.647\.955\.347\.950\.06\.45\.34\.356\.440\.443\.628\.758\.562\.880\.9L7 \(text\)9549\.537\.936\.841\.131\.633\.75\.31\.12\.149\.527\.429\.512\.673\.775\.880\.0\(b\) Precise prompt, mixed\-hop trainingL1 \(re\-ID\)9990\.977\.878\.898\.092\.996\.027\.316\.218\.292\.973\.773\.798\.080\.874\.769\.7L2 \(spatial\)9892\.977\.681\.693\.989\.890\.827\.610\.214\.386\.764\.370\.496\.981\.678\.668\.4L3 \(color\)9793\.879\.483\.596\.989\.793\.814\.44\.17\.284\.566\.070\.197\.985\.682\.578\.4L4 \(dead ends\)30090\.770\.073\.797\.388\.093\.714\.78\.09\.786\.764\.065\.399\.777\.775\.776\.7L5 \(twins\)10089\.067\.076\.098\.088\.095\.014\.06\.08\.086\.066\.073\.096\.078\.066\.074\.0L6 \(logo\)9457\.442\.647\.955\.347\.950\.06\.45\.34\.356\.440\.443\.650\.058\.562\.859\.6L7 \(text\)9549\.537\.936\.841\.131\.633\.75\.31\.12\.149\.527\.429\.527\.473\.775\.874\.7\(c\) Vague prompt,33\-hop trainingL1 \(re\-ID\)9946\.521\.227\.373\.745\.555\.624\.29\.110\.124\.213\.115\.247\.515\.214\.179\.8L2 \(spatial\)9856\.130\.640\.876\.559\.264\.317\.39\.210\.225\.513\.313\.345\.932\.720\.490\.8L3 \(color\)9655\.228\.135\.465\.647\.955\.211\.57\.37\.322\.910\.412\.533\.319\.815\.679\.2L4 \(dead ends\)30047\.019\.326\.366\.341\.347\.713\.36\.77\.718\.07\.79\.326\.321\.021\.377\.0L5 \(twins\)10034\.014\.019\.058\.031\.040\.010\.06\.06\.013\.07\.05\.031\.027\.020\.082\.0\(d\) Vague prompt, mixed\-hop trainingL1 \(re\-ID\)9946\.521\.227\.373\.745\.555\.624\.29\.110\.124\.213\.115\.287\.915\.214\.197\.0L2 \(spatial\)9856\.130\.640\.876\.559\.264\.317\.39\.210\.225\.513\.313\.393\.932\.720\.4100\.0L3 \(color\)9655\.228\.135\.465\.647\.955\.211\.57\.37\.322\.910\.412\.592\.719\.815\.696\.9L4 \(dead ends\)30047\.019\.326\.366\.341\.347\.713\.36\.77\.718\.07\.79\.389\.721\.021\.397\.0L5 \(twins\)10034\.014\.019\.058\.031\.040\.010\.06\.06\.013\.07\.05\.084\.027\.020\.095\.0
Table 11:Two\-hop success rate \(%\)\.Same corpora, checkpoints, and scoring conventions as Table[4](https://arxiv.org/html/2609.38743#S5.T4)\. L4 uses test worldsw1w\_\{1\}–w3w\_\{3\}; other levels use development worldw0w\_\{0\}\. Caption→\\toQwen3 retains the query image as pixels\. In block \(b\), the agent uses the learned per\-step retriever; the other trained\-agent blocks use the chain tool\. Vague prompts are evaluated on L1–L5\. Bold marks the best result within each training block on each side of the dashed rule\.Retrieval methodsAgentic searchEmbedding retrievalCaption→\\rightarrowembeddingVHop\-RouterGE2Qwen3\+\+VHop\-RouterGE2Qwen3GE2Qwen3Levelnn1\-shotGreedyHist\.1\-shotGreedyHist\.1\-shotGreedyHist\.1\-shotGreedyHist\.\(a\) Precise prompt,33\-hop trainingL1 \(re\-ID\)992\.013\.112\.11\.021\.212\.10\.02\.03\.00\.06\.15\.181\.890\.983\.896\.0L2 \(spatial\)982\.014\.39\.24\.120\.412\.20\.01\.03\.14\.15\.19\.244\.991\.888\.872\.4L3 \(color\)974\.113\.410\.32\.114\.45\.20\.00\.01\.01\.03\.14\.146\.484\.585\.679\.4L4 \(dead ends\)3002\.78\.38\.01\.012\.36\.00\.01\.00\.31\.04\.34\.345\.065\.366\.775\.0L5 \(twins\)1002\.08\.06\.02\.06\.04\.00\.00\.00\.00\.03\.04\.038\.050\.064\.070\.0L6 \(logo\)943\.23\.22\.14\.37\.44\.31\.10\.00\.02\.12\.11\.19\.623\.425\.528\.7L7 \(text\)955\.36\.34\.21\.12\.10\.01\.10\.00\.01\.11\.12\.12\.136\.843\.229\.5\(b\) Precise prompt, mixed\-hop trainingL1 \(re\-ID\)992\.013\.112\.11\.021\.212\.10\.02\.03\.00\.06\.15\.189\.990\.983\.889\.9L2 \(spatial\)982\.014\.39\.24\.120\.412\.20\.01\.03\.14\.15\.19\.292\.991\.888\.890\.8L3 \(color\)974\.113\.410\.32\.114\.45\.20\.00\.01\.01\.03\.14\.194\.884\.585\.692\.8L4 \(dead ends\)3002\.78\.38\.01\.012\.36\.00\.01\.00\.31\.04\.34\.390\.065\.366\.775\.3L5 \(twins\)1002\.08\.06\.02\.06\.04\.00\.00\.00\.00\.03\.04\.087\.050\.064\.077\.0L6 \(logo\)943\.23\.22\.14\.37\.44\.31\.10\.00\.02\.12\.11\.123\.423\.425\.526\.6L7 \(text\)955\.36\.34\.21\.12\.10\.01\.10\.00\.01\.11\.12\.16\.336\.843\.246\.3\(c\) Vague prompt,33\-hop trainingL1 \(re\-ID\)991\.08\.16\.12\.011\.13\.00\.02\.01\.00\.03\.01\.037\.423\.228\.379\.8L2 \(spatial\)490\.04\.14\.10\.014\.30\.02\.08\.24\.10\.04\.12\.040\.824\.530\.681\.6L3 \(color\)480\.06\.28\.30\.014\.66\.20\.00\.00\.00\.00\.00\.025\.06\.26\.272\.9L4 \(dead ends\)1500\.02\.71\.30\.77\.34\.00\.00\.00\.00\.01\.30\.020\.07\.36\.759\.3L5 \(twins\)500\.04\.02\.00\.010\.02\.00\.00\.00\.00\.00\.00\.022\.02\.010\.058\.0\(d\) Vague prompt, mixed\-hop trainingL1 \(re\-ID\)991\.08\.16\.12\.011\.13\.00\.02\.01\.00\.03\.01\.055\.623\.228\.380\.8L2 \(spatial\)490\.04\.14\.10\.014\.30\.02\.08\.24\.10\.04\.12\.059\.224\.530\.693\.9L3 \(color\)480\.06\.28\.30\.014\.66\.20\.00\.00\.00\.00\.00\.054\.26\.26\.283\.3L4 \(dead ends\)1500\.02\.71\.30\.77\.34\.00\.00\.00\.00\.01\.30\.053\.37\.36\.784\.7L5 \(twins\)500\.04\.02\.00\.010\.02\.00\.00\.00\.00\.00\.00\.048\.02\.010\.078\.0
### E\.2Agent Model and Retrieval Tool
##### Settings for Figure[1](https://arxiv.org/html/2609.38743#S1.F1)\.
Success rates use the same300300precise three\-hop L4 test queries\. Panel \(a\) compares Qwen3 andVHop\-Router, each standalone and as the Flash agent’s retrieval tool\.VHop\-Routeruses precise mixed\-hop training and checkpointmixh\_rl\_best\. Colors identify retrieval tools; circles denote standalone retrieval, and diamonds denote Flash agents\. The green dashed line connects the Pareto frontier among the four plotted configurations\.
The horizontal axis uses a linear scale\. It adds generated tokens, including thinking tokens, to retrieval hops, counting each hop as one token\. For agents, we sum tokens over all model calls and use the length of the recorded retrieval trajectory as the hop count\. These means come from a separate pilot on1010matched queries inw1w\_\{1\}\. Standalone Qwen3 uses five retrieval steps\. The standalone Router count is estimated as4\+2b4\+2b, wherebbis the mean number of backtracks, giving10\.1110\.11for the mixed\-hop policy\. The corresponding means for Flash with Qwen3 andVHop\-Routerare1,885\.81\{,\}885\.8and727\.9727\.9\. This is a token\-count convention, not a measure of latency or total computation\.
Untrained retrievers count an answer found anywhere in the returned images;VHop\-Routerallows1616policy actions and requires a correct answer atStop\. Agents are scored on their final answers\. Agents with GE2 or Qwen3 allow1616generation rounds; agents withVHop\-Routerallow88\. The generation limit is8,0008\{,\}000tokens per model call\. Panel \(b\) uses the same mixed\-hop Router and Flash \+ Qwen3 baseline: upgrading Flash to Pro adds3\.73\.7percentage points, while replacing Qwen3 withVHop\-Routeradds52\.752\.7\(Tables[4](https://arxiv.org/html/2609.38743#S5.T4)and[12](https://arxiv.org/html/2609.38743#A5.T12)\)\.
Figure[12](https://arxiv.org/html/2609.38743#A5.F12)extends Figure[1](https://arxiv.org/html/2609.38743#S1.F1)b to all seven levels and both tool training settings\. Table[12](https://arxiv.org/html/2609.38743#A5.T12)lists all model–tool combinations\. On L4, Flash \+ Qwen3 achieves32\.7%32\.7\\%success\. Upgrading Flash to Pro raises this to36\.3%36\.3\\%; replacing Qwen3 with the precise three\-hop or mixed\-hopVHop\-Routerraises it to59\.7%59\.7\\%or85\.4%85\.4\\%, respectively\. The gains are3\.73\.7,27\.027\.0, and52\.752\.7percentage points, computed before rounding\. With the mixed\-hop tool, Flash and Pro reach85\.4%85\.4\\%and84\.7%84\.7\\%\.
Table 12:Agent model and retrieval tool\.Three\-hop answer accuracy \(%\) with Gemini 3\.5 Flash or Gemini 3\.1 Pro, up to1616tool calls, and the same answer\-scoring convention\. Both Router tools use precise training prompts\. L4 uses the300300\-query test set; other levels use development worldw0w\_\{0\}\. The mean averages the seven levels\. Bold marks the highest value in each column\.Figure 12:Agent and tool upgrades across all seven levels\.Full version of Figure[1](https://arxiv.org/html/2609.38743#S1.F1)b, showing gains over Flash \+ Qwen3 on precise three\-hop queries, including both Router training configurations\. Router labels indicate training hops\. L4 uses the300300\-query test set; other levels use development worldw0w\_\{0\}\. Gains are computed before rounding\.
### E\.3Number of Images Returned per Call
Table[13](https://arxiv.org/html/2609.38743#A5.T13)reports accuracy and context usage for the retrieval\-width experiment\. Larger candidate sets improve the untrained tools, with GE2 reaching81\.3%81\.3\\%on L4 atk=50k=50\. The precise mixed\-hopVHop\-Routerreaches85\.4%85\.4\\%while returning fewer images\. The number of image occurrences sent to the model includes repeated conversation history; it is distinct from the number of retrieved images and from measured latency or monetary cost\.
Table 13:Full retrieval\-width comparison\.Three\-hop answer accuracy \(%\) with Flash and up to1616tool calls\. L4 uses test worldsw1w\_\{1\}–w3w\_\{3\}; other levels use development worldw0w\_\{0\}\. A chain call returns the policy’s active image chain\. L4 context statistics are per\-query means:*calls*counts tool calls,*images*counts returned images including repeats, and*sent*also counts repeated images in the conversation history\. Atk=50k=50, three GE2 and four Qwen3 conversations reached the request\-size limit and used the last top\-1 candidate as the fallback answer\. Dashes indicate settings not evaluated\.Retrieval toolkkL1L2L3L4L5L6L7callsimagessentGE2 \(untrained\)160\.658\.256\.719\.315\.06\.412\.614\.515120382\.881\.678\.446\.740\.010\.624\.211\.835249586\.986\.784\.561\.060\.019\.136\.89\.6483001092\.993\.988\.778\.072\.031\.942\.16\.96934450–––81\.3–––5\.42701084Qwen3 \(untrained\)173\.769\.477\.332\.728\.09\.614\.713\.013104366\.764\.371\.136\.726\.07\.422\.111\.936255575\.878\.682\.558\.054\.012\.832\.69\.3472951088\.985\.794\.869\.365\.025\.543\.27\.37339250–––78\.7–––5\.52751163Flash\+\+VHop\-Router\(33\-hop\)chain85\.962\.258\.859\.757\.012\.86\.32\.61442Flash\+\+VHop\-Router\(mixed\-hop\)chain80\.880\.689\.785\.484\.010\.65\.32\.11233
### E\.4Decision Penalty
Figure[13](https://arxiv.org/html/2609.38743#A5.F13)adds training reward and stopping behavior to the rollout length and test accuracy shown in the main text\. Table[14](https://arxiv.org/html/2609.38743#A5.T14)gives the checkpoint scores and prompt\-transfer results\. At iteration200200, theλ=0\\lambda=0andλ=0\.02\\lambda=0\.02runs reach48\.0%48\.0\\%and47\.7%47\.7\\%on L4\. On the vague prompts, the mean accuracy rises from26\.0%26\.0\\%to32\.3%32\.3\\%, although L4 decreases from20\.0%20\.0\\%to18\.0%18\.0\\%\. The mean stop rate increases and the mean number of backtracks decreases\. Each setting has one run, so these results do not establish statistical equivalence or seed\-level reliability\. Appendix[D\.4](https://arxiv.org/html/2609.38743#A4.SS4)records the optimization settings, including the gradient\-memory adjustments\.
Figure 13:Full training dynamics with a decision penalty\.Runs useλ∈\{0,0\.02,0\.05\}\\lambda\\in\\\{0,0\.02,0\.05\\\}with the same1616\-decision budget\.\(a\)Mean decisions per rollout\.\(b\)Unpenalized terminal rewardrr, defined identically across runs\.\(c\)Success rate on the preregistered300300\-question L4 test set\.\(d\)Fraction of sampled rollouts that issueStopbefore exhausting the budget\. Thin curves show per\-iteration values and thick curves show moving averages in \(a\), \(b\), and \(d\)\.Table 14:Decision\-penalty results\.Comparison ofλ=0\\lambda=0andλ=0\.02\\lambda=0\.02with gated decoding\.*In\-domain*rows use the L4 test set\.*Vague*rows measure stopping on a verified acceptable answer, using development worldw0w\_\{0\}except for L4, which uses the test worlds\.*Rephrased*rows use two natural\-language versions of the L4 development questions; they are separate from the human\-authored dataset\. Stop rate and backtracks are means over the five vague\-prompt levels\. Accuracy is in percent; stop rate is a fraction\. Bold marks the better value\. Each setting uses one seed\.
## Appendix FGeneralization and Limitations
### F\.1Three\-Object Scenes
The three\-object evaluation adds a distractor to each two\-object hop scene while preserving the question and target\. Query views and answer portraits are unchanged\. The added objects come from a separate pool\. The standaloneVHop\-Routerrows in Figure[6](https://arxiv.org/html/2609.38743#S5.F6)use the policy alone; Flash appears only in the rows explicitly labeled as agents\. No agent with the trained chain tool is included in that comparison\.
The precise three\-hop and mixed\-hop policies decrease from48\.0%48\.0\\%to30\.4%30\.4\\%and from76\.3%76\.3\\%to52\.2%52\.2\\%, respectively\. The perturbed results average the three test worlds, with9999,100100, and100100valid queries\. The drop shows sensitivity to competing visual content, but this experiment alone does not separate failures of object recognition from failures of search control\.
### F\.2Four\-Hop Extrapolation
The four\-hop evaluation extends L4 beyond the maximum training depth\. It uses9797questions over a600600\-image development corpus\. All policies receive precise prompts\. Table[15](https://arxiv.org/html/2609.38743#A6.T15)reports the better of two decoding settings for each checkpoint: a question\-specific minimum stop depth and stopping permitted after the first retrieval\. Thus the comparison gives each policy the benefit of either stopping constraint\.
Table 15:Four\-hop answer accuracy \(%\)\.Standalone policies on L4 four\-hop development queries, with a1616\-decision budget\. Values use the better of the two stopping constraints described above\.All four policies remain at or below8\.2%8\.2\\%\. For the precise three\-hop policy, the depth floor gives0\.0%0\.0\\%and removing it gives1\.0%1\.0\\%; the precise mixed\-hop policy obtains1\.0%1\.0\\%either way\. The depth floor therefore does not resolve the failure\. This test also increases the corpus from roughly450450to600600images and adds another layer of distractors\. It measures transfer to longer, harder searches, rather than isolating hop count alone\. Training on four\-hop worlds is a possible next experiment; recovery from such training has not been established here\.
### F\.3Realistic Queries and Settings
Table[16](https://arxiv.org/html/2609.38743#A6.T16)expands the results on realistic queries in Table[7](https://arxiv.org/html/2609.38743#S5.T7)\. All methods use the same5050questions and386386\-image corpus, without further training\. Annotators write the queries and specify everyday scenes; the images are re\-rendered while preserving the intended content\. These are human\-authored tasks with rendered images, so the experiment measures transfer beyond procedural prompts and scene templates, rather than performance on unmodified personal photographs\.
Table 16:Full results on realistic queries\.Answer accuracy and set success \(%\) for standaloneVHop\-Routerand Flash with each retrieval tool\. The metric definitions are given below\.##### Metrics\.
For standaloneVHop\-Router, answer accuracy \(answer\_acc\) requiresStopwith an acceptable image at the top of the stack\. Set success \(e2e\) checks whether an acceptable image is anywhere in the final active stack; popped images do not count\. For Flash, answer accuracy \(ans\_acc\) checks the final returned answer\. If there is no validFINALdeclaration, the evaluator uses the last answer candidate returned byVHop\-Router, or the last retrieved image for GE2 and Qwen3\. Set success \(final\_gt\) checks the cumulative returned image set across tool calls\. These set metrics therefore retain different histories for standalone policies and agents\.
##### Evaluation records and budgets\.
The precise three\-hop and mixed\-hop rows userloo10w10\_200andmixh\_rl\_best\. The archived human\-authored runs labelvgo\_rl\_bestas vague three\-hop andvgmix\_rl\_bestas vague mixed\-hop\. These are the checkpoints underlying the vague rows here; they differ from thevg\_rl\_best/vgo\_rl\_bestpair used for the vague blocks in the procedural main table\. The recorded Flash–VHop\-Routerruns allow up to eight chain calls, each with at most1616policy decisions\. GE2 and Qwen3 runs allow up to1616single\-step tool calls\. Each chain call may return multiple images, so a chain call and a single\-step call have different costs\.
The main table reports the agent’s set success:36\.0%36\.0\\%and34\.0%34\.0\\%for the two preciseVHop\-Routertools, compared with22\.0%22\.0\\%and26\.0%26\.0\\%for GE2 and Qwen3\. Their answer accuracies are20\.0%20\.0\\%,20\.0%20\.0\\%,12\.0%12\.0\\%, and8\.0%8\.0\\%\. The gap between finding an answer and returning it shows that answer selection remains a separate source of error\. With only5050queries, these results support a limited transfer claim, not reliable performance across arbitrary photo collections\.
### F\.4Extensions to Logo and Text Matching
The matching rule can extend beyond physical object identity\. L6 adds logo matching, and L7 adds printed\-text matching\. The rule is chosen independently for each hop and specified in the instruction; the actual logo or text must be read from the image\. These extensions retain the constraints and distractors from L1–L5\. Rendering checks also cover logos and printed text, and training and evaluation use disjoint logo and printed\-string pools\.
The L6 evaluation world contains9494queries,501501corpus images, and5050query images; L7 contains9595queries,481481corpus images, and5050query images\. Table[17](https://arxiv.org/html/2609.38743#A6.T17)reports the three\-hop results with precise prompts\. Tables[10](https://arxiv.org/html/2609.38743#A5.T10)and[11](https://arxiv.org/html/2609.38743#A5.T11)include the one\- and two\-hop results\. The trained policies use the same L4 checkpoints as the main evaluation, without further training on either extension\.
Table 17:Three\-hop success on logo and text matching \(%\)\.These rows extend Table[4](https://arxiv.org/html/2609.38743#S5.T4), with the same methods and scoring rules\. Both levels use precise prompts and development worldw0w\_\{0\}\. The strategy or training configuration is listed in the second column\.Training uses L4 identity links, so these results test transfer to matching rules absent from post\-training\. With precise mixed\-hop training, standaloneVHop\-Routerachieves5\.3%5\.3\\%on L6 and2\.1%2\.1\\%on L7; Flash with the same tool achieves10\.6%10\.6\\%and5\.3%5\.3\\%, respectively\. The results do not establish whether the main cause is reading the marks, applying the new matching rule, or controlling the resulting search\. Wider retrieval improves some L6/L7 agent results in Table[13](https://arxiv.org/html/2609.38743#A5.T13), but it does not resolve these questions\. Training with logo and text links, together with separate perception checks, would be needed to distinguish these possibilities\.
## Appendix GTrajectory Case Studies
These examples show the retrieval decisions behind the results\. In the figures,PushdenotesSelectandPopdenotesBacktrack; image identifiers refer to entries in the same world\.
### G\.1Recovering from a Wrong Branch
We compare checkpoints from precise three\-hop training on L4 development query2\_leftof\. The solver must find object Q to the right of a yellow object A, then A to the left of a green object B, and finally B alone\. The gold chain is00076→\\to00079→\\to00352\. Figure[14](https://arxiv.org/html/2609.38743#A7.F14)contrasts Online IL and RLVR after the same initial mistake; the following figures show the other methods on this query\.
Figure 14:Recovery after the same incorrect first choice\.Online IL and RLVR both first select00373\. Online IL continues along that branch and stops on a wrong answer\. RLVR backtracks withpbt=0\.636p\_\{\\mathrm\{bt\}\}=0\.636, retrieves the three gold images, and stops withpstop=0\.999p\_\{\\mathrm\{stop\}\}=0\.999\. Solid arrows show selections and the dashed return shows backtracking\. The lower rows summarize the other methods; the untrained and SFT trajectories use five fixed retrieval steps\.Figure 15:Untrained retrieval misses the first hop\.The complete five\-step greedy trajectory for2\_leftofends at00027\. Every retrieved image lies outside the gold chain \(red frames\)\. This decoder takes five steps without learned backtracking or stopping\.Figure 16:SFT retrieves the first hop, then leaves the chain\.The state\-encoder warm\-up retrieves00076\(green frame\), then selects00281and eventually ends at00278\. The full five\-step trajectory is shown with the same decoder as the untrained model\.Figure 17:Flash with GE2 on the same query\.All1616tool calls are shown, read down the left column and then down the right\. Each row gives the recorded query, image inputs, and retrieved image\. Call22retrieves the first gold hop, but neither the second hop nor the answer is retrieved\. The recorded outcome is incorrect\.Figure 18:Flash with Qwen3 on the same query\.The complete1616\-call trace uses the same layout as Figure[17](https://arxiv.org/html/2609.38743#A7.F17)\. None of the three gold images is retrieved\. The last thumbnail shows the final tool result; the figure does not supply a separate answer declaration\.
### G\.2Revising the Query and Selecting an Answer
An agent can help at two different points\. If a returned chain lacks the answer, it can call the policy again with revised text\. If the answer is already present, it can select that image even when the policy continues searching\. Figures[19](https://arxiv.org/html/2609.38743#A7.F19)and[20](https://arxiv.org/html/2609.38743#A7.F20)show these behaviors on two separate queries\.
Figure 19:Re\-querying the same policy recovers a valid chain\.On L4 development query19\_leftof, the initial policy call exhausts its budget and returns a chain without the answer\. Flash retries the same checkpoint and query image with revised text\. The eighth call retrieves00185→\\to00155→\\to00173and stops correctly\. Query strings are copied from the trace\. The recorded system answer is00173; the trace does not contain a separateFINALdeclaration\.Figure 20:Answer selection from the same returned chain\.On L4 development query33\_leftof, precise mixed\-hopVHop\-Routerretrieves the gold chain but continues searching until its1616\-action budget runs out\. Its returned stack contains six images and ends at the incorrect image00435\. Without re\-querying, Flash selects the correct third image,00247\.
### G\.3A Vague Query with Two Acceptable Answers
The vague question in Figure[21](https://arxiv.org/html/2609.38743#A7.F21)omits the relation that distinguishes two precise chains\. Both share the first hop, and both endpoints are verified answers\. The demonstration record specifies one branch; the outcome reward accepts either endpoint\. In this trace, RLVR first explores the demonstrated branch, backtracks, and stops at the other acceptable endpoint\. Online IL repeatedly backtracks and exhausts its budget without returning an answer\.
Figure 21:Two valid endpoints under one vague instruction\.Family55from L4 test worldw2w\_\{2\}\. Gray paths show the two verified chains; colored arrows show the recorded policy decisions\. Solid arrows are selections and dashed arrows are backtracks\. RLVR makes four selections, backtracks once, and stops at the alternative valid answer\. Online IL makes eight selections and eight backtracks, exhausting its1616\-decision budget\. Both are standalone policies\.Similar Articles
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
DeepVoyager-VL proposes a long-horizon multimodal deep-search framework that integrates visual evidence into intermediate reasoning, using a multimodal event graph for data synthesis and fine-tuning without reinforcement learning, achieving strong performance across ten benchmarks.
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Proposes Reroute, a training-free plug-in for vision-language models that replaces irreversible visual-token pruning with recoverable routing, allowing tokens to re-enter the pipeline later to improve grounding under aggressive token reduction while maintaining VQA performance.
Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
Q-RAG introduces a reinforcement learning-based fine-tuning approach for embedder models to enable efficient multi-step retrieval, achieving state-of-the-art results on long-context benchmarks up to 10M tokens. This method provides a resource-efficient alternative to fine-tuning small LLMs for complex multi-step search tasks.
VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents
This paper presents VLD-RAG, an agentic multimodal retrieval-augmented generation framework for question answering over long, visually-rich documents. It uses a page-preserving index and a verifier-guided agent workflow to improve cross-page evidence retrieval and reasoning, outperforming prior vision-based baselines on benchmarks like LongDocURL and MMLongBench-Doc.
RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
RetrievalRouter is a lightweight query-aware router that adaptively selects retrieval pipelines to improve accuracy and speed in document retrieval, outperforming static baselines.