空间策略而非动作:向量量化测地线作为 LLM 驱动智能体的工具

arXiv cs.AI 论文

摘要

本文提出了一种 LLM 驱动的智能体架构,将向量量化的测地线轨迹转化为可复用的工具,使快速的非推理模型能够达到链式思维(chain-of-thought)的性能,同时将决策延迟从分钟级降低到秒级。

arXiv:2610.00613v1 Announce Type: new Abstract: Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their spatial comprehension through an architecture combining geometrical tools with a LLM serving as a high-level orchestrator in grid-world environments. The agent first collects geodesic trajectories, which are then vector-quantized to extract a representative subset. Offline, the LLM associates a natural language description of the underlying behavioral patterns to each selected trajectory, making it a tool. Online, the LLM chooses the appropriate tool conditioned on the current state and goal. Low-level control is handled by primitive actions that execute the trajectory associated with the tool. From an agentic AI perspective, this approach separates learning into two levels: tool discovery is handled through unsupervised quantization of trajectories, while reasoning and decision-making are handled by the LLM. We test the approach in a partially observable dynamic 2D grid environment with an open vision-language model (Qwen3.6-35B-A3B). Pairing the geometry-derived tool library with an agent-centered zoom tool and a collision detection tool lets a fast, non-reasoning configuration match the goal-reaching rate of a much more costly chain-of-thought version, while cutting the cost of a decision from minutes to seconds.
查看原文
查看缓存全文

缓存时间: 2026/10/02 09:46

# Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents
Source: [https://arxiv.org/html/2610.00613](https://arxiv.org/html/2610.00613)
###### Abstract

Large language model \(LLM\) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns\. We investigate their spatial comprehension through an architecture combining geometrical tools with a LLM serving as a high\-level orchestrator in grid\-world environments\. The agent first collects geodesic trajectories, which are then vector\-quantized to extract a representative subset\. Offline, the LLM associates a natural language description of the underlying behavioral patterns to each selected trajectory, making it a tool\. Online, the LLM chooses the appropriate tool conditioned on the current state and goal\. Low\-level control is handled by primitive actions that execute the trajectory associated with the tool\. From an agentic AI perspective, this approach separates learning into two levels: tool discovery is handled through unsupervised quantization of trajectories, while reasoning and decision\-making are handled by the LLM\. We test the approach in a partially observable dynamic 2D grid environment with an open vision\-language model \(Qwen3\.6\-35B\-A3B\)\. Pairing the geometry\-derived tool library with an agent\-centered zoom tool and a collision detection tool lets a fast, non\-reasoning configuration match the goal\-reaching rate of a much more costly chain\-of\-thought version, while cutting the cost of a decision from minutes to seconds\.

## 1Introduction

Large Language Models \(LLMs\) have become the de\-facto controller in many recent agentic systems, chaining natural\-language reasoning steps with external tool calls to solve tasks that go beyond pure text generation\([Yao et al\., 2023](https://arxiv.org/html/2610.00613#bib.bib3);[Schick et al\., 2023](https://arxiv.org/html/2610.00613#bib.bib4);[Wei et al\., 2022](https://arxiv.org/html/2610.00613#bib.bib5)\)\. Spatial reasoning, however, remains a recognized weak spot: LLMs are trained on token sequences and have no built\-in notion of metric space, obstacles or continuous motion so they tend to fall back on superficial statistical regularities of text rather than genuine geometric reasoning about the environment they are asked to act in\. This is problematic for any application that requires an agent to move through, or reason about, physical or simulated space \- navigation, warehouse robotics, game\-playing agents, and embodied assistants among them\. Difficulties accumulate when the simulators used \(for instance in a reinforcement learning framework\) face uncertainties or use approximations about spatial movements or future scenarios\.

Interestingly, humans do not solve navigation problems by directly reasoning over raw coordinates either\. When giving directions inside a building, we do not say “move 1\.5 meters north, then 3\.2 meters east”; we say*“go straight ahead, take the elevator, go around the corner office, then turn left\.”*The instruction is a short composition of a handful of reusable, generic strategies \- go straight, take the elevator, turn, avoid an obstacle \- each of which packages several nontrivial low\-level motions into a single high\-level decision\. A closely related approach appears in recreational mathematics: solving a Rubik’s cube does not require deriving new moves from scratch for every configuration\. Instead, a small fixed library of move sequences, called*algorithms*in the cubing community \(e\.g\.,R U R’ U R U2 R’orF R U R’ U’ F’, see[https://alg\.cubing\.net](https://alg.cubing.net/)\), is reused across arbitrary starting positions; the solver’s actual skill lies in recognizing*which*algorithm applies to the current configuration, not in re\-deriving the moves themselves for each starting configuration\. We take this as our guiding analogy: if LLMs are weak at direct, low\-level geometric planning, perhaps they should not be asked to do it\. Instead, they can be asked to do what they are comparatively good at: recognizing patterns and selecting among a small set of well described, precomputed options\.

This paper explores exactly this division of labor for spatial tasks\. We equip an LLM\-driven agent with a library of*tools*, each one a short, reusable motion pattern extracted directly from the geometry of the environment rather than hand\-designed or learned through reinforcement learning; we view these geometry\-grounded modules as structured inductive biases\. Concretely, we \(i\) sample many*geodesic*trajectories \- shortest, obstacle\-aware paths \- inside a partially observable 2D grid world, capturing the environment’s topology and spatial structure, \(ii\) compress this large trajectory set into a small number of representative prototypes using a*vector quantization*\(VQ\) procedure, each prototype becoming a reusable strategy, \(iii\) in anLLM interpretationstep, use an LLM*offline*to translate each prototype into a natural\-language description of its geometric behavior, its strengths, and its risks, turning it into a callabletool, and \(iv\) use the same \(or another\) LLM*online*, at decision time, forLLM orchestration: picking the tool best suited to the current partial observation and goal\. Low\-level actuation is then simply “play out the chosen tool’s fixed action sequence,” mirroring the way a Rubik’s cube solver executes a memorized algorithm once it has recognized the pattern\.

From an agentic AI perspective, this design yields anagentic separationof two steps that are often entangled:*skill discovery*, handled by unsupervised geometric quantization and requiring no labels or reward signal, and*skill selection / reasoning*, handled by the LLM using its native strengths in language\-conditioned decision\-making\. Ourkey insightis that LLMs appear to be more effective as*orchestrators*of geometry\-derived skills than as direct spatial planners\.

We present the synthetic 2D grid world with moving obstacles and partial observability in Section[3](https://arxiv.org/html/2610.00613#S3)and the methodology in Section[4](https://arxiv.org/html/2610.00613#S4); results in Section[5](https://arxiv.org/html/2610.00613#S5)show that routing decisions through the geometry\-grounded tool library improves significantly the target\-reaching success rate over querying the LLM directly\.

## 2Related work

Our approach sits at the intersection of three lines of work that are, to our knowledge, rarely combined in this particular way\.

LLM agents and tool use\.LLM\-driven agents increasingly interleave natural\-language reasoning with calls to external tools or actions, as in ReAct\([Yao et al\., 2023](https://arxiv.org/html/2610.00613#bib.bib3)\)and Toolformer\([Schick et al\., 2023](https://arxiv.org/html/2610.00613#bib.bib4)\), and benefit from decomposing a task into intermediate reasoning steps, as in Chain\-of\-Thought paradigm of\([Wei et al\., 2022](https://arxiv.org/html/2610.00613#bib.bib5)\)\. These works treat tools mostly as external APIs \(calculators, search engines, code interpreters\); we instead*derive*the tool set itself from the geometry of the environment, so that the LLM orchestrates a set of skills that is neither given a priori not invented by the LLM alone\.

Quantization and skill discovery\.Vector quantization has a long history as a way to compress a large or continuous dataset into a small number of representative prototypes, from classical VQ codebooks to VQ\-VAE\-style discrete representation learning\([van den Oord et al\., 2017](https://arxiv.org/html/2610.00613#bib.bib6)\)and to measure\-theoretic versions\([Turinici, 2024b](https://arxiv.org/html/2610.00613#bib.bib2)\)which can handle signed measures \- a property that could be exploited to let new skills emerge from novel trajectory patterns rather than fixing the tool set in advance\.

Spatial reasoning and geodesics\.Geodesics \- minimal paths that reflect the intrinsic \(possibly curved, possibly obstacle\-constrained\) geometry of a space, as opposed to paths that are straight only in some extrinsic embedding \- are a classical tool in geometry and robotics for capturing the true topology of an environment\. A growing body of work documents that LLMs and vision\-language models struggle with exactly this kind of geometric reasoning when it is left implicit in raw or visual input:[Patel and Pavlick \(2022\)](https://arxiv.org/html/2610.00613#bib.bib16)find that only the largest LMs learn to ground directional/spatial concepts in a grid world, and only from explicit examples;[Aghzal et al\. \(2023\)](https://arxiv.org/html/2610.00613#bib.bib15)show with their PPNL benchmark that GPT\-4 plans spatially only when prompted to interleave reasoning and acting, and still fails at longer\-horizon temporal reasoning;[Martorell \(2025\)](https://arxiv.org/html/2610.00613#bib.bib9)show that grid\-navigation success depends heavily on how coordinates are textually encoded, scaling with model size;[Sandfuchs et al\. \(2026\)](https://arxiv.org/html/2610.00613#bib.bib11)find that under strict partial observability only reasoning\-tuned LLMs reliably reach the goal in text gridworlds, and still less efficiently than an oracle path;[Gao et al\. \(2025\)](https://arxiv.org/html/2610.00613#bib.bib10)use a Rubik’s\-cube benchmark to isolate long\-horizon spatial state tracking under partial observation, reporting a uniform0%0\\%pass rate for current LLM agents once external solver tools are removed; and[Rodríguez Salgado Gonzalo \(2026\)](https://arxiv.org/html/2610.00613#bib.bib17)show that vision\-enabled models solving maze images are largely performing brute\-force path enumeration in token space rather than genuine visual planning\. Complementary work instead sidesteps this weakness by keeping the LLM out of low\-level geometric computation:[Zhu et al\. \(2023\)](https://arxiv.org/html/2610.00613#bib.bib14)plan long\-horizon Minecraft tasks over structured text actions and memory rather than raw pixels;[Song et al\. \(2024\)](https://arxiv.org/html/2610.00613#bib.bib12)and[Rahimi et al\. \(2025\)](https://arxiv.org/html/2610.00613#bib.bib8)each pair the LLM with a dedicated non\-LLM module \- a text\-based topological map, or a vision\-based localizer \- that handles metric geometry so the LLM only reasons over a simplified representation;[Dang et al\. \(2025\)](https://arxiv.org/html/2610.00613#bib.bib13)equip an LLM wayfinding agent with externally maintained spatial\-memory representations instead of inferred geometry; and[Zhao et al\. \(2025\)](https://arxiv.org/html/2610.00613#bib.bib7)tackle city\-scale embodied question answering via a hierarchical planner/manager/actor agent built around an explicit object\-centric cognitive map rather than a single LLM call over raw city\-scale geometry\. Rather than asking an LLM to reason directly about such geometric objects, which is precisely the kind of task token\-based models handle poorly, we use geodesics purely as an offline data\-generation and skill\-discovery mechanism, and reserve the LLM’s role for the language\-native tasks of describing and selecting among the resulting skills\.

## 3Grid\-world environment

To study this approach in a controlled setting we use a 2D grid environment loosely inspired by lane\-crossing games such as Frogger\. The environment illustrated in Figure[1](https://arxiv.org/html/2610.00613#S3.F1)is defined as follows:

- •Agent:a single controlled cell \(blue\)\.
- •Cars:dynamic obstacles that move upward by one cell per time step, representing traffic the agent must avoid\.
- •Free space:cells that can be traversed safely at the current time\.
- •Grey cells:unobserved cells, which may contain additional cars outside the agent’s observation window\.
- •Target:the goal cell the agent must reach\.
- •Yellow highlight:marks the time steps at which a decision was taken\.
- •Actions:wait, up, down, left, right \- five primitive actions available at every time step\.
- •Partial observability:only a local window centered on the agent \(\) is visible; cells outside this window are grey and unknown\.

Theobjectiveis to reach the goal cell, if possible with minimum expected number of movements and elapsed time but imperatively avoiding collisions with cars\.

To this end we focus onmulti\-step decisions; we cannot take one decision at a time because the environment evolves quickly relative to the decision frequency of the LLM \- the same constraint that forces landers and deep\-space rovers to commit to sequences of actions rather than await step\-by\-step commands from Earth\. To keep the number of costly \(in terms of money or wall\-clock time\) LLM queries manageable, at each decision point the agent commits to and executes a fixed sequence of five future actions regardless of how the environment evolves in between; only then the next decision is taken\. This is precisely the setting that motivates our tool design: each tool introduced below is such a five\-action sequence, pre\-computed offline so that, online, the LLM only needs to pick one rather than plan it step by step\.

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_0_full_view.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_0_obs_view_for_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_1_zoom_view_for_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_10_full_view_pre_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_10_obs_view_for_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_10_zoom_view_for_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_8_full_view_pre_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_8_obs_view_for_llm.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/turn_8_zoom_view_for_llm.png)

Figure 1:All panels use the LLM Qwen3\.6\-35B\-A3B\. Top row, left to right: full grid att=0t=0, the partial observation given to the LLM att=0t=0, and its agent\-centered close\-up; then the full grid, partial observation and close\-up at the last decision point, just before the target is reached \(turn 10\)\. Bottom row: an intermediate decision \(turn 8\) – full grid, partial observation, and close\-up\. The agent only ever sees the observation window and the close\-up; the full\-grid panels are shown only to indicate what it could not see\.
## 4Method

Our architecture is built in several stages, executed in the order below: geodesic collection, offline quantization, tool creation, and online orchestration\.

### 4\.1Geodesics as shortest\-path connections

A geodesic captures the intrinsic geometry of a space by defining the minimal path between two locations; curvature and obstacles influence which path is actually shortest\. On a flat, obstacle\-free plane a geodesic is just a straight line, but as soon as the space is curved or obstacles are present, the shortest path can look very different from the naive straight\-line, see examples in Figure[2](https://arxiv.org/html/2610.00613#S4.F2)\.

![Refer to caption](https://arxiv.org/html/2610.00613v1/comparison_geodesic_par_sfo.png)

![Refer to caption](https://arxiv.org/html/2610.00613v1/env_geodesic2.png)

Figure 2:Left:geodesic between two cities on the 3D Earth surface: great\-circle geodesic in green versus the naive straight line on a 2D Mercator projection in red: the shortest path in the underlying geometry is not the visually straight one\.Right:a geodesic in our environment that has to dodge cars \(in red\), assuming for graphical purposes that cars are static \(do not move\) : initial cell is blue, target green and geodesic in yellow\. More direct routes are not available\.In our grid world, we treat the agent’s start cell, the goal cell, and the \(possibly moving\) obstacles as defining a discrete metric space, and compute geodesics as shortest admissible paths in this space \- i\.e\., paths that reach the goal in as few steps as possible while never crossing a car\. Because obstacles move and the environment is only partially observed, many such geodesics are sampled under different obstacle realizations and different partial\-information conditions, so that the resulting set reflects the range of spatial strategies that are useful across plausible configurations of the environment, rather than a single trajectory tailored to one specific instance\.

### 4\.2Vector quantization of trajectories

Sampling many geodesics is useful for capturing the environment’s topology, but it also produces far more trajectories than can be handed, individually, to an LLM as candidate actions: the resulting library would be too large to describe or to choose from efficiently\. We therefore compress the sampled trajectories using*vector quantization*\(VQ\): given a large set of trajectories, VQ selects a small number of representative prototypes that best summarize the whole set, in the same way that classical VQ or VQ\-VAE codebooks summarize a high\-dimensional dataset with a handful of codewords\([van den Oord et al\., 2017](https://arxiv.org/html/2610.00613#bib.bib6)\)\.

Concretely, we treat the empirical distribution of sampled geodesics as a \(possibly signed, cf\.\([Turinici, 2024b](https://arxiv.org/html/2610.00613#bib.bib2);[Turinici, 2024a](https://arxiv.org/html/2610.00613#bib.bib1)\)\) distribution and apply a measure quantization procedure \(like K\-medoids\) which findsKKrepresentative trajectories in the given set\. Figure[3](https://arxiv.org/html/2610.00613#S4.F3)illustrates the general idea on a simpler, well\-understood test case \- quantization of standard Brownian motion paths, also known as cubature\.

Figure 3:Quantization of standard Brownian motion \(also called cubature\), used here as a sanity\-check illustration of the VQ machinery on a classical stochastic\-process benchmark before applying it to grid geodesics\. Left:55representative trajectories selected out of1111\-step paths of a discrete Brownian motion\. Right:200200representative trajectories selected out of6464\-step Brownian motions, illustrating that the same procedure scales to larger, denser trajectory sets\.\([Turinici, 2024b](https://arxiv.org/html/2610.00613#bib.bib2);[Turinici, 2024a](https://arxiv.org/html/2610.00613#bib.bib1)\)\. of a Brownian motion\.Applied to the grid environment, we sample250250geodesic trajectories under varying obstacle configurations and quantize them down toK=5K=5representative prototypes \(Figure[4](https://arxiv.org/html/2610.00613#S4.F4)\)\.

Each of theKKprototypes gives a candidate tool, described in natural language in Section[4\.3](https://arxiv.org/html/2610.00613#S4.SS3)and made available to the online orchestrator in Section[4\.4](https://arxiv.org/html/2610.00613#S4.SS4)\. To this library we also add the primitive actions \(a primitive action followed by four “wait” actions\) in case the target is already close or a precise move is needed\.

![Refer to caption](https://arxiv.org/html/2610.00613v1/filled_cluster_DRDRD.png)![Refer to caption](https://arxiv.org/html/2610.00613v1/filled_cluster_RRRRR.png)![Refer to caption](https://arxiv.org/html/2610.00613v1/filled_cluster_DDDDD.png)![Refer to caption](https://arxiv.org/html/2610.00613v1/filled_cluster_DRRRR.png)![Refer to caption](https://arxiv.org/html/2610.00613v1/filled_cluster_DDRDD.png)

Figure 4:Left:the250250sampled geodesic trajectories in the grid environment, overlaid; each color is one geodesic\. Trajectories are drawn here from their true start cells; in the pipeline all start positions are reset to the origin before quantization\.Right:their vector quantization intoK=5K=5representative five\-step trajectories, one prototype per panel \(DRDRD,RRRRR,DDDDD,DRRRR,DDRDD– named by their action letters U/D/L/R/W\), each of which becomes a reusable tool\. To this library we add the five primitive toolsUWWWW,DWWWW,LWWWW,RWWWW,WWWWW\(one primitive move then four waits\)\. These tools are then used for the results in Figure[5](https://arxiv.org/html/2610.00613#S5.F5)\.
### 4\.3From quantized trajectories to LLM tools

Each of theKKquantized five\-step action sequences is, on its own, meaningless to an LLM: it is just a list of moves such asdown, right, down, right, down\. To make it usable as a decision\-time tool, we run an LLM*offline*, once per prototype, and prompt it to describe the trajectory in natural language: what net displacement it produces, what tactical advantages and risks it carries with respect to the moving\-car dynamics of the environment, and under what circumstances it should or should not be selected\. This offline pass converts a purely geometric object into a self\-containedTOOL; note that in principle this callable unit can, at run time, also report spatial metrics such as the remaining distance to the goal and locally observed obstacles, in addition to the natural\-language description attached to it\. We report in appendix one such tool description\. When the deterministic collision check of Section[5](https://arxiv.org/html/2610.00613#S5)is used, this offline description is shortened to just the net displacement and a one\-line strategic summary, since the collision and timing reasoning is then handled in code rather than by the LLM\.

Each of theKKtools is described in the same way, so that the online orchestrator always has access to a small library ofKKshort, natural\-language tool descriptions, rather than raw trajectories\. As discussed above, we also add the primitive actions to this library\.

### 4\.4Online orchestration

At decision time, the LLM does not plan a path from scratch\. Instead, it is given

- •the current partial observation i\.e\., the local window around the agent, including visible cars and, if within range, the target; in practice we use a vision\-enabled LLM, Qwen3\.6\-35B\-A3B, and input both the partial observation and an agent\-centered close\-up \(“zoom”\) of it, cf\. Figure[1](https://arxiv.org/html/2610.00613#S3.F1)\. The close\-up \(second input image\) was required for Qwen3\.6\-35B\-A3B to produce good results\.
- •the description of the goal \(i\.e\., its color\),
- •the natural\-language tool descriptions produced offline in Section[4\.3](https://arxiv.org/html/2610.00613#S4.SS3);
- •optionally, a collision detection tool that removes tools that would collide with observed cars \(but not with unobserved ones\)\.

Its task is purely one of*selection*: pick the single tool whose described behavior best matches the current situation, e\.g\., preferring a down\-right staircase tool when the lower\-right is clear and the target lies down\-right, or the all\-wait tool when no safe move is currently available\. Once a tool is selected, its fixed action sequence is executed open\-loop by the low\-level controller, irrespective of how the environment evolves during those five steps, and the cycle repeats at the next decision time\.

This separation is the central agentic design choice of the paper: geometric skill discovery \(Sections[4\.1](https://arxiv.org/html/2610.00613#S4.SS1)\-[4\.3](https://arxiv.org/html/2610.00613#S4.SS3)\) is unsupervised and requires no LLM involvement at all, while the LLM is only ever asked to do two language\-native tasks, describe a trajectory, and choose among descriptions, rather than to reason directly about coordinates, distances, or collision geometry\.

## 5Results

We evaluate the full pipeline \(geodesic sampling→\\toVQ intoKKtools→\\tooffline LLM tool descriptions→\\toonline LLM orchestration\) against a baseline in which the same LLM is queried directly on the raw grid observation, with no geodesic tools, no quantization, and no natural\-language skill descriptions i\.e\., the LLM must propose the next five actions itself from the same visual/textual input\. Implementation is provided in the github repository[https://github\.com/gabriel\-turinici/quantized\_geodesics\_agents](https://github.com/gabriel-turinici/quantized_geodesics_agents)\.

Figure[5](https://arxiv.org/html/2610.00613#S5.F5)shows a representative episode \(LLM: Qwen3\.6\-35B\-A3B\): for a sequence of decision points \("turns"\), it displays the full grid state just before the LLM decision \(Pre\-LLM\), the partial observation actually given to the LLM \(Obs\) together with its agent\-centered close\-up \(Zoom\), and the full grid state after the selected tool’s five\-action sequence has been executed \(Post\-LLM\)\. The agent \(blue cell\) is only ever shown its local observation window and the close\-up of it; the full\-grid panels are shown here purely to illustrate what the agent could not see\.

We tested several LLMs but most failed to ever reach the target, so we focus on the vision\-enabled modelQwen3\.6\-35B\-A3B, used both for offline tool description and for online tool selection in all runs\. A second, larger frontier vision model was also tried on a small number of episodes and showed the same qualitative trends \- tools helping over direct querying \- but our compute budget did not allow a full evaluation, so we report quantitative results for Qwen3\.6\-35B\-A3B only\. The results \(Table[1](https://arxiv.org/html/2610.00613#S5.T1)\) can be summarized in three observations:

- •first two lines: when thinking mode is disabled and the collision tool absent the model fails\. The detailed logs \(not shown here but reproducible with the github code\) show that the model tries to reproduce the collision detection computations manually, reasoning one step at a time but gets lost because of the complexity of the task
- •lines 3 and 4: when thinking mode is enabled: the absence of the collision\-detection tool is compensated by a high effort of around 20k tokens per turn and takes a non\-negligible amount of time; the geodesic\-VQ\-tool pipeline reaches the target successfully more often than the same LLM queried directly on the raw grid observation under matched data and query budget;
- •lines 5 and 6 : once a collision validation tool filters out unsafe tools before the LLM decides \(see also Appendix[A\.2](https://arxiv.org/html/2610.00613#A1.SS2)\), results come back to optimal and a fast non\-thinking configuration matches the success rate of full thinking mode at a small fraction of the cost\.

Success rate for Qwen3\.6\-35B\-A3BCollisionZoomAverageAverageThinkingGeodesiccheck\(close\-up\)timetokensSuccesstoolstooltoolpergenerated%\\%, \(NN\)turnper turn×\\times×\\times×\\times✓∼20\.5\\sim 20\.5sec\.12002%​\(100\)2\\%\\ \(100\)×\\times✓×\\times✓∼17\\sim 17sec\.8502%​\(100\)2\\%\\ \(100\)✓×\\times×\\times✓5 min\.19k54%​\(100\)54\\%\\ \(100\)✓✓×\\times✓5\.5 min\.20k66%​\(100\)66\\%\\ \(100\)✓✓✓✓∼9\.5\\sim 9\.5sec\.45065%​\(100\)65\\%\\ \(100\)×\\times✓✓✓∼5\\sim 5sec\.20068% \(100\)Table 1:Grid\-navigation success rate for Qwen3\.6\-35B\-A3B across reasoning mode and the geometry\-derived components: the geodesic tool library, a code collision check that filters unsafe tools before the LLM decides, and the agent\-centered close\-up image\. “Time / turn” is the wall\-clock cost of one decision\. All rows use the close\-up image, which Qwen3\.6\-35B\-A3B requires, without it results are very poor\. The final row shows that, with the collision check and the close\-up in place, the geodesic tools reach top quality \(comparable to full thinking mode\) at a small fraction of the time\.Figure 5:Sequence of decision turns: full grid before the decision, the partial observation given to the LLM, its agent\-centered close\-up \(“zoom”\), and the full grid after executing the selected tool’s five\-action sequence, until the target is reached at turn 10\. The LLM is Qwen3\.6\-35B\-A3B\.This improvement in target\-reaching success supports our central hypothesis: the bottleneck for LLM spatial reasoning is not necessarily the LLM’s decision\-making*per se*, but the representation on which that decision is made\. When the LLM is asked to reason directly over a raw grid of colored cells, it must implicitly perform geometric operations \(shortest paths, alternatives, collision timing\) that are known to be a poor match for token\-based reasoning\. When the same decision is reduced to choosing among a handful of pre\-computed, geometrically grounded, natural\-language\-described options, the LLM’s comparative strength \- language\-conditioned selection under a stated objective \- is what actually gets exercised, and performance improves markedly under matched conditions\. A further consequence is on cost: once a deterministic collision check removes unsafe tools before the query, the remaining choice is easy enough that chain\-of\-thought reasoning is no longer needed, and a non\-thinking model reaches the same success rate as full thinking mode while cutting the per\-decision cost from minutes to seconds \(Table[1](https://arxiv.org/html/2610.00613#S5.T1), last row\)\.

## 6Discussion and limitations

The grid\-world setting used here is intentionally simple: it isolates the core question of whether decoupling geometric skill discovery from LLM\-based orchestration helps spatial reasoning, without the confound of high\-dimensional perception or continuous control\. Several simplifications should be kept in mind when extrapolating the result\. First, the environment is fully discrete and low\-dimensional \(a 2D grid with a handful of moving obstacles\), so the offline quantization problem is comparatively easy; scaling to continuous state spaces, 3D navigation, or robotic manipulation will require replacing the discrete geodesic search and quantization with continuous analogues, which is a natural direction for future work\. Second, the tool library size \(K=5K=5\) was chosen for tractability of natural\-language description and orchestration, not tuned systematically; how success rate and description quality scale withKK, and how the signed\-measure mechanism performs at triggering new skills in a longer\-running, non\-stationary environment, are open questions\. Third, our comparison to a “direct\-query” baseline uses the same LLM in both conditions, which isolates the effect of the tool\-based representation but does not yet test whether smaller, possibly non\-vision LLMs would benefit similarly from the same architecture\. Fourth, the offline LLM tool descriptions were generated once and used verbatim; whether iteratively refining descriptions based on downstream orchestration failures would further improve performance is left for future work\. Finally, we report a single LLM \(Qwen3\.6\-35B\-A3B\); a larger frontier model was spot\-checked with consistent trends but not fully evaluated for budget reasons\.

A further observation, based on data not shown here, is that a major failure mode of the online orchestrator is perception rather than planning: the vision\-enabled LLM occasionally misreads the position of a nearby car by one cell, which is enough to turn an otherwise\-safe tool choice into a collision\. Even large vision\-enabled models remain unreliable at this kind of precise, cell\-level reading of the observation image, and prompt engineering alone did not remove these errors\.

## 7Conclusion

We presented a protocol that separates spatial skill discovery from spatial decision\-making in LLM\-driven agents\. Geodesic trajectories sampled from a partially observable grid environment are compressed via vector quantization into a small set of representative, reusable tools; an LLM is used offline to translate each such strategy into a natural\-language tool description, and online to select among these descriptions given the current state and goal\. This pipeline improves the target\-reaching success rate compared to querying the same LLM directly on raw observations and, when paired with a deterministic collision filter, matches the accuracy of an expensive thinking\-mode configuration at a fraction of its wall\-clock cost, suggesting that LLMs are more effective as orchestrators of geometry\-derived skills than as direct spatial planners\. We see this as an argument supporting a broader agentic\-AI design principle \- let unsupervised, domain\-appropriate methods discover skills, and let the LLM do what it does best: describe and choose between them \- and hope to test its generality on higher\-dimensional, continuous, and robotic navigation and manipulation settings in future work\.

## References

- M\. Aghzal, E\. Plaku, and Z\. YaoCan Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial\-temporal Reasoning\.Note:arXiv:2310\.03249Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Danget al\.\(2025\)P\. Dang, J\. Zhu, W\. Li, and J\. LaiA large language model\-based agent for wayfinding: simulation of spatial perception and memory\.Cartography and Geographic Information Science52\(4\),pp\. 350–369\.External Links:[Document](https://dx.doi.org/10.1080/15230406.2024.2405596)Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Gaoet al\.\(2025\)H\. Gao, Z\. Zhang, T\. Luo, K\. Yang, X\. Juan, J\. Qiu, T\. Chen, B\. He, H\. Zhao, H\. Zhou, S\. Liu, and M\. WangCubeBench: Diagnosing Interactive, Long\-Horizon Spatial Reasoning under Partial Observations\.Note:arXiv:2512\.23328Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Martorell \(2025\)N\. MartorellFrom Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid\-World Navigation Task\.InWorld Conference on Explainable Artificial Intelligence,pp\. 268–291\.Note:arXiv:2502\.16690Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Patel and Pavlick \(2022\)R\. Patel and E\. PavlickMapping Language Models to Grounded Conceptual Spaces\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Rahimiet al\.\(2025\)K\. Rahimi, Md\. W\. Haque, S\. Dasgupta, and M\. RahmanVision\-Based Localization and LLM\-based Navigation for Indoor Environments\.Note:arXiv:2508\.08120Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Rodríguez Salgado Gonzalo \(2026\)A\. Rodríguez Salgado GonzaloFrom Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning\.Note:arXiv:2603\.26839Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Sandfuchset al\.\(2026\)S\. Sandfuchs, M\. Melchert, and J\. FrochteLLMs for Text\-Based Exploration and Navigation under Partial Observability\.Note:arXiv:2604\.09604Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: Language Models Can Teach Themselves to Use Tools\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2302\.04761Cited by:[§1](https://arxiv.org/html/2610.00613#S1.p1.1),[§2](https://arxiv.org/html/2610.00613#S2.p2.1)\.
- Songet al\.\(2024\)S\. Song, S\. Kodagoda, A\. Gunatilake, M\. G\. Carmichael, K\. Thiyagarajan, and J\. MartinGuide\-LLM: An Embodied LLM Agent and Text\-Based Topological Map for Robotic Guidance of People with Visual Impairments\.Note:arXiv:2410\.20666Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Turinici \(2024a\)G\. TuriniciDiversity in Deep Generative Models and Generative AI\.InMachine Learning, Optimization, and Data Science,G\. Nicosia and al\. \(Eds\.\),Cham,pp\. 84–93\.Note:arXiv:2202\.09573External Links:ISBN 978\-3\-031\-53966\-4Cited by:[Figure 3](https://arxiv.org/html/2610.00613#S4.F3),[§4\.2](https://arxiv.org/html/2610.00613#S4.SS2.p2.1)\.
- Turinici \(2024b\)G\. TuriniciHuber\-energy measure quantization\.Statistics and Computing35\(1\),pp\. 5\.External Links:ISSN 1573\-1375,[Link](https://doi.org/10.1007/s11222-024-10540-3),[Document](https://dx.doi.org/10.1007/s11222-024-10540-3)Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p3.1),[Figure 3](https://arxiv.org/html/2610.00613#S4.F3),[§4\.2](https://arxiv.org/html/2610.00613#S4.SS2.p2.1),[Remark 1](https://arxiv.org/html/2610.00613#Thmtheorem1.p1.1.1)\.
- van den Oordet al\.\(2017\)A\. van den Oord, O\. Vinyals, and K\. KavukcuogluNeural Discrete Representation Learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:1711\.00937Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p3.1),[§4\.2](https://arxiv.org/html/2610.00613#S4.SS2.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. ZhouChain\-of\-Thought Prompting Elicits Reasoning in Large Language Models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2201\.11903Cited by:[§1](https://arxiv.org/html/2610.00613#S1.p1.1),[§2](https://arxiv.org/html/2610.00613#S2.p2.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: Synergizing Reasoning and Acting in Language Models\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2210\.03629Cited by:[§1](https://arxiv.org/html/2610.00613#S1.p1.1),[§2](https://arxiv.org/html/2610.00613#S2.p2.1)\.
- Zhaoet al\.\(2025\)Y\. Zhao, K\. Xu, Z\. Zhu, Y\. Hu, Z\. Zheng, Y\. Chen, Y\. Ji, C\. Gao, Y\. Li, and J\. HuangCityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space\.Note:arXiv:2502\.12532Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.
- Zhuet al\.\(2023\)X\. Zhu, Y\. Chen, H\. Tian, W\. Tao, W\. Su, C\. Yang, G\. Huang, B\. Li, L\. Lu, X\. Wang, Y\. Qiao, Z\. Zhang, and J\. DaiGhost in the Minecraft: Generally Capable Agents for Open\-World Environments via Large Language Models with Text\-based Knowledge and Memory\.Note:arXiv:2305\.17144Cited by:[§2](https://arxiv.org/html/2610.00613#S2.p4.1)\.

## Appendix ATechnical appendices and supplementary material

### A\.1Example of tool description

The box below shows the tool description, generated automatically\. The LLM is Qwen3\.6\-35B\-A3B; when the deterministic collision check is enabled the offline prompt is shortened to just the net displacement and a one\-line strategic summary; prompts are in the repository\.

```
Each tool is a list of actions that affect the position of the agent. When
  a tool is selected all its actions are implemented until collision,
  reaching the goal, or exhaustion of the list. Each tool is named by its
  actions, one letter per step: U=up (-1,0), D=down (1,0), L=left (0,-1),
  R=right (0,1), W=wait (0,0).

  Available tools are:

 ====================tool WWWWW ====================:
 total displacement: no movement
action list: hold five steps
strategy: does nothing
 ====================end description tool WWWWW ====================:

 ====================tool UWWWW ====================:
 total displacement: one row up, no sideways change
action list: up once, then hold four steps
strategy: a single step up followed by waiting
 ====================end description tool UWWWW ====================:

 ====================tool DWWWW ====================:
 total displacement: one row down, no sideways change
action list: down once, then hold four steps
strategy: a single step down followed by waiting
 ====================end description tool DWWWW ====================:

 ====================tool LWWWW ====================:
 total displacement: one column to the left, no vertical change
action list: left once, then hold four steps
strategy: a single step left followed by waiting
 ====================end description tool LWWWW ====================:

 ====================tool RWWWW ====================:
 total displacement: one column to the right, no vertical change
action list: right once, then hold four steps
strategy: a single step right followed by waiting
 ====================end description tool RWWWW ====================:

 ====================tool DRDRD ====================:
 total displacement: three rows down and two columns to the right
action list: down once, then right once, then down once, then right once,
then down once strategy: a staircase pattern moving steadily toward the
bottom-right
 ====================end description tool DRDRD ====================:

 ====================tool RRRRR ====================:
 total displacement: five columns to the right, no row change
action list: right five times
strategy: a straight run towards the right side
 ====================end description tool RRRRR ====================:

 ====================tool DDDDD ====================:
 total displacement: five rows down, no sideways change
action list: down five times
strategy: a straight vertical run downwards
 ====================end description tool DDDDD ====================:

 ====================tool DRRRR ====================:
 total displacement: one row down and four columns to the right
action list: down once, then right four times
strategy: a single step down followed by a straight run to the right
 ====================end description tool DRRRR ====================:

 ====================tool DDRDD ====================:
 total displacement: four rows down and one column to the right
action list: down twice, then right once, then down twice
strategy: a vertical descent with a single rightward step
 ====================end description tool DDRDD ====================:
```

### A\.2Reproducibility details

Environment\.A25×2525\\times 25grid; the agent starts near the top\-left and the target is a fixed cell near the bottom\-right\. Cars move up by one row per step and wrap from the top row back to the bottom row; the agent does not wrap and grid edges are hard walls \(a move into a wall is absorbed\)\. A collision is a car entering the agent’s cell, or the agent and a car swapping cells, on the same step\. The agent observes only a local window \(half\-size1010\) centered on it; everything else is unobserved\. Each decision commits the next five primitive actions, executed open\-loop, before the next decision is taken\.

Tools\.250250geodesics are sampled under randomized car configurations and partial information, reset to a common origin, and quantized with K\-medoids toK=5K=5prototypes \(Figure[4](https://arxiv.org/html/2610.00613#S4.F4)\), named by their action letters \(U/D/L/R/W\)\. The five one\-move primitives \(one move then four waits\) are added, giving a1010\-tool library\.

Collision\-check tool\.When enabled, before the LLM is queried a deterministic routine replays each candidate tool against only the cars currently inside the observation window and discards any tool that would collide\. If no tool survives, the agent waits; if exactly one survives it is taken with no LLM call; otherwise the surviving tools \(and their descriptions\) are the only options offered to the LLM, whose answer is further constrained to that set\. This removes collision geometry from the LLM’s task and is what allows the fast non\-thinking configuration in Table[1](https://arxiv.org/html/2610.00613#S5.T1)\. Note that there are situations when cars are not visible so that collision detection tool is not always guaranteed to work, it only gives the best conclusion given available data\.

Close\-up image\.The “zoom” is a fixed\-size crop of the observation centered on the agent, rendered at higher resolution; it is passed alongside the wide observation whenever the observation is passed \(Figures[1](https://arxiv.org/html/2610.00613#S3.F1),[5](https://arxiv.org/html/2610.00613#S5.F5)\)\.

Evaluation\.An episode succeeds if the agent occupies the target cell at any step, and fails on collision or after3030decisions\. Each table cell reports successes over episodes \(NNgiven\); episode counts differ across conditions because of the compute budget\. Both the offline tool descriptions and the online selection use the same model, Qwen3\.6\-35B\-A3B\.

## Appendix BMore detailed trajectory plot

We present in Figures[6](https://arxiv.org/html/2610.00613#A2.F6)\-[7](https://arxiv.org/html/2610.00613#A2.F7)a detailed version of Figure[5](https://arxiv.org/html/2610.00613#S5.F5)\(the same episode, LLM Qwen3\.6\-35B\-A3B\)\. Each row also shows the agent\-centered close\-up \(“zoom”\) next to the observation\.

Figure 6:Detailed version of Figure[5](https://arxiv.org/html/2610.00613#S5.F5)\(columns: full grid pre\-decision, LLM observation, agent\-centered close\-up, full grid post\-execution\)\. Turns 1–5\. LLM: Qwen3\.6\-35B\-A3B\.Figure 7:Detailed version of Figure[5](https://arxiv.org/html/2610.00613#S5.F5), turns 6–10 \(target reached at turn 10\)\. LLM: Qwen3\.6\-35B\-A3B\.

相似文章

SpatialAct: 探索VLM智能体在3D场景中的空间推理到行动的能力

Hugging Face Daily Papers

SpatialAct是一个新的基于模拟器的基准,用于探索VLM智能体是否能在多轮反馈设置下进行连贯的空间推理并将其转化为3D环境中的行动。实验揭示了一个显著的推理到行动差距:当前的VLM尽管在孤立推理任务上表现良好,但难以维持空间信念并产生可靠的行为。