Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models
Summary
Introduces Autoregressive Mosaics (AM-Bench), a benchmark to evaluate whether text-only LLMs have genuine 2D spatial reasoning abilities, distinct from code generation. Findings show spatial reasoning varies among models and is influenced by output medium like SVG vs. code.
View Cached Full Text
Cached at: 09/03/26, 11:52 AM
Paper page - Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models
Source: https://huggingface.co/papers/2608.30751 TL;DR We introduce Autoregressive Mosaics (AM-Bench) to evaluate whether text-only LLMs possess genuine 2D spatial reasoning or merely translate spatial text into syntax.
The Core Problem When LLMs generate code that draws an image, it remains unclear if this reflects an internal understanding of 2D spatial layouts or simply an adept translation of explicit geometric instructions into code.
Methodology AM-Bench decouples spatial reasoning from code-generation ability through two targeted tasks. The translation task provides a fully specified geometric description and requires the model to generate corresponding code. The layout task provides an underspecified prompt, forcing the model to autonomously compose the 2D spatial arrangement.
Key Findings Across eight open-weight models, baseline translation ability is universally high, yet open-ended layout performance varies drastically, proving that spatial reasoning is distinct from pure code generation. The output medium also heavily dictates performance; replacing procedural code with raw SVG universally improves layout scores. Finally, activation probing reveals that models form only a coarse initial plan based on the prompt, dynamically tracking the evolving geometric state during autoregressive generation rather than executing a fixed structural plan.
Similar Articles
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
SpatialWorld is a unified benchmark for evaluating interactive spatial reasoning in multimodal agents across diverse real-world tasks, revealing that even the strongest models achieve low task success rates.
Spatial Reasoning via Modality Switching Between Language and Symbolic Representation
This paper explores grounding multi-hop textual-spatial stories into geometry-aware modalities like grids, showing a 42% performance improvement when switching from language-only to grid-based reasoning, and introduces a switching metric for modality selection in LLMs.
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
OVO-S-Bench introduces a comprehensive human-annotated benchmark of 1,680 questions across 348 videos to evaluate streaming spatial intelligence in multimodal LLMs, revealing that even the best model (Gemini-3.1-Pro) trails human experts by 27 points. The benchmark exposes key limitations including allocentric mapping as a major bottleneck and chain-of-thought reasoning amplifying spatial errors.
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
SpatialAct is a new simulator-grounded benchmark that probes whether VLM agents can perform coherent spatial reasoning and translate it into actions in 3D environments across multi-turn feedback settings. Experiments reveal a significant reasoning-to-action gap, with current VLMs struggling to maintain spatial beliefs and produce reliable actions despite performing well on isolated reasoning tasks.
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
This paper introduces ReasonMatch-Bench, a benchmark for wide-baseline matching in multimodal LLMs, and proposes Dynamic Correspondence Reinforcement Learning (DCRL) to improve spatial reasoning. Experiments show significant gains on the benchmark while maintaining general performance.