Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Hugging Face Daily Papers Papers

Summary

Introduces Autoregressive Mosaics (AM-Bench), a benchmark to evaluate whether text-only LLMs have genuine 2D spatial reasoning abilities, distinct from code generation. Findings show spatial reasoning varies among models and is influenced by output medium like SVG vs. code.

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translate spatial descriptions into code. We introduce Autoregressive Mosaics (AM-Bench), a benchmark that separates these factors: First, a translation task gives a model a fully specified geometry of a picture in words as a prompt and asks for the code that produces it. Second, a layout task requires the model to compose an image from an underspecified prompt. Across eight open-weight text-and-code-only models, all models reliably translate specified geometry into code, but their open-ended layout performance differs substantially, indicating that these differences are not explained by code-generation ability alone. An output-medium ablation further shows that the interface or medium of expression that the model uses matters: replacing procedural code with raw SVG improves layout scores across all models. Finally, probing model activations shows that a coarse layout plan is present before generation, but reflects only the layout implied by the prompt. During generation, models track the evolving geometric state instead of executing an initially fixed plan. Overall, these results show that 2D spatial performance in text-only LLMs depends on both the model and the output medium, and is not explained by code-generation ability alone.
Original Article
View Cached Full Text

Cached at: 09/03/26, 11:52 AM

Paper page - Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Source: https://huggingface.co/papers/2608.30751 TL;DR We introduce Autoregressive Mosaics (AM-Bench) to evaluate whether text-only LLMs possess genuine 2D spatial reasoning or merely translate spatial text into syntax.

The Core Problem When LLMs generate code that draws an image, it remains unclear if this reflects an internal understanding of 2D spatial layouts or simply an adept translation of explicit geometric instructions into code.

Methodology AM-Bench decouples spatial reasoning from code-generation ability through two targeted tasks. The translation task provides a fully specified geometric description and requires the model to generate corresponding code. The layout task provides an underspecified prompt, forcing the model to autonomously compose the 2D spatial arrangement.

Key Findings Across eight open-weight models, baseline translation ability is universally high, yet open-ended layout performance varies drastically, proving that spatial reasoning is distinct from pure code generation. The output medium also heavily dictates performance; replacing procedural code with raw SVG universally improves layout scores. Finally, activation probing reveals that models form only a coarse initial plan based on the prompt, dynamically tracking the evolving geometric state during autoregressive generation rather than executing a fixed structural plan.

Similar Articles

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

Papers with Code Trending

OVO-S-Bench introduces a comprehensive human-annotated benchmark of 1,680 questions across 348 videos to evaluate streaming spatial intelligence in multimodal LLMs, revealing that even the best model (Gemini-3.1-Pro) trails human experts by 27 points. The benchmark exposes key limitations including allocentric mapping as a major bottleneck and chain-of-thought reasoning amplifying spatial errors.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

Hugging Face Daily Papers

SpatialAct is a new simulator-grounded benchmark that probes whether VLM agents can perform coherent spatial reasoning and translate it into actions in 3D environments across multi-turn feedback settings. Experiments reveal a significant reasoning-to-action gap, with current VLMs struggling to maintain spatial beliefs and produce reliable actions despite performing well on isolated reasoning tasks.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

Hugging Face Daily Papers

This paper introduces ReasonMatch-Bench, a benchmark for wide-baseline matching in multimodal LLMs, and proposes Dynamic Correspondence Reinforcement Learning (DCRL) to improve spatial reasoning. Experiments show significant gains on the benchmark while maintaining general performance.