Tag
Introduces Autoregressive Mosaics (AM-Bench), a benchmark to evaluate whether text-only LLMs have genuine 2D spatial reasoning abilities, distinct from code generation. Findings show spatial reasoning varies among models and is influenced by output medium like SVG vs. code.
This paper reproduces the phenomenon of answer pre-commitment in an open-weight LLM (Qwen3-8B) using a minimal car-wash question and provides preliminary activation-level evidence that the commitment is encoded in hidden states before the answer text is emitted.