Tag
The paper proposes Soft Spatial Reasoning, a post-training framework that introduces adaptive "soft thinking" for spatial tasks in Large Vision-Language Models, where intermediate reasoning steps mix token embeddings instead of committing to a single discrete token. An AdaptSoft controller dynamically adjusts softness based on hidden state and predictive uncertainty, outperforming hard and fixed-soft CoT baselines on spatial benchmarks.
SpatialCORE is a post-training framework that uses a model's confidence in its generated bounding-box grounding as a learning signal for spatial reasoning in large vision-language models, achieving state-of-the-art results among open-source models on spatial reasoning benchmarks.
LeRF introduces a method to enhance perspective-taking reasoning in vision-language models by learning reference coordinate frames, improving performance on benchmarks through supervised fine-tuning and reinforcement learning.
SpatialSpeak introduces a two-stage framework for spatial reasoning in vision-language models, using QA-native reconstruction pretraining and spatial chain-of-thought learning to achieve state-of-the-art performance on benchmarks.
Epoch AI created the Furniture Assembly Benchmark (FAB) to test AI models' visual reasoning in spotting mistakes in IKEA assembly photos. OpenAI's GPT-6 Astra leads with 80% accuracy, showing major improvements over previous models.
Spatial-Interactor is a framework that trains vision-language models to enhance spatial reasoning through interaction with the physical world, employing a three-level curriculum and two-stage training strategy to improve state transition modeling and long-horizon integration.
HarnessVLN is a zero-shot, training-free framework for embodied navigation that unifies perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface, achieving state-of-the-art results on benchmarks like R2R and RxR.
Astra achieved the highest score on Blueprint-Bench 2, a benchmark testing AI agents' ability to convert apartment photos into accurate 2D floor plans using spatial reasoning and cross-apartment learning.
Astra achieves state-of-the-art performance on MazeBench, significantly outperforming GPT-6 in a 3D open world spatial reasoning evaluation.
The author discusses demos of GPT-6 Astra demonstrating advanced iterative 3D modeling capabilities, including reasoning about geometry and progressively improving results, and asks what technical changes are driving this leap.
Astra reportedly surpasses other leading models in spatial and 3D understanding, achieving first place on the VoxelBench benchmark with an Elo rating exceeding 2600 and a lead of over 300 points.
ZDTaichu5.0-9B is a multimodal foundation model that combines a Qwen3.5-9B language backbone with a C-RADIOv4-H vision encoder, excelling in general visual understanding, spatial reasoning, and agent tasks among 10B-scale VLMs.
The paper introduces RoboSPA, a large-scale benchmark for evaluating vision-language-action models on fine-grained spatial reasoning and long-horizon procedural planning in robotic manipulation.
Playco used GPT-6 Astra to build Playbot, an AI-powered IDE for game developers, cutting manual fixes by 50% and improving spatial reasoning and prototyping efficiency.
This paper introduces FactoSR, a factorized reinforcement learning framework that enhances spatial reasoning in Vision-Language Models by decomposing 4D properties into orthogonal sub-objectives, achieving significant performance boosts on multi-view and video benchmarks.
Introduces Autoregressive Mosaics (AM-Bench), a benchmark to evaluate whether text-only LLMs have genuine 2D spatial reasoning abilities, distinct from code generation. Findings show spatial reasoning varies among models and is influenced by output medium like SVG vs. code.
UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior.
The article discusses how AI models struggle with intelligence tests like spatial reasoning and memory puzzles, highlighting gaps compared to human cognition and inviting readers to test their own skills.
The paper introduces GUI-Primitives, a benchmark of 994 contrastive instruction pairs to diagnose spatial reasoning failures in vision-language models for GUI grounding, revealing that most failures stem from candidate localization rather than relation understanding.
StateSight is a new benchmark for evaluating spatial-state reconstruction in vision-language models, showing that models like GPT-5.5 and Claude Sonnet 5 struggle with spatial reasoning tasks compared to human performance.