SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
Summary
SpatialCORE is a post-training framework that uses a model's confidence in its generated bounding-box grounding as a learning signal for spatial reasoning in large vision-language models, achieving state-of-the-art results among open-source models on spatial reasoning benchmarks.
View Cached Full Text
Cached at: 10/01/26, 04:23 PM
Paper page - SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision–Language Models
Source: https://huggingface.co/papers/2609.38716
Abstract
LargeVision-LanguageModels(LVLMs)havemaderemarkableprogressacrossvisualperceptiontasks,yetspatialreasoningremainsapersistentweakness,especiallyforquestionsthatrequirereasoningovervisualspace.Recentspatial-reasoningmethodsincorporategeneratedgrounding,wheremodelspredictboundingboxes,masks,orotherlocalizationoutputsfortask-relevantobjectsaspartoftheirreasoningtrace.However,theseapproachestypicallyoptimizefinal-answercorrectnessalone,allowingcorrectanswerstoberewardedevenwhenthemodeldoesnotreasonfromconfidentlylocalizedtask-relevantobjects.WeintroduceSpatialCORE(SpatiallyCOnfidentREasoning),apost-trainingframeworkthatturnsthemodel’sownconfidenceingeneratedgroundingintoalearningsignalforspatialreasoning.Itscentralideaistoreinforcegroundingthatisbothaccurateandconfident,encouragingthemodeltoreasonfromconfidentlylocalizedtask-relevantobjects.SpatialCORErealizesthisthroughaself-regulatingspatialrewardthatweightseachpredictedboundingbox’smatchingqualitybyitscoordinate-tokenconfidence.Ananswergatefurthertiesgroundingoptimizationtofinal-answercorrectness.SpatialCOREachievesstate-of-the-artresultsamongopen-sourceandspecializedspatialreasoningmodelsacrossdiversebenchmarks,andtransferseffectivelyinzero-shotsettingstounseendatadistributions.Thesourcecodeisavailableathttps://github.com/rafiibnsultan/SpatialCORE.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.38716
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38716 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.38716 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38716 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
This paper introduces SR-REAL, a unified framework for spatial vision-language models that combines linguistic deduction and 3D geometric reasoning via reinforcement learning, enabling robust multi-step spatial reasoning across diverse tasks.
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
SpatialCLI presents a framework that trains vision-language models to use specialist spatial tools and then internalize those capabilities, boosting Qwen3-VL-8B-Instruct on the MindCube benchmark from 29.3% to 84.6% with tools.
Soft Spatial Reasoning
The paper proposes Soft Spatial Reasoning, a post-training framework that introduces adaptive "soft thinking" for spatial tasks in Large Vision-Language Models, where intermediate reasoning steps mix token embeddings instead of committing to a single discrete token. An AdaptSoft controller dynamically adjusts softness based on hidden state and predictive uncertainty, outperforming hard and fixed-soft CoT baselines on spatial benchmarks.
SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
SpatialSpeak introduces a two-stage framework for spatial reasoning in vision-language models, using QA-native reconstruction pretraining and spatial chain-of-thought learning to achieve state-of-the-art performance on benchmarks.
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
This paper introduces RIS, a framework for spatial-semantic grounded latent visual reasoning in Multimodal Large Language Models to overcome information bottlenecks. It proposes anchoring latent tokens to spatial and semantic evidence, showing improvements on benchmarks like V* and HRBench.