SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

Hugging Face Daily Papers Papers

Summary

SpatialCORE is a post-training framework that uses a model's confidence in its generated bounding-box grounding as a learning signal for spatial reasoning in large vision-language models, achieving state-of-the-art results among open-source models on spatial reasoning benchmarks.

Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:23 PM

Paper page - SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision–Language Models

Source: https://huggingface.co/papers/2609.38716

Abstract

LargeVision-LanguageModels(LVLMs)havemaderemarkableprogressacrossvisualperceptiontasks,yetspatialreasoningremainsapersistentweakness,especiallyforquestionsthatrequirereasoningovervisualspace.Recentspatial-reasoningmethodsincorporategeneratedgrounding,wheremodelspredictboundingboxes,masks,orotherlocalizationoutputsfortask-relevantobjectsaspartoftheirreasoningtrace.However,theseapproachestypicallyoptimizefinal-answercorrectnessalone,allowingcorrectanswerstoberewardedevenwhenthemodeldoesnotreasonfromconfidentlylocalizedtask-relevantobjects.WeintroduceSpatialCORE(SpatiallyCOnfidentREasoning),apost-trainingframeworkthatturnsthemodel’sownconfidenceingeneratedgroundingintoalearningsignalforspatialreasoning.Itscentralideaistoreinforcegroundingthatisbothaccurateandconfident,encouragingthemodeltoreasonfromconfidentlylocalizedtask-relevantobjects.SpatialCORErealizesthisthroughaself-regulatingspatialrewardthatweightseachpredictedboundingbox’smatchingqualitybyitscoordinate-tokenconfidence.Ananswergatefurthertiesgroundingoptimizationtofinal-answercorrectness.SpatialCOREachievesstate-of-the-artresultsamongopen-sourceandspecializedspatialreasoningmodelsacrossdiversebenchmarks,andtransferseffectivelyinzero-shotsettingstounseendatadistributions.Thesourcecodeisavailableathttps://github.com/rafiibnsultan/SpatialCORE.

View arXiv pageView PDFGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.38716

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38716 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38716 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38716 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Hugging Face Daily Papers

This paper introduces SR-REAL, a unified framework for spatial vision-language models that combines linguistic deduction and 3D geometric reasoning via reinforcement learning, enabling robust multi-step spatial reasoning across diverse tasks.

Soft Spatial Reasoning

Hugging Face Daily Papers

The paper proposes Soft Spatial Reasoning, a post-training framework that introduces adaptive "soft thinking" for spatial tasks in Large Vision-Language Models, where intermediate reasoning steps mix token embeddings instead of committing to a single discrete token. An AdaptSoft controller dynamically adjusts softness based on hidden state and predictive uncertainty, outperforming hard and fixed-soft CoT baselines on spatial benchmarks.