iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Summary
Introduces iVGR, a reinforcement learning framework that internalizes visual localization into textual reasoning for multimodal language models, eliminating the need for explicit visual grounding during inference while improving fine-grained perception performance.
View Cached Full Text
Cached at: 06/01/26, 11:20 AM
Paper page - iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Source: https://huggingface.co/papers/2605.31096
Abstract
A reinforcement learning framework called iVGR is introduced to transfer visual localization capabilities into textual reasoning, improving fine-grained perception in multimodal language models without requiring explicit visual grounding during inference.
While visually groundedChain-of-Thought(CoT) has emerged as a promising paradigm to enhancefine-grained perceptioninmultimodal large language models(MLLMs), its efficacy during the inference phase remains underexplored. In this work, we empirically find that mandating explicit object boxes in visually grounded CoT during inference often degrades performance compared to standard textual CoT, which reasons without explicit visual grounding. We hypothesize that the visual localization capability can be internalized into the textual CoT and that the mandatory explicit grounding introduces unnecessary interference with the model’s primary objective of answer prediction. To address this problem, we propose InternalizingVisually Grounded Reasoning(iVGR), a novelreinforcement learningframework that transfers localization capabilities into the textual reasoning process. We employ adual-stream trainingstrategy, where a textual stream is aligned with a high-quality visually grounded stream via a proposedconsistency reward, enabling the model to localize accurately without explicit grounding during inference. Extensive experiments demonstrate that our method significantly outperforms existing baselines on fine-grained benchmarks, while maintaining the flexibility to support tool-assisted inference workflows.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2605\.31096
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.31096 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.31096 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.31096 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning
This paper proposes PIVOT, a dual-level learning framework that enhances visually-grounded reasoning in large vision-language models by using self-calibrated experience replay and vision-guided advantage allocation to optimize reinforcement learning.
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
This paper introduces RIS, a framework for spatial-semantic grounded latent visual reasoning in Multimodal Large Language Models to overcome information bottlenecks. It proposes anchoring latent tokens to spatial and semantic evidence, showing improvements on benchmarks like V* and HRBench.
Region-Level Policy Optimization for Fine-grained MLLM Perception
Vision-RL² is a method for improving fine-grained perception in multimodal large language models by using region-level reinforcement learning to compress visual tokens and enhance performance across multiple benchmarks without full model fine-tuning.
Thinking with Visual Grounding
This paper introduces visually grounded thinking, a method for vision-language models to interleave natural-language reasoning with explicit visual evidence grounding using points or boxes. A scalable synthesis pipeline and grounding-aware reinforcement learning improve reasoning accuracy, enabling a 4B model to match or surpass a 27B model on spatial and counting benchmarks.
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
This paper introduces SR-REAL, a unified framework for spatial vision-language models that combines linguistic deduction and 3D geometric reasoning via reinforcement learning, enabling robust multi-step spatial reasoning across diverse tasks.