Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

arXiv cs.AI Papers

Summary

BRAID is a framework that formulates interleaved text-image-text reasoning as a unified Markov decision process, enabling joint optimization of textual and visual generation via reinforcement learning with a VLM judge providing dense turn-level feedback.

arXiv:2607.03748v1 Announce Type: new Abstract: Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegating image generation to supervised surrogates, preventing policy gradients from propagating through the full interleaved trajectory across heterogeneous modalities. This leaves the potential of RL for UMMs largely untapped. In the paper, we introduce \textbf{BRAID} (\textbf{B}ridging inte\textbf{R}le\textbf{A}ved mult\textbf{I}-modal reasoning as a unified \textbf{D}ecision process), a simple framework that casts multi-turn text-image-text reasoning as a unified Markov decision process (MDP), enabling joint optimization of textual and visual generation via a single, principled RL objective. BRAID computes a shared trajectory-level advantage and propagates it coherently into both text tokens and image denoising paths, each optimized through its modality-native policy gradient mechanism. To further address long-horizon credit assignment, BRAID employs a vision-language model (VLM) judge that scores each intermediate image on its reasoning utility, supplying dense turn-level feedback to sharpen learning at critical visual branches. Experiments on spatial reasoning and visual perception benchmarks show that BRAID consistently outperforms various baselines, confirming that a unified MDP formulation with vision-thinking guidance is essential for effective multi-modal reasoning.
Original Article
View Cached Full Text

Cached at: 07/07/26, 04:35 AM

# Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
Source: [https://arxiv.org/html/2607.03748](https://arxiv.org/html/2607.03748)
Zican Hu12∗\\astXuyang Hu3∗\\astYiming Liu4Zuwei Long2Wei Liu2 Yunzhuo Hao5Jiawei Gu6Linjie Li6Yu Cheng7Zhenhong Sun8 Weibo Gu2Xing Sun2Zhi Wang1🖂 1Nanjing University2Tencent Youtu Lab3Shanghai AI Laboratory 4Tsinghua University5Zhejiang University6University of Washington 7The Chinese University of Hong Kong8Australian National University Contact:zicanhu@smail\.nju\.edu\.cnzhiwang@nju\.edu\.cn

###### Abstract

Unified multi\-modal models \(UMMs\) have shown promising interleaved text\-image reasoning capabilities, yet effectively optimizing such multi\-turn generation via reinforcement learning \(RL\) remains an open challenge\. Existing approaches apply RL exclusively to text steps, relegating image generation to supervised surrogates, preventing policy gradients from propagating through the full interleaved trajectory across heterogeneous modalities\. This leaves the potential of RL for UMMs largely untapped\. In the paper, we introduceBRAID\(Bridging inteRleAved multI\-modal reasoning as a unifiedDecision process\), a simple framework that casts multi\-turn text\-image\-text reasoning as a unified Markov decision process \(MDP\), enabling joint optimization of textual and visual generation via a single, principled RL objective\. BRAID computes a shared trajectory\-level advantage and propagates it coherently into both text tokens and image denoising paths, each optimized through its modality\-native policy gradient mechanism\. To further address long\-horizon credit assignment, BRAID employs a vision\-language model \(VLM\) judge that scores each intermediate image on its reasoning utility, supplying dense turn\-level feedback to sharpen learning at critical visual branches\. Experiments on spatial reasoning and visual perception benchmarks show that BRAID consistently outperforms various baselines, confirming that a unified MDP formulation with vision\-thinking guidance is essential for effective multi\-modal reasoning\.

††footnotetext:∗Equal contributions\. This work was conducted during internship at Tencent Youtu Lab\.🖂\{\}^\{\\textrm\{\\Letter\}\}Corresponding authors\.### 1Introduction

![Refer to caption](https://arxiv.org/html/2607.03748v1/x1.png)Figure 1:SFT, Half\-optimized RL, and BRAID\.Recent emergence of unified multi\-modal models \(UMMs\)\(Chameleon Team,[2024](https://arxiv.org/html/2607.03748#bib.bib1); Denget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib4)\)that seamlessly bridge understanding and generation across modalities within a single framework has opened a new frontier in multi\-modal reasoning: interleaved Text\-Image\-Text Chain\-of\-Though \(TIT\-CoT\), where verbal thinking alternates with intermediate image generation to externalize perceptual steps beyond the reach of language\(Gaoet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib19); Wanget al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib2); Cuiet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib3)\)\. The underlying premise, often termed thevisual superiority hypothesis\(Wuet al\.,[2026](https://arxiv.org/html/2607.03748#bib.bib20)\), is that for physically grounded tasks, visual generation affords a more faithful world model than verbal reasoning, whose expressiveness is inherently bounded by the symbolic nature of language\. So far, this premise has been pursued exclusively through supervised fine\-tuning \(SFT\) on curated TIT\-CoT traces\(Guet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib17); Chernet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib18)\), an approach bottlenecked by the cost and scarcity of high\-quality interleaved data\. Reinforcement learning \(RL\) offers a principled alternative: by learning from task rewards rather than curated data, it enables models to autonomously discover novel reasoning strategies that go beyond what static data can teach, as already shown in language\(Guoet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib21); Shaoet al\.,[2024b](https://arxiv.org/html/2607.03748#bib.bib22)\)and multi\-modal reasoning\(Menget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib23); Yanget al\.,[2025b](https://arxiv.org/html/2607.03748#bib.bib24)\)\.

However, existing formulations leave TIT\-CoT half\-optimized: policy gradients flow through text tokens alone, while intermediate image generation remains outside the training loop, undermining the very advantage on which interleaved reasoning is premised\. COOPERZhanget al\.\([2025](https://arxiv.org/html/2607.03748#bib.bib27)\)optimizes visual generation, but through a separate supervised stage\. While DeepEyesZhenget al\.\([2025b](https://arxiv.org/html/2607.03748#bib.bib26)\)is trained end\-to\-end with RL, its policy gradients reach only the visual tool calling rather than the image generation process, still leaving visual reasoning outside the optimization loop\. Meantime, RL for image generation is well established: DDPO\(Blacket al\.,[2023](https://arxiv.org/html/2607.03748#bib.bib31)\)and DPOK\(Fanet al\.,[2023](https://arxiv.org/html/2607.03748#bib.bib32)\)cast diffusion denoising as an MDP, and subsequent work extends this to flow matching in FlowGRPO\(Liuet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib33)\)and DiffusionNFT\(Zhenget al\.,[2025a](https://arxiv.org/html/2607.03748#bib.bib34)\)\. These formulations provide a ready\-made RL toolkit for the visual branch of a unified model, yet none has been integrated into an interleaved reasoning setting\. The concurrent work, UniGRPO\(Liuet al\.,[2026a](https://arxiv.org/html/2607.03748#bib.bib36)\), applies RL to both modalities, but limited to single\-turn text\-to\-image generation\. These observations point to a critical missing piece:a unified RL framework that jointly optimizes textual and visual reasoning within an interleaved TIT\-CoT trajectory, enabling end\-to\-end credit assignment across modality boundaries and exploration over the full combinatorial space of interleaved strategies\.

To close this gap, we proposeBRAID\(Bridging inteRleAved multI\-modal reasoning as a unifiedDecision process\), which casts multi\-turn text\-image\-text reasoning as a unified MDP\. This formulation allows both textual and visual generation turns to be optimized jointly under one RL objective, without requiring separate training stages or modality\-specific reward designs\. Built on BAGELDenget al\.\([2025](https://arxiv.org/html/2607.03748#bib.bib4)\), BRAID formulates every turn in a\{text→image→text→⋯→answer\}\\\{\\texttt\{text\}\\\!\\to\\\!\\texttt\{image\}\\\!\\to\\\!\\texttt\{text\}\\\!\\to\\\!\\cdots\\\!\\to\\\!\\texttt\{answer\}\\\}trajectory as a macro\-action\. However, unifying the two modalities under a single objective is non\-trivial: the text branch yields per\-token autoregressive likelihoods while the image branch traverses continuous denoising paths via flow matching, and naively combining them collapses the optimization due to incompatible scales\. BRAID addresses the modality boundary by deriving a shared trajectory\-level advantage and channeling it intomodality\-native policy gradients: autoregressive token generation for text and a likelihood\-free flow\-matching objective for image\. Further, BRAID introduces a vision\-thinking process reward to address credit assignment over the long\-horizon TIT\-CoT rollout: a VLM judge scores each intermediate image on its utility to the reasoning trajectory, providing turn\-level feedback to sharpen learning at critical visual branches\. We hope BRAID serves as a step toward simple, general\-purpose RL frameworks for UMMs, and inspires further investigation into principled, modality\-agnostic policy optimization paradigms\.

Experiments show that BRAID improves BAGEL by\+5\.73avg\. across seven benchmarks, surpassing GPT\-4o with only 7B parameters, with pronounced gains on SAT \(\+14\.00\) and V\*Bench \(\+10\.76\)\. Notably, Maj@nnshows BRAID’s solution space scales efficiently with sampling budget while baselines plateau, and case studies reveal that the optimized image branch produces faithful intermediates and also transfers strong instruction\-following capability to text\-to\-image generation\.

### 2Related Work

Interleaved Multi\-Modal Reasoning\.Chain\-of\-Thought prompting\(Weiet al\.,[2022](https://arxiv.org/html/2607.03748#bib.bib9); Kojimaet al\.,[2022](https://arxiv.org/html/2607.03748#bib.bib10)\)decomposes complex reasoning into intermediate steps, yet pure text limits expressiveness for visually grounded tasks where spatial/geometric relations are more naturally captured by images\(Huet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib11); Menonet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib12)\)\. Recent unified architectures address this gap by integrating generation and understanding within a single backbone, enabling interleaved text\-image reasoning traces with intermediate visual states\(Menget al\.,[2023](https://arxiv.org/html/2607.03748#bib.bib16); Chenget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib53)\)\. These architectures follow two designs: fully autoregressive \(AR\) models that treat both modalities as discrete tokens under a unified next\-token objective\(Chameleon Team,[2024](https://arxiv.org/html/2607.03748#bib.bib1); Xieet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib7)\), and hybrid AR\-diffusion models that employ selective activation of modality\-specific parameters \(AR for text, diffusion for images\)\(Denget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib4); Liet al\.,[2025a](https://arxiv.org/html/2607.03748#bib.bib14)\)\. Both lines demonstrate that cross\-modal reasoning can be made intrinsic to a single model through dedicated objectives and curated interleaved data\(Guet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib17); Liet al\.,[2025b](https://arxiv.org/html/2607.03748#bib.bib15); Tonget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib13)\)\. BRAID builds on the hybrid backbone for its superiority in delegating each modality to its best\-suited generative mechanism\(Wuet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib5); Zhouet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib8); Donget al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib55)\)\.

RL Fine\-tuning\.Verifiable\-reward RL has emerged as a principled paradigm for enhancing reasoning beyond the ceiling of supervised data, first in LLMs\(Guoet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib21); Shaoet al\.,[2024b](https://arxiv.org/html/2607.03748#bib.bib22)\)and subsequently in VLMs that are still limited to textual action space\(Menget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib23); Huanget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib25)\)\. A parallel line adapts RL to image generation by casting diffusion denoising as an MDP\(Blacket al\.,[2023](https://arxiv.org/html/2607.03748#bib.bib31); Fanet al\.,[2023](https://arxiv.org/html/2607.03748#bib.bib32); Liuet al\.,[2026b](https://arxiv.org/html/2607.03748#bib.bib54)\), with subsequent work generalizing to flow\-matching models in FlowGRPO\(Liuet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib33)\), accelerating learning through direct forward\-process optimization in DiffusionNFT\(Zhenget al\.,[2025a](https://arxiv.org/html/2607.03748#bib.bib34)\), and unifying these variants under a single policy\-gradient framework in DanceGRPO\(Xueet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib35)\)\. However, these two lines of work operate within single\-modality generation, and neither accommodates the setting where a unified model emits interleaved text\-image outputs\. BRAID aims to bridge this divide by formulating RL over the joint multi\-modal reasoning trajectory\.

RL for Interleaved Multi\-Modal Reasoning\.A growing body of work applies RL to fine\-tune unified models\. For AR backbones, Emu3\.5\(Cuiet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib3)\)scales post\-training with verifiable\-reward RL, and other studies adopt the same paradigm with a hybrid reward design\(Genget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib28); Nieet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib56)\)\. For hybrid architectures, prior work typically excludes the visual branch from the RL loop: COOPER\(Zhanget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib27)\)trains image generation in a separate supervised stage, and DeepEyes\(Zhenget al\.,[2025b](https://arxiv.org/html/2607.03748#bib.bib26)\)backpropagates policy gradients to the visual tool calling rather than the image generation process itself\. Further, the concurrent work UniGRPO\(Liuet al\.,[2026a](https://arxiv.org/html/2607.03748#bib.bib36)\)does jointly optimize both modalities under a single RL objective, but is limited to single\-turn text\-to\-image generation\. In summary, no existing pipeline applies RL to the multi\-turn interleaved reasoning setting where each text/image turn conditions subsequent generation\. BRAID fills this gap by casting the full interleaved trajectory as a unified decision process and deriving modality\-native gradients end\-to\-end within a single policy\.

### 3Preliminaries

##### Group Relative Policy Optimization\.

GRPOShaoet al\.\([2024b](https://arxiv.org/html/2607.03748#bib.bib22)\)samples, for each promptqq, a group ofGGrollouts\{τi\}i=1G\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{G\}from the behavior policyπθold\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}, scores eachτi\\tau\_\{i\}with a scalar rewardr​\(τi\)r\(\\tau\_\{i\}\), and assigns a group\-normalized advantage asA^​\(τi\)=r​\(τi\)−μ​\(r​\(τ\)\)σ​\(r​\(τ\)\)\+κ\\hat\{A\}\(\\tau\_\{i\}\)=\\frac\{r\(\\tau\_\{i\}\)\-\\mu\(r\(\\tau\)\)\}\{\\sigma\(r\(\\tau\)\)\+\\kappa\}, whereμ​\(⋅\)\\mu\(\\cdot\)andσ​\(⋅\)\\sigma\(\\cdot\)denote the mean and standard derivation, andκ\\kappais a small constant for numerical stability\. The policyπθ\\pi\_\{\\theta\}is then updated by a clipped surrogate objective with a KL penalty against a frozen reference modelπref\\pi\_\{\\mathrm\{ref\}\}as

𝒥GRPO​\(πθ\)=𝔼​\[1G​∑i=1G1\|τi\|​∑t=1\|τi\|min⁡\(ρi,t​A^i,clip​\(ρi,t,1−ε,1\+ε\)​A^i\)\]−η​DKL​\[πθ∥πref\],\\mathcal\{J\}\_\{\\text\{GRPO\}\}\(\\pi\_\{\\theta\}\)=\\mathbb\{E\}\\\!\\Biggl\[\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\frac\{1\}\{\|\\tau\_\{i\}\|\}\\sum\_\{t=1\}^\{\|\\tau\_\{i\}\|\}\\min\\bigl\(\\rho\_\{i,t\}\\,\\hat\{A\}\_\{i\},\\;\\mathrm\{clip\}\(\\rho\_\{i,t\},1\\\!\-\\\!\\varepsilon,1\\\!\+\\\!\\varepsilon\)\\,\\hat\{A\}\_\{i\}\\bigr\)\\Biggr\]\-\\eta D\_\{\\mathrm\{KL\}\}\\\!\\bigl\[\\pi\_\{\\theta\}\\\|\\pi\_\{\\mathrm\{ref\}\}\\bigr\],\(1\)where\|τi\|\|\\tau\_\{i\}\|is the length of rolloutτi\\tau\_\{i\},ρi,t=πθ​\(τi,t\|q,τi,<t\)/πθold​\(τi,t\|q,τi,<t\)\\rho\_\{i,t\}=\\pi\_\{\\theta\}\(\\tau\_\{i,t\}\|q,\\tau\_\{i,<t\}\)/\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\tau\_\{i,t\}\|q,\\tau\_\{i,<t\}\)is the per\-token importance ratio,ε\\varepsilonis the clipping threshold, andη\\etacontrols the KL regularization strength\.

##### RL for Image generation\.

Image steps are modeled by flow matchingLipmanet al\.\([2022](https://arxiv.org/html/2607.03748#bib.bib30)\)\. Given contextcc, the policy parameterizes a conditional velocity fieldvθ​\(xt,c,t\)v\_\{\\theta\}\(x\_\{t\},c,t\)that transports a noisy sampleϵ∼𝒩​\(0,I\)\\epsilon\\sim\\mathcal\{N\}\(0,I\)to a clean imagex0x\_\{0\}via the probability\-flow ODE along the continuous time axist∈\[0,1\]t\\in\[0,1\]\. On interpolated samplesxt=\(1−t\)​x0\+t​ϵx\_\{t\}=\(1\-t\)x\_\{0\}\+t\\epsilon, flow matching regressesvθv\_\{\\theta\}onto the ground\-truth velocity:

ℒFM​\(θ\)=𝔼t,x0,x1,c​‖vθ​\(xt,c,t\)−\(ϵ−x0\)‖2\.\\mathcal\{L\}\_\{\\text\{FM\}\}\(\\theta\)=\\mathbb\{E\}\_\{t,x\_\{0\},x\_\{1\},c\}\\bigl\\lVert v\_\{\\theta\}\(x\_\{t\},c,t\)\-\(\\epsilon\-x\_\{0\}\)\\bigr\\rVert^\{2\}\.\(2\)
Applying RL to flow\-matching generators requires injecting reward signals into this objective\. Rather than estimating the intractable likelihoodlog⁡pθ​\(x0\|c\)\\log p\_\{\\theta\}\(x\_\{0\}\|c\), DiffusionNFTZhenget al\.\([2025a](https://arxiv.org/html/2607.03748#bib.bib34)\)turns Eq\. \([2](https://arxiv.org/html/2607.03748#S3.E2)\) into a reward\-aware regression by splitting each sample into a positive and a negative branch weighted by a soft labelr​\(x0,c\)∈\[0,1\]r\(x\_\{0\},c\)\\in\[0,1\], i\.e\., the reward representing the optimality probability:

ℒNFT​\(θ\)=𝔼​\[r​\(x0,c\)​‖vθ\+​\(xt,c,t\)−\(ϵ−x0\)‖2\+\(1−r​\(x0,c\)\)​‖vθ−​\(xt,c,t\)−\(ϵ−x0\)‖2\],\\mathcal\{L\}\_\{\\text\{NFT\}\}\(\\theta\)=\\mathbb\{E\}\\Bigl\[\\,r\(x\_\{0\},c\)\\,\\bigl\\lVert v^\{\+\}\_\{\\theta\}\(x\_\{t\},c,t\)\-\(\\epsilon\-x\_\{0\}\)\\bigr\\rVert^\{2\}\+\(1\-r\(x\_\{0\},c\)\)\\,\\bigl\\lVert v^\{\-\}\_\{\\theta\}\(x\_\{t\},c,t\)\-\(\\epsilon\-x\_\{0\}\)\\bigr\\rVert^\{2\}\\Bigr\],\(3\)with implicit predictorsvθ±:=\(1∓β\)​vθold±β​vθv^\{\\pm\}\_\{\\theta\}:=\(1\\mp\\beta\)\\,v\_\{\\theta\_\{\\mathrm\{old\}\}\}\\pm\\beta\\,v\_\{\\theta\}that linearly mix the current policyvθv\_\{\\theta\}with a frozen behavior velocityvθoldv\_\{\\theta\_\{\\mathrm\{old\}\}\}, whereβ\\betais a hyperparameter\. High\-reward samples pullvθv\_\{\\theta\}toward the data target, while low\-reward ones push it away, injecting the RL signal directly into the forward training loss without any likelihood or score estimation\.

### 4Method

![Refer to caption](https://arxiv.org/html/2607.03748v1/x2.png)Figure 2:Overview of BRAID\. \(a\) Interleaved TIT\-CoT reasoning formulated as a two\-level MDP\. \(b\) Unified policy optimization via GRPO for text and DiffusionNFT for image\. \(c\) Decoupled advantage estimation combining terminal reward and vision\-thinking process reward\.Here, we first cast TIT\-CoT reasoning as a two\-level MDP that unifies text and image generation in \(§[4\.2](https://arxiv.org/html/2607.03748#S4.SS2)\)\. Then, we introduce the unified optimization mechanism that channels the shared trajectory\-level advantage into modality\-native policy gradients in \(§[4\.3](https://arxiv.org/html/2607.03748#S4.SS3)\)\. Finally, we present the vision\-thinking process reward design that enables efficient credit assignment for image turns in \(§[4\.4](https://arxiv.org/html/2607.03748#S4.SS4)\)\.

#### 4\.1Problem Formulation

We formulate a UMM policyπθ\\pi\_\{\\theta\}that unifies text and image generation within a single modelDenget al\.\([2025](https://arxiv.org/html/2607.03748#bib.bib4)\)\. Given a promptc=\(q,\{ℐqn\}i=1M\)c=\(q,\\\{\\mathcal\{I\}^\{n\}\_\{q\}\\\}\_\{i=1\}^\{M\}\)pairing a textual queryqqwithM≥1M\\\!\\geq\\\!1images, the model produces an*interleaved*reasoning trajectoryτ∼πθ\(⋅\|c\)\\tau\\sim\\pi\_\{\\theta\}\(\\cdot\\,\|\\,c\)as

τ=\(T1,I1,T2,I2,…,TK,IK,Tans\),\\tau=\\bigl\(T\_\{1\},\\,I\_\{1\},\\,T\_\{2\},\\,I\_\{2\},\\,\\ldots,\\,T\_\{K\},\\,I\_\{K\},\\,T\_\{\\mathrm\{ans\}\}\\bigr\),\(4\)which alternates between text chunksTkT\_\{k\}and intermediate imagesIkI\_\{k\}across multiple turns, and ends with a textual answerTansT\_\{\\mathrm\{ans\}\}\. Both modalities are generated fromπθ\\pi\_\{\\theta\}using a shared multi\-modal attention backboneGuet al\.\([2025](https://arxiv.org/html/2607.03748#bib.bib17)\)\. We denote this structure as*Text\-Image\-Text Chain\-of\-Thought*\(TIT\-CoT\)\.

#### 4\.2Interleaved Reasoning as a Unified MDP

To optimize text and image turns under a joint RL framework, we formulate the generation of a TIT\-CoT trajectory as a two\-level MDP asℳ=\(ℳturn,ℳstep\)\\mathcal\{M\}=\(\\mathcal\{M\}\_\{\\mathrm\{turn\}\},\\mathcal\{M\}\_\{\\mathrm\{step\}\}\)\. The outer levelℳturn\\mathcal\{M\}\_\{\\mathrm\{turn\}\}abstracts each reasoning turn as a*macro\-action*that emits either a text chunkTkT\_\{k\}or an intermediate imageIkI\_\{k\}\. The inner levelℳstep\\mathcal\{M\}\_\{\\mathrm\{step\}\}unrolls each macro\-action into a sequence of*micro\-actions*, i\.e\., tokens for text generation and denoising paths for image generation, where the turn\-level advantage can be consistently propagated to every micro\-action within that turn\.

Turn\-level MDP\.The interleaved reasoning over2​K\+12K\{\+\}1turns is mapped to the following MDP:

sk≜\(c,τ<k\)\\displaystyle s\_\{k\}\\triangleq\(c,\\,\\tau\_\{<k\}\)\\qquadπ​\(ak∣sk\)≜πθ​\(ak∣c,τ<k\)\\displaystyle\\pi\(a\_\{k\}\\mid s\_\{k\}\)\\triangleq\\pi\_\{\\theta\}\(a\_\{k\}\\mid c,\\,\\tau\_\{<k\}\)\\qquadP​\(sk\+1∣sk,ak\)≜δsk⊕ak\\displaystyle P\(s\_\{k\+1\}\\mid s\_\{k\},a\_\{k\}\)\\triangleq\\delta\_\{\\,s\_\{k\}\\oplus a\_\{k\}\}ak≜\{Tk∈𝒜txtIk∈𝒜img\\displaystyle a\_\{k\}\\triangleq\\begin\{cases\}T\_\{k\}\\in\\mathcal\{A\}\_\{\\text\{txt\}\}\\\\\[2\.0pt\] I\_\{k\}\\in\\mathcal\{A\}\_\{\\text\{img\}\}\\end\{cases\}\\qquadρ​\(s0\)≜δ\(c,∅\)\\displaystyle\\rho\(s\_\{0\}\)\\triangleq\\delta\_\{\(c,\\,\\varnothing\)\}\\qquadR​\(sk,ak\)≜\{r​\(τ\)if​k=2​K\+10otherwise\\displaystyle R\(s\_\{k\},a\_\{k\}\)\\triangleq\\begin\{cases\}r\(\\tau\)&\\text\{if \}k\\\!=\\\!2K\\\!\+\\\!1\\\\\[2\.0pt\] 0&\\text\{otherwise\}\\end\{cases\}The statesks\_\{k\}concatenates the promptccwith all preceding turnsτ<k\\tau\_\{<k\}, andρ​\(⋅\)\\rho\(\\cdot\)is the initial state distribution\. The actionaka\_\{k\}is either a text chunkTkT\_\{k\}or an imageIkI\_\{k\}, sampled from the policyπθ​\(ak\|sk\)\\pi\_\{\\theta\}\(a\_\{k\}\|s\_\{k\}\)whose text and image experts share a unified backbone\. The transitionPPis the deterministic appendsk\+1=sk⊕aks\_\{k\+1\}\\\!=\\\!s\_\{k\}\\\!\\oplus\\\!a\_\{k\}, the rewardRRis a sparse rule\-based signal assigned only at the final answerTansT\_\{\\mathrm\{ans\}\}\. By collapsing each text chunk or intermediate image into a single macro\-action, this abstraction renders the two modalitieshomogeneousat the outer level: text and image turns are seamlessly integrated into a single MDP that shares a unified state space, transition dynamics, and reward signal, thereby enabling a coherent RL formulation over the entire interleaved reasoning trajectory\.

Step\-level MDP\.Each turn\-level actionaka\_\{k\}is realized by rolling out a lower\-level MDPℳstep\\mathcal\{M\}\_\{\\mathrm\{step\}\}that operates at a fine\-grained scale\. The specific form depends on the modality:

- •For a text turnak=Tka\_\{k\}\\\!=\\\!T\_\{k\},ℳstep\\mathcal\{M\}\_\{\\mathrm\{step\}\}is the standard autoregressive token generation: at stepii, the micro\-state is the running prefixski:=\(sk,ak<i\)s\_\{k\}^\{i\}:=\\left\(s\_\{k\},a\_\{k\}^\{<i\}\\right\), the micro\-action is the next tokenaki:=Tkia\_\{k\}^\{i\}:=T\_\{k\}^\{i\}, andπθ\\pi\_\{\\theta\}is factorized asπθ​\(Tk\|sk\)=∏i=1\|Tk\|πθ​\(Tki∣sk,Tk<i\)\\pi\_\{\\theta\}\(T\_\{k\}\|s\_\{k\}\)=\\prod\_\{i=1\}^\{\|T\_\{k\}\|\}\\pi\_\{\\theta\}\\left\(T\_\{k\}^\{i\}\\mid s\_\{k\},T\_\{k\}^\{<i\}\\right\), where\|Tk\|\|T\_\{k\}\|is the text length\.
- •For an image turnak=Ika\_\{k\}\\\!=\\\!I\_\{k\},ℳstep\\mathcal\{M\}\_\{\\mathrm\{step\}\}is the iterative denoising procedure via flow matching: at denoising timestept∈\[0,1\]t\\\!\\in\\\!\[0,1\], the micro\-state is the noisy imageskt:=xts\_\{k\}^\{t\}:=x\_\{t\}, the micro\-action is the predicted flow velocityakt:=vθ​\(xt,sk,t\)a\_\{k\}^\{t\}:=v\_\{\\theta\}\(x\_\{t\},s\_\{k\},t\), and the state transition follows the probability\-flow ODE asxtn−1=xtn\+vθ​\(xtn,sk,tn\)⋅d​tx\_\{t\_\{n\-1\}\}\\\!=\\\!x\_\{t\_\{n\}\}\+v\_\{\\theta\}\(x\_\{t\_\{n\}\},s\_\{k\},t\_\{n\}\)\\cdot\\mathrm\{d\}t, with the clean imageIk=x0I\_\{k\}=x\_\{0\}as the terminal state, wheren=N,N−1,…,1n=N,N\\\!\-\\\!1,\.\.\.,1is the denoising step along the time axis\.

Let𝒦txt\\mathcal\{K\}\_\{\\mathrm\{txt\}\}and𝒦img\\mathcal\{K\}\_\{\\mathrm\{img\}\}index text and image turns, respectively\. Because all turns reside in a single MDP, the joint policy over a complete trajectory factorizes naturally along the turn sequence:

log⁡πθ​\(τ∣c\)=∑k=12​K\+1log⁡πθ​\(ak∣sk\)\.\\log\\pi\_\{\\theta\}\(\\tau\\mid c\)\\;=\\;\\sum\\nolimits\_\{k=1\}^\{2K\+1\}\\log\\pi\_\{\\theta\}\(a\_\{k\}\\mid s\_\{k\}\)\\,\.\(5\)Expanding each macro\-action into its step\-level form then reveals twomodality\-nativebranches:

log⁡πθ​\(τ∣c\)=∑k∈𝒦txt∑i=1\|Tk\|log⁡πθ​\(Tki∣sk,Tk<i\)⏟text branch\+∑k∈𝒦img∑n=1Nlog⁡πθ​\(xtn−1k∣xtnk,sk\)⏟image branch \(flow\-matching\)\.\\log\\pi\_\{\\theta\}\(\\tau\\mid c\)\\;=\\;\\underbrace\{\\sum\_\{k\\in\\mathcal\{K\}\_\{\\mathrm\{txt\}\}\}\\sum\_\{i=1\}^\{\|T\_\{k\}\|\}\\log\\pi\_\{\\theta\}\\bigl\(T\_\{k\}^\{i\}\\mid s\_\{k\},T\_\{k\}^\{<i\}\\bigr\)\}\_\{\\text\{text branch\}\}\\;\+\\;\\underbrace\{\\sum\_\{k\\in\\mathcal\{K\}\_\{\\mathrm\{img\}\}\}\\sum\_\{n=1\}^\{N\}\\log\\pi\_\{\\theta\}\\bigl\(x^\{k\}\_\{t\_\{n\-1\}\}\\mid x^\{k\}\_\{t\_\{n\}\},\\,s\_\{k\}\\bigr\)\}\_\{\\text\{image branch \(flow\-matching\)\}\}\.\(6\)Although the two branches differ in computational mechanism \(autoregressive token generation for text and flow\-matching denoising for image generation\), their structural parallelism ensures that a shared trajectory\-level advantage can be consistently back\-propagated through both modality\-native policy gradients, a property we exploit in the unified RL objective in §[4\.3](https://arxiv.org/html/2607.03748#S4.SS3)\.

#### 4\.3Unified Policy Optimization

We train the unified policyπθ=\(πθtxt,πθimg\)\\pi\_\{\\theta\}=\(\\pi\_\{\\theta\}^\{\\text\{txt\}\},\\,\\pi\_\{\\theta\}^\{\\text\{img\}\}\)by jointly optimizing both branches on grouped rollouts, leveraging the structural parallelism established in Eq\. \([6](https://arxiv.org/html/2607.03748#S4.E6)\): a shared turn\-level advantageA^​\(sk,ak\)\\hat\{A\}\(s\_\{k\},a\_\{k\}\)is back\-propagated through each modality\-native surrogate, a GRPO\-style clipped objective for the text branch and a DiffusionNFT loss\(Zhenget al\.,[2025a](https://arxiv.org/html/2607.03748#bib.bib34)\)for the image branch, yielding the unified objective:

𝒥​\(θ\)=∑k∈𝒦txt𝒥GRPO​\(πθtxt​\(Tk\|sk\);A^​\(sk,Tk\)\)−λ⋅∑k∈𝒦imgℒNFT​\(πθimg​\(Ik\|sk\);A^​\(sk,Ik\)\),\\mathcal\{J\}\(\\theta\)\\\!=\\\!\\sum\_\{k\\in\\mathcal\{K\}\_\{\\mathrm\{txt\}\}\}\\\!\\\!\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\\bigl\(\\pi^\{\\mathrm\{txt\}\}\_\{\\theta\}\(T\_\{k\}\|s\_\{k\}\);\\hat\{A\}\(s\_\{k\},T\_\{k\}\)\\bigr\)\\;\-\\;\\lambda\\cdot\\\!\\\!\\sum\_\{k\\in\\mathcal\{K\}\_\{\\mathrm\{img\}\}\}\\\!\\mathcal\{L\}\_\{\\mathrm\{NFT\}\}\\bigl\(\\pi^\{\\mathrm\{img\}\}\_\{\\theta\}\(I\_\{k\}\|s\_\{k\}\);\\hat\{A\}\(s\_\{k\},I\_\{k\}\)\\bigr\),\(7\)whereλ\\lambdabalances the two modalities\. Eqs\. \([5](https://arxiv.org/html/2607.03748#S4.E5)\)\-\([7](https://arxiv.org/html/2607.03748#S4.E7)\) together realize BRAID’s core principle: despite distinct generative mechanisms, text and image branches are optimized within a unified RL loop, sharing one interleaved trajectory, one turn\-level advantage, and one joint backward pass\.

To injectA^\\hat\{A\}into flow matching without likelihood estimation, we map it to a soft reward labelrk=σ​\(A^​\(sk,Ik\)/υ\)∈\[0,1\]r\_\{k\}=\\sigma\(\\hat\{A\}\(s\_\{k\},I\_\{k\}\)/\\upsilon\)\\in\[0,1\]with temperatureυ\\upsilon, and plug it into the DiffusionNFT loss of Eq\. \([3](https://arxiv.org/html/2607.03748#S3.E3)\):

ℒNFT​\(θ\)=𝔼​\[rk​‖vθ\+​\(xt,sk,t\)−\(ϵ−x0\)‖2\+\(1−rk\)​‖vθ−​\(xt,sk,t\)−\(ϵ−x0\)‖2\],\\mathcal\{L\}\_\{\\mathrm\{NFT\}\}\(\\theta\)=\\mathbb\{E\}\\Bigl\[r\_\{k\}\\,\\bigl\\lVert v^\{\+\}\_\{\\theta\}\\left\(x\_\{t\},s\_\{k\},t\\right\)\-\(\\epsilon\-x\_\{0\}\)\\bigr\\rVert^\{2\}\+\(1\-r\_\{k\}\)\\,\\bigl\\lVert v^\{\-\}\_\{\\theta\}\\left\(x\_\{t\},s\_\{k\},t\\right\)\-\(\\epsilon\-x\_\{0\}\)\\bigr\\rVert^\{2\}\\Bigr\],\(8\)wherex0=Ikx\_\{0\}=I\_\{k\}is the clean image,ϵ\\epsilonis Gaussian noise,xt=\(1−t\)​x0\+t​ϵx\_\{t\}=\(1\-t\)x\_\{0\}\+t\\epsilonis the interpolated sample, andvθ±=\(1∓β\)​vθold±β​vθv^\{\\pm\}\_\{\\theta\}=\(1\\\!\\mp\\\!\\beta\)v\_\{\\theta\_\{\\mathrm\{old\}\}\}\\pm\\beta v\_\{\\theta\}are implicit predictors as described in §[3](https://arxiv.org/html/2607.03748#S3)\. Intuitively, a highA^k\\hat\{A\}\_\{k\}steers the velocity fieldvθv\_\{\\theta\}toward the generated imageIkI\_\{k\}, while a lowA^k\\hat\{A\}\_\{k\}steers it away, thereby propagating the terminal reward signal to image generation directly via the flow\-matching loss\.

#### 4\.4Vision\-Thinking Process Reward

The trajectory\-level advantageA^​\(τ\)\\hat\{A\}\(\\tau\)broadcasts a single terminal\-reward scalar uniformly to every turn\. As the TIT\-CoT horizon grows, this undifferentiated signal increasingly obscures which intermediate steps truly matter for the final answer\. The issue is amplified for image turns: unlike autoregressive text, image generation must produce coherent, pixel\-level content in each non\-autoregressive pass, demanding a far sharper learning signal than a sparse, delayed reward can provide\.

To address this, we introduce a*vision\-thinking process reward*that scores each intermediate image on its utility to the reasoning trajectory, providing dense credit assignment for image turns\. Specifically, we employ a VLM judge111We use GPT\-5\.2 in all experiments; see Appendix[A\.4](https://arxiv.org/html/2607.03748#A1.SS4)for the full prompt\.to score each intermediate imageIkI\_\{k\}according to a set of complementary criteria𝒞\\mathcal\{C\}, yielding a per\-turn*vector*of process rewards:𝒓vis=\[r\(c\)\]c∈𝒞∈ℝ\|𝒞\|\\bm\{r\}^\{\\mathrm\{vis\}\}=\\bigl\[r^\{\(c\)\}\\bigr\]\_\{c\\in\\mathcal\{C\}\}\\in\\mathbb\{R\}^\{\|\\mathcal\{C\}\|\}\. In practice, we instantiate𝒞=\{vc,vf,ru,tr\}\\mathcal\{C\}=\\\{\\texttt\{vc\},\\texttt\{vf\},\\texttt\{ru\},\\texttt\{tr\}\\\}, covering four salient aspects of a visual thought:

- •Visual Correctnessr\(vc\)r^\{\(\\mathrm\{vc\}\)\}: whether the intended visual operation is executed accurately;
- •Visual Fidelityr\(vf\)r^\{\(\\mathrm\{vf\}\)\}: whether the image is coherent and free of hallucinated content;
- •Reasoning Utilityr\(ru\)r^\{\(\\mathrm\{ru\}\)\}: whether the image provides useful evidence to the final answer;
- •Trustworthinessr\(tr\)r^\{\(\\mathrm\{tr\}\)\}: whether the image is unlikely to mislead a reasonable reasoner\.

Decoupled Advantage Estimation\.Since the visual process reward is intrinsically multi\-objective, its components typically operate on disparate and dynamically shifting scales during training\. This complicates the development of a flexible weighting scheme for aggregating the process reward components\. To circumvent this issue, we pool all image turns within the group of rollouts into a visual setℐ𝒢\\mathcal\{I\}\_\{\\mathcal\{G\}\}, compute the advantages for each reward component independently, and simply merge them with equal weight at the advantage level:

A^vis​\(sk,Ik\)=∑c∈𝒞rk\(c\)−μℐ𝒢​\(r\(c\)\)σℐ𝒢​\(r\(c\)\),Ik∈ℐ𝒢,\\hat\{A\}^\{\\mathrm\{vis\}\}\(s\_\{k\},I\_\{k\}\)\\;=\\;\\sum\\nolimits\_\{c\\in\\mathcal\{C\}\}\\frac\{r^\{\(c\)\}\_\{k\}\-\\mu\_\{\\mathcal\{I\}\_\{\\mathcal\{G\}\}\}\(r^\{\(c\)\}\)\}\{\\sigma\_\{\\mathcal\{I\}\_\{\\mathcal\{G\}\}\}\(r^\{\(c\)\}\)\},\\qquad I\_\{k\}\\in\\mathcal\{I\}\_\{\\mathcal\{G\}\},\(9\)whereμℐ𝒢​\(⋅\)\\mu\_\{\\mathcal\{I\}\_\{\\mathcal\{G\}\}\}\(\\cdot\)andσℐ𝒢​\(⋅\)\\sigma\_\{\\mathcal\{I\}\_\{\\mathcal\{G\}\}\}\(\\cdot\)denote the mean and standard derivation calculated fromℐ𝒢\\mathcal\{I\}\_\{\\mathcal\{G\}\}\. Analogously, we incorporate the terminal rewardr​\(τ\)r\(\\tau\)with the visual process reward in a decoupled way as

A^​\(sk,ak\)=\{A^ans​\(sk,Tk\),for text turns,A^ans​\(sk,Ik\)\+λvis⋅A^vis​\(sk,Ik\),for image turns,\\hat\{A\}\(s\_\{k\},a\_\{k\}\)\\;=\\;\\begin\{cases\}\\hat\{A\}^\{\\mathrm\{ans\}\}\(s\_\{k\},T\_\{k\}\),&\\text\{for text turns\},\\\\\[2\.0pt\] \\hat\{A\}^\{\\mathrm\{ans\}\}\(s\_\{k\},I\_\{k\}\)\+\\lambda\_\{\\text\{vis\}\}\\cdot\\hat\{A\}^\{\\mathrm\{vis\}\}\(s\_\{k\},I\_\{k\}\),&\\text\{for image turns\},\\end\{cases\}\(10\)whereA^ans​\(⋅\)\\hat\{A\}^\{\\mathrm\{ans\}\}\(\\cdot\)is the advantage calculated from the terminal rewardr​\(τ\)r\(\\tau\), andλvis\\lambda\_\{\\text\{vis\}\}balances two advantages\. By decoupling advantage estimation, each reward signal is normalized relative to its own baseline before aggregation\. This architecture mitigates sensitivity to scale mismatch and temporal fluctuations, facilitating more robust and interpretable control over the process reward integration\.

### 5Experiments

Dataset and Evaluation\.Following the literature\(Guet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib17)\), our training corpus spans a wide range of eight task families:Matterport3D\(Changet al\.,[2017](https://arxiv.org/html/2607.03748#bib.bib37)\),RealSee3D\(Liet al\.,[2025c](https://arxiv.org/html/2607.03748#bib.bib38)\),Perspective,DirectionalQuery,360\+x360\{\+\}x\(Chenet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib41)\),PositionedDirection,Jigsaw\(Wanget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib43)\), andVisualSearch\(Shaoet al\.,[2024a](https://arxiv.org/html/2607.03748#bib.bib42)\)\. In each task, the dataset is partitioned into SFT and RL pools under three ratios:𝒟2:1\\mathcal\{D\}^\{2\{:\}1\},𝒟1:1\\mathcal\{D\}^\{1\{:\}1\}, and𝒟∩\\mathcal\{D\}^\{\\cap\}\. We default to the partial\-overlap regime𝒟∩\\mathcal\{D\}^\{\\cap\}, which bridges supervised initialization and on\-policy exploration\. We evaluate on seven benchmarks spanning two categories: 1\) spatial reasoning, includingMMSI\-Bench\(Yanget al\.,[2025a](https://arxiv.org/html/2607.03748#bib.bib47)\),SAT\(Rayet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib48)\),MMVP\(Tonget al\.,[2024b](https://arxiv.org/html/2607.03748#bib.bib45)\), andCV\-Bench 3D\(Tonget al\.,[2024a](https://arxiv.org/html/2607.03748#bib.bib49)\); 2\) visual perception, coveringBLINK\(Fuet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib46)\),V∗Bench\(Wu and Xie,[2024](https://arxiv.org/html/2607.03748#bib.bib44)\), andCV\-Bench 2D\(Tonget al\.,[2024a](https://arxiv.org/html/2607.03748#bib.bib49)\)\. All evaluations are conducted under theVLMEvalKitframework\(Duanet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib50)\)\. Appendix[A\.1](https://arxiv.org/html/2607.03748#A1.SS1)shows details of datasets and benchmarks\.

Baselines and Training Recipe\.We compare our method to two categories of baselines: 1\) proprietary and open\-source VLMs:GPT\-5\.4,GPT\-4o\(Hurstet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib51)\), andQwen2\.5\-VL\(Baiet al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib52)\); 2\) competitive UMMs:Janus\-Pro\(Wuet al\.,[2024](https://arxiv.org/html/2607.03748#bib.bib5)\),Chameleon\(Chameleon Team,[2024](https://arxiv.org/html/2607.03748#bib.bib1)\), andBAGEL\(Denget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib4)\)\. For RL training, we remove the KL loss term, and set the key hyperparameters: rollout batch size6464, update batch size3232,88rollouts per prompt\. For DiffusionNFT, we setβ=0\.4\\beta=0\.4and use2020sampling steps for training and evaluation\. All experimental details are documented in Appendix[A](https://arxiv.org/html/2607.03748#A1)\.

#### 5\.1Main Results

Table 1:Performance comparison on multimodal spatial reasoning and visual perception benchmarks\. Best results inboldand second bestunderlined\.Comparison among Unified Models\.Table[1](https://arxiv.org/html/2607.03748#S5.T1)summarizes the performance of all tested methods on all benchmarks\. Among UMMs of comparable scale, BAGEL already establishes a strong baseline \(59\.46 avg\.\), substantially outperforming Janus\-Pro \(46\.05\) and Chameleon \(27\.49\), which validates its hybrid AR\-diffusion architecture as a capable backbone for multi\-modal reasoning \(even from limited data\) and motivates our choice of building upon it\. Nevertheless, a notable gap remains between BAGEL and understanding\-only VLMs: even the same\-scale Qwen2\.5\-VL\-7B \(62\.11\) surpasses it by 2\.65 points, with the margin widening further against Qwen2\.5\-VL\-72B \(70\.53\)\. It suggests that the generative versatility of unified architectures does not automatically translate into strong visual reasoning without dedicated post\-training\.

Spatial Reasoning and Visual Perception Performance\.Effective on\-policy exploration presupposes a policy that can already generate reasonable rollouts\. Our SFT stage fulfills precisely this prerequisite: it improves over BAGEL by\+2\.53points on average, with gains observed on nearly every benchmark, thereby furnishing a reliable initialization for the subsequent RL phase\. From this initialization, RL yields an additional\+2\.73\-point average gain over BAGEL, lifting BRAID to65\.19and surpassing GPT\-4o \(62\.46\) despite employing only 7B parameters\. Notably, the most pronounced gains concentrate on SAT \(\+14\.00\) and V\*Bench \(\+10\.76\), which demand mental viewpoint transformation or fine\-grained visual discrimination – precisely the capabilities where unified models have historically trailed understanding\-only counterparts\. This confirms that BRAID can narrow such deficits without any architectural change\. The sole regression appears on CV\-Bench 3D \(\-1\.24\), which we attribute to a distributional gap in the SFT data; encouragingly, BRAID recovers part of this loss, reducing the deficit from \-1\.74 \(SFT\) to \-1\.24\.

![Refer to caption](https://arxiv.org/html/2607.03748v1/x3.png)Figure 3:Maj@nnand Avg@nncurves on VStar and SAT benchmarks with varying number of samplesn=\{3,5,9,17\}n=\\\{3,5,9,17\\\}\. Complete results are available in Table[5](https://arxiv.org/html/2607.03748#A2.T5)\.Reasoning Capacity Boundary\.Figure[3](https://arxiv.org/html/2607.03748#S5.F3)examines the reasoning ceiling via Maj@nnand Avg@nnmetrics across varying samples\. On VStar, BRAID’s Maj@nncurve climbs steadily from 66\.5 to 70\.7 asnngrows from 3 to 17, while curves of SFT and BAGEL plateau early, suggesting that RL endows the model with a broader and more exploitable solution space\. SAT exhibits the same pattern: BRAID attains 61\.3 at Maj@17, exceeding SFT \(56\.7\) and BAGEL \(52\.3\) by\+4\.6and\+9\.0, respectively\. The Avg@nncurves reinforce this interpretation: BRAID’s mean per\-sample quality is consistently higher, ruling out the possibility that its majority\-vote advantage stems solely from greater output diversity\. Taken together, these scaling trends confirm that RL not only aligns the model with the reward signal but also incentivizes genuinely stronger reasoning capabilities\.

#### 5\.2Ablation Study

To isolate the contribution of each component, we evaluate two ablated variants under the same training recipe: \(1\)BRAID w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}, which freezes the image generation parameters during RL and only optimizes the text branch, following the design of COOPER\(Zhanget al\.,[2025](https://arxiv.org/html/2607.03748#bib.bib27)\); and \(2\)BRAID w/orvisr^\{\\text\{vis\}\}, which removes the vision\-thinking process reward and relies solely on the terminal rewardr​\(τ\)r\(\\tau\)for advantage estimation\. Results are reported in Table[4](https://arxiv.org/html/2607.03748#A2.T4)and Figure[4](https://arxiv.org/html/2607.03748#S5.F4)with key findings as

- •Droppingrvisr^\{\\text\{vis\}\}\(BRAIDvs\.w/orvisr^\{\\text\{vis\}\}\) results in a moderate decline \(65\.19 vs\. 62\.68 avg\.\), mainly on VStar \(−\-4\.48\) and SAT \(−\-2\.67\), two benchmarks reliant on fine\-grained visual reasoning\. Training accuracy, however, does not collapse: the shared trajectory\-level advantage still couples the two modalities, allowing the terminal reward signal to backpropagate to image turns\.
- •Excluding the image branch \(w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}\) incurs a steeper decline, notably on VStar \(−\-6\.00\) and CV\-Bench 2D \(−\-4\.32\)\. However, its format reward curve converges faster, presumably because the text branch receives an undiluted reward as the sole trainable modality\. In contrast, BRAID trades slower early format convergence for a higher accuracy ceiling, confirming that jointly optimizing the image branch produces richer visual intermediates that benefit downstream reasoning\.
- •ℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}opens the gradient pathway for image\-turn optimization, whilervisr^\{\\text\{vis\}\}supplies a dedicated signal that steers those gradients toward reasoning\-relevant outputs\. Neither alone suffices: BRAID outperforms the two ablations by\+3\.09and\+2\.51on average, confirming that both components are essential\. Notably, removingℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}is more damaging than removingrvisr^\{\\text\{vis\}\}, underscoring that bringing the image branch into the RL loop \(BRAID’s central premise\) is the more critical factor\.

![Refer to caption](https://arxiv.org/html/2607.03748v1/x4.png)Figure 4:Ablation results\. Left: Final test accuracy across seven benchmarks\. Middle: Training average accuracy\. Right: Training average format reward\. Complete results are available in Table[4](https://arxiv.org/html/2607.03748#A2.T4)\.
#### 5\.3Analysis

![Refer to caption](https://arxiv.org/html/2607.03748v1/x5.png)Figure 5:Complete results are available in Table[7](https://arxiv.org/html/2607.03748#A2.T7)\.Rollout\-to\-Update Ratio\.We study how the rollout\-to\-update batch ratio affects training stability\. Our default \(2:1, rollout 64 / update 32\) maintains a rollout buffer that stabilizes gradient estimates\. We compare this against a more aggressive setting \(Ratio = 1:1, rollout and update batch is 64\) that immediately consumes all rollouts in a single update\. As shown in Table[7](https://arxiv.org/html/2607.03748#A2.T7)and Figure[5](https://arxiv.org/html/2607.03748#S5.F5), the 1:1 ratio leads to a notable decline of avg\. 65\.19 to 61\.93\. The accuracy curve exhibits pronounced oscillations and recovers slowly, indicating that aggressive updates can destabilize the policy\. Interestingly, despite the overall degradation, CV\-Bench 3D improves to84\.83under the 1:1 ratio, even surpassing the default 2:1 setting \(82\.2\)\. We conjecture that higher update frequency encourages aggressive exploration that coincidentally favors this task but destabilizes others; the 2:1 ratio therefore better balances exploration and stable improvement\.

![Refer to caption](https://arxiv.org/html/2607.03748v1/x6.png)Figure 6:Complete results are available in Table[6](https://arxiv.org/html/2607.03748#A2.T6)\.Data Allocation Strategy\.We compare three data\-allocation regimes \(detailed in Appendix[A\.1](https://arxiv.org/html/2607.03748#A1.SS1)\): a disjoint split with equal budget \(𝒟1:1\\mathcal\{D\}^\{1:1\},ω=0\\omega\{=\}0\), a disjoint split favoring SFT \(𝒟2:1\\mathcal\{D\}^\{2:1\},ω=0\\omega\{=\}0\), and a replay regime \(𝒟∩\\mathcal\{D\}^\{\\cap\},ω≈0\.32\\omega\{\\approx\}0\.32\) where the RL pool partially overlaps with SFT data\. Results are reported in Table[6](https://arxiv.org/html/2607.03748#A2.T6)and Figure[6](https://arxiv.org/html/2607.03748#S5.F6)\. The𝒟1:1\\mathcal\{D\}^\{1:1\}SFT checkpoint already exhibits weakness on several benchmarks \(e\.g\., CV\-Bench 2D drops to 65\.24\), indicating that an insufficient SFT budget yields a suboptimal initialization that constrains downstream RL gains\. We therefore did not proceed with RL training from this checkpoint\. With adequate SFT initialization \(𝒟2:1\\mathcal\{D\}^\{2:1\}\), RL brings consistent gains and reaches62\.63avg\. The replay regime𝒟∩\\mathcal\{D\}^\{\\cap\}further improves this to65\.19avg\., with pronounced gains on SAT \(\+9\.34over𝒟2:1\\mathcal\{D\}^\{2:1\}RL\) and VStar \(\+3\.43\)\. This confirms that replaying a portion of SFT data during RL bridges imitation learning and on\-policy exploration, enabling the model to revisit familiar prompts under its own policy and thereby consolidate supervised knowledge while continuing to discover novel reasoning strategies\.

#### 5\.4Case Study

Figure 7:Text\-to\-image generation comparison\.Interleaved Reasoning\.We qualitatively compare BRAID and its ablations on a visual search task \(Appendix[C](https://arxiv.org/html/2607.03748#A3)\)\. BRAID generates diverse, high\-fidelity intermediate images that directly support correct reasoning\. In contrast, w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}produces distorted visuals that induce hallucinated text reasoning, and w/orvisr^\{\\text\{vis\}\}introduces fabricated artifacts\. Although all variants reach the correct answer on this instance, only BRAID’s reasoning chain is grounded in faithful visual evidence, consistent with its superior robustness under Maj@nnsampling\.

Text\-to\-Image Generation\.Figure[7](https://arxiv.org/html/2607.03748#S5.F7)shows whether BRAID’s RL training transfers to standalone generation on demo prompts requiring spatial compositionality and numerical accuracy\. WithoutℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}, the image branch fails to follow instructions \(cat placed beside the fishbowl; balloon count exceeds five\)\. Removingrvisr^\{\\text\{vis\}\}improves spatial relations but still lacks numerical precision\. Only BRAID correctly satisfies both prompts, indicating that the vision\-thinking reward provides optimization direction that generalizes beyond interleaved reasoning \(full comparison in Appendix[D](https://arxiv.org/html/2607.03748#A4)\)\.

### 6Conclusions, Limitations, and Future Work

We presented BRAID, a framework that casts interleaved text\-image\-text reasoning as a unified MDP, enabling joint optimization over both textual and visual generation within a unified policy\. By deriving a shared trajectory\-level advantage and channeling it into modality\-native policy gradients, BRAID propagates reward signals end\-to\-end across modality boundaries\. A complementary vision\-thinking process reward, scored by a VLM judge, supplies turn\-level feedback that sharpens credit assignment at critical image branches\. Experiments on seven spatial\-reasoning and visual\-perception benchmarks demonstrate that BRAID consistently outperforms various baselines\.

Several limitations point to promising future directions\. First, BRAID is built upon BAGEL’s hybrid AR\-diffusion backbone; extending to fully autoregressive UMMsWanget al\.\([2024](https://arxiv.org/html/2607.03748#bib.bib2)\), where discrete image tokens are generated via next\-token prediction, would broaden the framework’s applicability\. Second, the vision\-thinking reward currently relies on an external VLM judge, inducing inference cost and dependency; distilling it into a compact, jointly trained PRM would improve both efficiency and reproducibility\. Finally, the current formulation assumes a fixed interleaving pattern; exploring adaptive turn structures and richer credit\-assignment mechanisms, such as hierarchical advantage decomposition across turns, could further unlock RL’s potential for interleaved multi\-modal reasoning\.

### References

- S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang, H\. Zhong, Y\. Zhu, M\. Yang, Z\. Li, J\. Wan, P\. Wang, W\. Ding, Z\. Fu, Y\. Xu, J\. Ye, X\. Zhang, T\. Xie, Z\. Cheng, H\. Zhang, Z\. Yang, H\. Xu, and J\. Lin \(2025\)Qwen2\.5\-vl technical report\.External Links:2502\.13923,[Link](https://arxiv.org/abs/2502.13923)Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p2.5)\.
- Training diffusion models with reinforcement learning\.arXiv preprint arXiv:2305\.13301\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- Chameleon Team \(2024\)Chameleon: mixed\-modal early\-fusion foundation models\.arXiv preprint arXiv:2405\.09818\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§2](https://arxiv.org/html/2607.03748#S2.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p2.5)\.
- A\. Chang, A\. Dai, T\. Funkhouser, M\. Halber, M\. Niessner, M\. Savva, S\. Song, A\. Zeng, and Y\. Zhang \(2017\)Matterport3d: learning from rgb\-d data in indoor environments\.arXiv preprint arXiv:1709\.06158\.Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- H\. Chen, Y\. Hou, C\. Qu, I\. Testini, X\. Hong, and J\. Jiao \(2024\)360\+ x: a panoptic multi\-modal scene understanding dataset\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 19373–19382\.Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- Z\. Cheng, Q\. Chen, X\. Xu, J\. WANG, W\. Wang, H\. Fei, Y\. Wang, A\. J\. Wang, Z\. Chen, W\. Che,et al\.\(2025\)Visual thoughts: a unified perspective of understanding multimodal chain\-of\-thought\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- E\. Chern, Z\. Hu, S\. Chern, S\. Kou, J\. Su, Y\. Ma, Z\. Deng, and P\. Liu \(2025\)Thinking with generated images\.arXiv preprint arXiv:2505\.22525\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1)\.
- Y\. Cui, H\. Chen, H\. Deng, X\. Huang, X\. Li, J\. Liu, Y\. Liu, Z\. Luo, J\. Wang, W\. Wang,et al\.\(2025\)Emu3\.5: native multimodal models are world learners\.arXiv preprint arXiv:2510\.26583\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§2](https://arxiv.org/html/2607.03748#S2.p3.1)\.
- C\. Deng, D\. Zhu, K\. Li, C\. Gou, F\. Li, Z\. Wang, S\. Zhong, W\. Yu, X\. Nie, Z\. Song, G\. Shi, and H\. Fan \(2025\)Emerging properties in unified multimodal pretraining\.arXiv preprint arXiv:2505\.14683\.Cited by:[§A\.3](https://arxiv.org/html/2607.03748#A1.SS3.p1.1),[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§1](https://arxiv.org/html/2607.03748#S1.p3.1),[§2](https://arxiv.org/html/2607.03748#S2.p1.1),[§4\.1](https://arxiv.org/html/2607.03748#S4.SS1.p1.5),[§5](https://arxiv.org/html/2607.03748#S5.p2.5)\.
- R\. Dong, C\. Han, Y\. Peng,et al\.\(2024\)DreamLLM: synergistic multimodal comprehension and creation\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- H\. Duan, J\. Yang, Y\. Qiao, X\. Fang, L\. Chen, Y\. Liu, X\. Dong, Y\. Zang, P\. Zhang, J\. Wang,et al\.\(2024\)Vlmevalkit: an open\-source toolkit for evaluating large multi\-modality models\.InProceedings of the 32nd ACM international conference on multimedia,pp\. 11198–11201\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- Y\. Fan, O\. Watkins, Y\. Du, H\. Liu, M\. Ryu, C\. Boutilier, P\. Abbeel, M\. Ghavamzadeh, K\. Lee, and K\. Lee \(2023\)Dpok: reinforcement learning for fine\-tuning text\-to\-image diffusion models\.Advances in Neural Information Processing Systems36,pp\. 79858–79885\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- X\. Fu, Y\. Hu, B\. Li, Y\. Feng, H\. Wang, X\. Lin, D\. Roth, N\. A\. Smith, W\. Ma, and R\. Krishna \(2024\)Blink: multimodal large language models can see but not perceive\.InEuropean Conference on Computer Vision,pp\. 148–166\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- J\. Gao, Y\. Li, Z\. Cao, and W\. Li \(2024\)Interleaved\-modal chain\-of\-thought\.arXiv preprint arXiv:2411\.19488\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1)\.
- Z\. Geng, Y\. Wang, Y\. Ma, C\. Li, Y\. Rao, S\. Gu, Z\. Zhong, Q\. Lu, H\. Hu, X\. Zhang,et al\.\(2025\)X\-omni: reinforcement learning makes discrete autoregressive image generative models great again\.arXiv preprint arXiv:2507\.22058\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p3.1)\.
- J\. Gu, Y\. Hao, H\. W\. Wang, L\. Li, M\. Q\. Shieh, Y\. Choi, R\. Krishna, and Y\. Cheng \(2025\)ThinkMorph: emergent properties in multimodal interleaved chain\-of\-thought reasoning\.arXiv preprint arXiv:2510\.27492\.Cited by:[§A\.1](https://arxiv.org/html/2607.03748#A1.SS1.p1.12),[§A\.3](https://arxiv.org/html/2607.03748#A1.SS3.p2.4),[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§2](https://arxiv.org/html/2607.03748#S2.p1.1),[§4\.1](https://arxiv.org/html/2607.03748#S4.SS1.p1.9),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi,et al\.\(2025\)DeepSeek\-r1: incentivizing reasoning capability in LLMs via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- Y\. Hu, W\. Shi, X\. Fu, D\. Roth, M\. Ostendorf, L\. Zettlemoyer, N\. A\. Smith, and R\. Krishna \(2024\)Visual sketchpad: sketching as a visual chain of thought for multimodal language models\.Advances in Neural Information Processing Systems37,pp\. 139348–139379\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- W\. Huang, B\. Jia, Z\. Zhai, S\. Cao, Z\. Ye, F\. Zhao, Z\. Xu, X\. Tang, Y\. Hu, and S\. Lin \(2025\)Vision\-r1: incentivizing reasoning capability in multimodal large language models\.arXiv preprint arXiv:2503\.06749\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.\(2024\)Gpt\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p2.5)\.
- T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. Iwasawa \(2022\)Large language models are zero\-shot reasoners\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 22199–22213\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- A\. Li, C\. Wang, D\. Fu, K\. Yue, Z\. Cai, W\. B\. Zhu, O\. Liu, P\. Guo, W\. Neiswanger, F\. Huang,et al\.\(2025a\)Zebra\-cot: a dataset for interleaved vision language reasoning\.arXiv preprint arXiv:2507\.16746\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- C\. Li, W\. Wu, H\. Zhang, Y\. Xia, S\. Mao, L\. Dong, I\. Vulić, and F\. Wei \(2025b\)Imagine while reasoning in space: multimodal visualization\-of\-thought\.arXiv preprint arXiv:2501\.07542\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- L\. Li, Y\. Wu, X\. Li, L\. Wang, T\. Rao, J\. Zhou, C\. Pan, and X\. Hui \(2025c\)Realsee3D: a large\-scale multi\-view rgb\-d dataset of indoor scenes \(version 1\.0\)\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.17826243),[Link](https://doi.org/10.5281/zenodo.17826243)Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- Y\. Lipman, R\. T\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le \(2022\)Flow matching for generative modeling\.arXiv preprint arXiv:2210\.02747\.Cited by:[§3](https://arxiv.org/html/2607.03748#S3.SS0.SSS0.Px2.p1.7)\.
- J\. Liu, G\. Liu, J\. Liang, Y\. Li, J\. Liu, X\. Wang, P\. Wan, D\. Zhang, and W\. Ouyang \(2025\)Flow\-grpo: training flow matching models via online RL\.arXiv preprint arXiv:2505\.05470\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- J\. Liu, Z\. Ye, L\. Yuan, S\. Zhu, Y\. Gao, J\. Wu, K\. Li, X\. Wang, X\. Nie, W\. Huang, and W\. Ouyang \(2026a\)UniGRPO: unified policy optimization for reasoning\-driven visual generation\.arXiv preprint arXiv:2603\.23500\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p3.1)\.
- J\. Liu, H\. Li, Z\. Sun, C\. Chen, Y\. Bian, B\. Wang, D\. Dong, C\. Chen, and Z\. Wang \(2026b\)Beyond the Dirac Delta: mitigating diversity collapse in reinforcement fine\-tuning for versatile image generation\.arXiv preprint arXiv:2601\.12401\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- F\. Meng, L\. Du, Z\. Liu, Z\. Zhou, Q\. Lu, D\. Fu, T\. Han, B\. Shi, W\. Wang, J\. He, K\. Zhang, P\. Luo, Y\. Qiao, Q\. Zhang, and W\. Shao \(2025\)MM\-eureka: exploring the frontiers of multimodal reasoning with rule\-based reinforcement learning\.arXiv preprint arXiv:2503\.07365\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- F\. Meng, H\. Yang, Y\. Wang, and M\. Zhang \(2023\)Chain of images for intuitively reasoning\.arXiv preprint arXiv:2311\.09241\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- S\. Menon, R\. Zemel, and C\. Vondrick \(2024\)Whiteboard\-of\-thought: thinking step\-by\-step across modalities\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 20016–20031\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- M\. Nie, C\. Wang, J\. Han, H\. Xu, and L\. Zhang \(2025\)Towards unified multimodal interleaved generation via group relative policy optimization\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p3.1)\.
- A\. Ray, J\. Duan, R\. Tan, D\. Bashkirova, R\. Hendrix, K\. Ehsani, A\. Kembhavi, B\. A\. Plummer, R\. Krishna, K\. Zeng,et al\.\(2024\)Sat: spatial aptitude training for multimodal language models\.arXiv preprint arXiv:2412\.077553\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- H\. Shao, S\. Qian, H\. Xiao, G\. Song, Z\. Zong, L\. Wang, Y\. Liu, and H\. Li \(2024a\)Visual cot: unleashing chain\-of\-thought reasoning in multi\-modal language models\.arXiv preprint arXiv:2403\.169992\.Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu,et al\.\(2024b\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1),[§3](https://arxiv.org/html/2607.03748#S3.SS0.SSS0.Px1.p1.12)\.
- S\. Tong, E\. Brown, P\. Wu, S\. Woo, M\. Middepogu, S\. C\. Akula, J\. Yang, S\. Yang, A\. Iyer, X\. Pan,et al\.\(2024a\)Cambrian\-1: a fully open, vision\-centric exploration of multimodal llms\.Advances in Neural Information Processing Systems37,pp\. 87310–87356\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- S\. Tong, D\. Fan, J\. Li, Y\. Xiong, X\. Chen, K\. Sinha, M\. Rabbat, Y\. LeCun, S\. Xie, and Z\. Liu \(2025\)Metamorph: multimodal understanding and generation via instruction tuning\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 17001–17012\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- S\. Tong, Z\. Liu, Y\. Zhai, Y\. Ma, Y\. LeCun, and S\. Xie \(2024b\)Eyes wide shut? exploring the visual shortcomings of multimodal llms\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 9568–9578\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- X\. Wang, X\. Zhang, Z\. Luo, Q\. Sun, Y\. Cui, J\. Wang, F\. Zhang, Y\. Wang, Z\. Li, Q\. Yu,et al\.\(2024\)Emu3: next\-token prediction is all you need\.arXiv preprint arXiv:2409\.18869\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1),[§6](https://arxiv.org/html/2607.03748#S6.p2.1)\.
- Z\. Wang, J\. Zhu, B\. Tang, Z\. Li, F\. Xiong, J\. Yu, and M\. B\. Blaschko \(2025\)Jigsaw\-r1: a study of rule\-based visual reinforcement learning with jigsaw puzzles\.arXiv preprint arXiv:2505\.23590\.Cited by:[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24824–24837\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- C\. Wu, X\. Chen, Z\. Wu, Y\. Ma, X\. Liu, Z\. Pan, W\. Liu, Z\. Xie, X\. Yu, C\. Ruan, and P\. Luo \(2024\)Janus: decoupling visual encoding for unified multimodal understanding and generation\.arXiv preprint arXiv:2410\.13848\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p2.5)\.
- J\. Wu, X\. Zhang, H\. Yuan, X\. Zhang, T\. Huang, C\. He, C\. Deng, R\. Zhang, Y\. Wu, and M\. Long \(2026\)Visual generation unlocks human\-like reasoning through multimodal world models\.arXiv preprint arXiv:2601\.19834\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1)\.
- P\. Wu and S\. Xie \(2024\)V\*: guided visual search as a core mechanism in multimodal llms\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13084–13094\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- J\. Xie, W\. Mao, Z\. Bai, D\. J\. Zhang, W\. Wang, K\. Q\. Lin, Y\. Gu, Z\. Chen, Z\. Yang, and M\. Z\. Shou \(2024\)Show\-o: one single transformer to unify multimodal understanding and generation\.arXiv preprint arXiv:2408\.12528\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.
- Z\. Xue, J\. Wu, Y\. Gao, F\. Kong, L\. Zhu, M\. Chen, Z\. Liu, W\. Liu, Q\. Guo, W\. Huang,et al\.\(2025\)Dancegrpo: unleashing grpo on visual generation\.arXiv preprint arXiv:2505\.07818\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p2.1)\.
- S\. Yang, R\. Xu, Y\. Xie, S\. Yang, M\. Li, J\. Lin, C\. Zhu, X\. Chen, H\. Duan, X\. Yue,et al\.\(2025a\)Mmsi\-bench: a benchmark for multi\-image spatial intelligence\.arXiv preprint arXiv:2505\.23764\.Cited by:[§A\.2](https://arxiv.org/html/2607.03748#A1.SS2.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.03748#S5.p1.6)\.
- Y\. Yang, X\. He, H\. Pan, X\. Jiang, Y\. Deng, X\. Yang, H\. Lu, D\. Yin, F\. Rao, M\. Zhu,et al\.\(2025b\)R1\-onevision: advancing generalized multimodal reasoning through cross\-modal formalization\.arXiv preprint arXiv:2503\.10615\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p1.1)\.
- Z\. Zhang, X\. Hao, H\. Tang, Z\. Zhang, J\. Sheng, X\. Li, Z\. Li, L\. Gao, D\. Shi, D\. Yin, and T\. Liu \(2025\)COOPER: a unified model for cooperative perception and reasoning in spatial intelligence\.arXiv preprint arXiv:2512\.04563\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p3.1),[§5\.2](https://arxiv.org/html/2607.03748#S5.SS2.p1.3)\.
- K\. Zheng, H\. Chen, H\. Ye, H\. Wang, Q\. Zhang, K\. Jiang, H\. Su, S\. Ermon, J\. Zhu, and M\. Liu \(2025a\)DiffusionNFT: online diffusion reinforcement with forward process\.arXiv preprint arXiv:2509\.16117\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p2.1),[§3](https://arxiv.org/html/2607.03748#S3.SS0.SSS0.Px2.p2.2),[§4\.3](https://arxiv.org/html/2607.03748#S4.SS3.p1.2)\.
- Z\. Zheng, M\. Yang, J\. Hong, C\. Zhao, G\. Xu, L\. Yang, C\. Shen, and X\. Yu \(2025b\)DeepEyes: incentivizing “thinking with images” via reinforcement learning\.arXiv preprint arXiv:2505\.14362\.Cited by:[§1](https://arxiv.org/html/2607.03748#S1.p2.1),[§2](https://arxiv.org/html/2607.03748#S2.p3.1)\.
- C\. Zhou, L\. Yu, A\. Babu, K\. Tirumala, M\. Yasunaga, L\. Shamis, J\. Kahn, X\. Ma, L\. Zettlemoyer, and O\. Levy \(2024\)Transfusion: predict the next token and diffuse images with one multi\-modal model\.arXiv preprint arXiv:2408\.11039\.Cited by:[§2](https://arxiv.org/html/2607.03748#S2.p1.1)\.

## Appendix

### Appendix AExperimental Details

#### A\.1Data Configuration

Following ThinkMorph and its follow\-ups\[[16](https://arxiv.org/html/2607.03748#bib.bib17)\], our training corpus spans eight task families \(Matterport3D,RealSee3D,Perspective,DirectionalQuery,360\+x360\{\+\}x,PositionedDirection,Jigsaw,VisualSearch\), each contributing both SFT and RL split\. To study how the SFT/RL data budget interacts with our two\-stage pipeline, we define an*overlap ratio*ω≜\|𝒟SFT∩𝒟RL\|/\|𝒟RL\|∈\[0,1\]\\omega\\triangleq\|\\mathcal\{D\}\_\{\\text\{SFT\}\}\\cap\\mathcal\{D\}\_\{\\text\{RL\}\}\|/\|\\mathcal\{D\}\_\{\\text\{RL\}\}\|\\in\[0,1\], and consider three data\-allocation regimes \(per\-task counts in Table[2](https://arxiv.org/html/2607.03748#A1.T2)\):\(1\)a*disjoint*regime \(𝒟2:1\\mathcal\{D\}^\{2\{:\}1\},ω=0\\omega\{=\}0\) with a2:12\{:\}1SFT/RL budget, favoring supervised initialization;\(2\)a*disjoint*regime \(𝒟1:1\\mathcal\{D\}^\{1\{:\}1\},ω=0\\omega\{=\}0\) with a1:11\{:\}1budget, isolating the effect of equal\-sized splits; and\(3\)a*replay*regime \(𝒟∩\\mathcal\{D\}^\{\\cap\}\) where the RL pool augments the original RL split with additional samples drawn from the SFT pool \(20,00020\{,\}000/15,20015\{,\}200,ω≈0\.32\\omega\{\\approx\}0\.32\), so that part of the RL prompts revisit SFT examples under on\-policy rollouts\. We adopt regime\(3\)as our default\.

Table 2:Per\-task sample counts under the three data\-allocation regimes\. In𝒟∩\\mathcal\{D\}^\{\\cap\}, the RL pool extends the original RL split with4,8004\{,\}800samples drawn from the SFT pool, yielding an overlap ratioω≈0\.32\\omega\\approx 0\.32\.
#### A\.2Evaluation Benchmarks

We evaluate on seven benchmarks spanning two categories\. All evaluations are conducted under the VLMEvalKit\[[11](https://arxiv.org/html/2607.03748#bib.bib50)\]framework with accuracy as the primary metric\.

##### Spatial Reasoning\.

\(1\)MMSI\-Bench\[[47](https://arxiv.org/html/2607.03748#bib.bib47)\]is a multi\-image spatial intelligence benchmark comprising 1,000 multiple\-choice questions across 10 fundamental spatial tasks \(e\.g\., object motion tracking, camera ego\-motion, scene reconstruction\), curated from over 120,000 images by 3D\-vision experts\. \(2\)SAT\[[33](https://arxiv.org/html/2607.03748#bib.bib48)\]evaluates dynamic and static spatial reasoning, including egocentric movement, allocentric perspective shifts, depth relations, and action consequence prediction\. It contains over 218,000 QA pairs across 22,000 synthetic scenes generated with the ProcTHOR engine, plus 150 real\-image QA pairs for out\-of\-distribution testing\. \(3\)MMVP\[[38](https://arxiv.org/html/2607.03748#bib.bib45)\]targets systematic failures of vision encoders by constructing “CLIP\-blind pairs”—image pairs that CLIP encodes similarly despite clear visual differences in orientation, counting, color, or viewpoint\. It contains approximately 150 image pairs formulated as multiple\-choice questions\. \(4\)CV\-Bench 3D\[[36](https://arxiv.org/html/2607.03748#bib.bib49)\]is the 3D split of the Cambrian Vision\-Centric Benchmark, evaluating depth ordering and relative distance understanding via natural language questions derived from OMNI3D annotations\.

##### Visual Perception\.

\(1\)BLINK\[[13](https://arxiv.org/html/2607.03748#bib.bib46)\]contains 3,807 multiple\-choice questions reformulated from 14 classic computer vision tasks \(relative depth, visual correspondence, multi\-view reasoning, forensic detection, spatial relations, etc\.\) that humans can solve “in a blink” but multimodal LLMs find challenging\. \(2\)V∗Bench\[[44](https://arxiv.org/html/2607.03748#bib.bib44)\]evaluates fine\-grained visual perception in high\-resolution, visually crowded images, comprising 191 multiple\-choice questions on attribute recognition \(color, material\) and spatial relationship reasoning for small or occluded objects that require targeted visual search\. \(3\)CV\-Bench 2D\[[36](https://arxiv.org/html/2607.03748#bib.bib49)\]is the 2D split of the Cambrian Vision\-Centric Benchmark, assessing spatial relationship understanding and object counting via questions derived from ADE20k and COCO annotations\.

#### A\.3Training Recipe

Table 3:Key hyperparameters for RL training\.HyperparameterValueHyperparameterValueRollout batch size64NFTβ\\beta0\.4Update batch size32NFT training timesteps1Rollouts per prompt8Denoising steps20PPO epochs1Timestep shift3\.0Training steps120CFG text scale4\.0Clip ratio \(high / low\)0\.28 / 0\.20CFG image scale2\.0Entropy coefficient0\.0CFG interval\[0\.4,1\.0\]\[0\.4,1\.0\]KL lossdisabledNoise level0\.7Sampling temperature1\.0Max latent size64Max interleaved rounds3VLM judgeGPT\-5\.2Max thinking tokens2048Reward weights0\.25 eachMax total tokens32768We build on BAGEL\-7B\[[9](https://arxiv.org/html/2607.03748#bib.bib4)\]and train with a two\-stage pipeline: SFT followed by RL\. Both stages use FSDP2 with hybrid sharding, gradient checkpointing, and the VAE frozen throughout\. During RL, all other modules \(LLM, ViT, understanding head, generation head\) are unfrozen and jointly optimized\.

For the SFT stage, we follow ThinkMorph\[[16](https://arxiv.org/html/2607.03748#bib.bib17)\]and train on curated TIT\-CoT trajectories under the𝒟2:1\\mathcal\{D\}^\{2:1\}data regime \(detailed in Section[A\.1](https://arxiv.org/html/2607.03748#A1.SS1)\)\. For the RL stage, we adopt the replay regime𝒟∩\\mathcal\{D\}^\{\\cap\}\(ω≈0\.32\\omega\{\\approx\}0\.32\) as default, where the RL pool partially overlaps with SFT data to bridge imitation learning and on\-policy exploration\. We sample rollouts with temperature 1\.0 and allow up to 3 interleaved image rounds per trajectory, with a maximum of 2048 thinking tokens per round and 32768 total tokens\. Image generation uses 20 denoising steps \(for both training and evaluation\) with a timestep shift of 3\.0, classifier\-free guidance scales of 4\.0 \(text\) and 2\.0 \(image\), and a guidance interval of\[0\.4,1\.0\]\[0\.4,1\.0\]\. The noise level for intermediate images is set to 0\.7\. Table[3](https://arxiv.org/html/2607.03748#A1.T3)summarizes the key hyperparameters\.

#### A\.4Vision\-Thinking Process Reward Prompt

To instantiate the per\-turn visual process reward𝒓vis\\bm\{r\}^\{\\text\{vis\}\}, we employGPT\-5\.2as the VLM judge\. Given the question, the problem image\(s\), and the generated intermediate imageIkI\_\{k\}, the judge follows a three\-step protocol \(question understanding→\\rightarrowimage analysis→\\rightarrowscoring\) and emits four integer scores in\[1,10\]\[1,10\], corresponding one\-to\-one to the four criteria\{r\(vc\),r\(vf\),r\(ru\),r\(tr\)\}\\\{r^\{\(\\text\{vc\}\)\},r^\{\(\\text\{vf\}\)\},r^\{\(\\text\{ru\}\)\},r^\{\(\\text\{tr\}\)\}\\\}defined in Sec\.[4\.4](https://arxiv.org/html/2607.03748#S4.SS4)\. The full system prompt is shown below\.

You are an expert evaluator for a vision\-language model \(VLM\) that reasons by generating intermediate images as visual tools\. Your task: evaluate one generated intermediate image within a multi\-step reasoning trajectory\.STEP 1 — UNDERSTAND THE QUESTION & PROBLEM IMAGE\(S\)Read the question carefully\. Identify:•What specific visual information is needed to answer it?•What is provided by the problem image\(s\), and what remains ambiguous?•What type of visual operation would best help answer this question?STEP 2 — ANALYZE THE GENERATED IMAGEDescribe concretely what it shows\. Compare it carefully against the problem image\(s\) — what is preserved, what is changed, what is new, what might be fabricated? Be very precise: zoom in on small details, distinguish between nearby objects, and verify that what the model claims to target is actually what the generated image shows\.STEP 3 — SCORE on the 4 dimensions below \(1–10 each, integers only\)\.1\.action\_correctness— Did the model choose the RIGHT visual operation and execute it correctly? 10 = perfect execution\. 1 = completely wrong\.2\.visual\_faithfulness— Does the generated image faithfully represent the actual content of the problem image\(s\)? Only penalize fabrications that could mislead reasoning\. 10 = pixel\-accurate\. 1 = entirely fabricated\.3\.reasoning\_contribution— Does this image provide NEW, CORRECT information that could advance reasoning toward the correct answer? 10 = decisive new evidence\. 1 = useless or actively misleading\.4\.misleading\_risk— Could this image cause a reasonable model to reach a \*\*wrong\*\* answer? 10 = completely safe\. 1 = virtually guarantees wrong answer\.Be willing to use the full 1–10 range\. A score of 8\+ means genuinely good\.Return \*\*only\*\* valid \*\*json\*\* \(no markdown, no extra text\):\{"action\_correctness": <int\>, "visual\_faithfulness": <int\>, "reasoning\_contribution": <int\>, "misleading\_risk": <int\>\}

### Appendix BAdditional Results

#### B\.1Ablation Study

Table 4:Ablation study on different components of BRAID across seven spatial reasoning and perception benchmarks\. Best results inboldand second bestunderlined\.Table[4](https://arxiv.org/html/2607.03748#A2.T4)reports the complete numerical results of the ablation study discussed in Section[5\.2](https://arxiv.org/html/2607.03748#S5.SS2), covering all seven benchmarks for BAGEL, SFT, and the two ablated variants\. Several per\-benchmark patterns are worth noting\. On MMSI, all RL variants outperform SFT, indicating that even withoutrvisr^\{\\text\{vis\}\}orℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}, the text\-branch RL alone provides gains on multi\-image spatial tasks\. On CV\-Bench 3D, both BRAID and w/orvisr^\{\\text\{vis\}\}recover the SFT regression \(82\.17 vs\. 81\.67\), whereas w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}does not \(81\.50\), suggesting that optimizing the image branch is necessary to recover 3D spatial understanding\. The perception benchmarks \(BLINK, VStar, CV\-Bench 2D\) show the largest gap between the full model and ablations, confirming that fine\-grained visual tasks benefit most from the combination of image\-branch RL and dense process reward\.

#### B\.2Analysis of Reasoning Capacity

Table 5:Performance comparison of different methods on Vstar and SAT benchmarks under majority voting \(maj@n\) and average \(avg@n\) atn∈\{3,5,9,17\}n\\in\\\{3,5,9,17\\\}\.Table[5](https://arxiv.org/html/2607.03748#A2.T5)provides the full Maj@nnand Avg@nnresults acrossn∈\{3,5,9,17\}n\\in\\\{3,5,9,17\\\}on VStar and SAT, corresponding to the curves in Figure[3](https://arxiv.org/html/2607.03748#S5.F3)\. BRAID consistently achieves the highest scores at all sampling budgets, with the gap widening asnnincreases\. Notably, on VStar the Maj@17 gap between BRAID and SFT is \+3\.15, whereas the Avg@17 gap is \+2\.34, indicating that BRAID benefits from both higher per\-sample quality and greater diversity\. On SAT, BRAID’s Avg@nnimproves monotonically from 58\.11 to 60\.75 asnngrows, while SFT and BAGEL plateau or decline, suggesting that RL produces a policy whose samples are more complementary and less redundant\.

#### B\.3Analysis of Training Recipe

Table 6:Effect of data\-allocation regimes \(𝒟2:1\\mathcal\{D\}^\{2:1\},𝒟1:1\\mathcal\{D\}^\{1:1\},𝒟∩\\mathcal\{D\}^\{\\cap\}\) at the SFT and RL stages across seven benchmarks\. Best results inboldand second bestunderlined\. “–” indicates that we did not run RL from this SFT checkpoint due to its noticeable weaknesses on several benchmarks\.Table[6](https://arxiv.org/html/2607.03748#A2.T6)presents the full results under the three data\-allocation regimes discussed in Section[5\.3](https://arxiv.org/html/2607.03748#S5.SS3)\. The𝒟1:1\\mathcal\{D\}^\{1:1\}SFT checkpoint underperforms BAGEL on CV\-Bench 2D \(65\.24 vs\. 73\.53\), a−\-8\.29 drop that we attribute to the reduced SFT budget failing to cover sufficient visual perception data\. By contrast, the𝒟2:1\\mathcal\{D\}^\{2:1\}regime provides a stable SFT initialization that enables effective downstream RL\. The replay regime𝒟∩\\mathcal\{D\}^\{\\cap\}further improves upon𝒟2:1\\mathcal\{D\}^\{2:1\}RL on all benchmarks except BLINK \(tied at 57\.76\), with the largest gains on SAT \(\+9\.34\) and VStar \(\+3\.43\)\. This pattern suggests that replaying SFT prompts under on\-policy rollouts is particularly beneficial for tasks requiring multi\-step spatial reasoning, where the model can consolidate learned strategies while exploring new ones\.

Table 7:Effect of the rollout\-to\-update batch size ratio across seven spatial reasoning and perception benchmarks\. “Ratio” denotes the rollout batch size to update batch size ratio used during RL training\. Best results inboldand second bestunderlined\.Table[7](https://arxiv.org/html/2607.03748#A2.T7)reports the complete benchmark results for the rollout\-to\-update ratio analysis\. The 2:1 ratio achieves the best average \(65\.19 vs\. 61\.93\), with pronounced advantages on MMSI \(\+7\.20\), SAT \(\+8\.67\), and VStar \(\+8\.15\)\. The 1:1 ratio, despite its overall lower performance, achieves the highest score on CV\-Bench 3D \(84\.83\) and BLINK \(59\.07\), both exceeding even the 2:1 setting\. We hypothesize that the more aggressive update frequency acts as implicit exploration pressure, which benefits benchmarks where the SFT initialization is already close to a local optimum but fails to generalize on tasks requiring stable policy improvement over many steps\.

### Appendix CInterleave Reasoning Case

We present a qualitative comparison of BRAID and its ablated variants on a visual search task from the VStar benchmark\.

Question and AnswerProblem Image: ![[Uncaptioned image]](https://arxiv.org/html/2607.03748v1/Figs/case/problem_image.png)Question:What is the license plate number of the police van?Options:\(A\) OU68 EPD \(B\) QU68 EPO \(C\) OV68 EDP \(D\) OU86 EPDGround Truth:\(A\) OU68 EPD

BRAID — Rollout 1\(✓ Correct\)<think\>To find the correct answer, first examine the image and locate the police van, which is the main vehicle in the scene\. Focus on the rear section of the van, as the license plate is always displayed at the back of public service vehicles\. On the bottom center of the rear doors, just above the yellow safety strip and below the right door handle, there is a yellow plate with black text\. This plate is the vehicle’s registration number\. Carefully read the sequence of characters on this plate: it consists of a two\-letter vanity plate followed by three digits, then another three characters\. The style and spacing match a typical license plate, and the characters are clearly separated by spaces\. That specific yellow plate, with the visible text inside it, is where you should place the red bounding box\.</think\><image\_start\> Generated Image: ![[Uncaptioned image]](https://arxiv.org/html/2607.03748v1/Figs/case/braid_rollout1.png) <image\_end\><think\>With the yellow license plate at the lower center of the rear doors now highlighted, you can focus on the black text printed on the yellow background\. The sequence is a two big characters, then a gap, followed by three digits, then another gap and three more characters\. Comparing that exact sequence to the options, only the option \(A\) OU688EPO matches this pattern: the digit in the first sequence is the same as the second one in the options, the second sequence in the second and the two last characters match in the two letters and you can see on the plate\. Because the highlighted area is the only area you can clearly read the plate, this identified is the correct answer: A\.</think\><answer\>A</answer\>

BRAID — Rollout 2\(✓ Correct\)<think\>To locate the relevant region, first look at the police van in the image\. The question asks for the license plate, so focus on the rear of the vehicle\. At the bottom center of the back, just above the bumper and between the two rear door sections, there is a rectangular plate with a black background and yellow border\. Inside this rectangle, there are black characters arranged in a combination of letters and numbers\. This rectangular area, containing the black\-and\-yellow license plate, is where the red bounding box should be drawn\.</think\><image\_start\> Generated Image: ![[Uncaptioned image]](https://arxiv.org/html/2607.03748v1/Figs/case/braid_rollout2.png) <image\_end\><think\>With the correct license plate now clearly highlighted, you can concentrate on reading the characters inside the rectangle\. The plate shows a two\-letter country code starting with “OU”, then the number “68” repeated twice, and the letters “EPD” at the end\. Matching this full sequence against the answer choices, it corresponds exactly to option \(A\), so the correct answer is A\.</think\><answer\>A</answer\>

BRAID w/orvisr^\{\\text\{vis\}\}— Rollout 1\(✓ Correct, but hallucinated content\)<think\>To locate the relevant region, first focus on the back of the police van, since license plates are typically placed there\. Scan the lower central part of the rear doors, just above the bumper and below the rear windows\. There you can see a rectangular yellow plate with a black border and black text on it\. This plate contains the vehicle’s registration number\. The red bounding box should be drawn tightly around this yellow rectangular plate at the bottom middle of the rear doors, enclosing all the black alphanumeric characters printed on it\.</think\><image\_start\> Generated Image: ![[Uncaptioned image]](https://arxiv.org/html/2607.03748v1/Figs/case/wo_rvis.png) <image\_end\><think\>With the license plate now clearly highlighted, you can read the characters printed on it\. The plate shows a combination starting with the letters “OU68” followed by “EPD” and ending with the same letter “D” at the far right\. Comparing this complete sequence with the multiple\-choice options, it matches exactly option \(A\), so the correct answer is A\.</think\><answer\>A</answer\>

BRAID w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}— Rollout 1\(✓ Correct, but misleading image\)<think\>To locate the relevant region, first focus on the police van in the image\. The question asks about the license plate, so look at the rear of the vehicle\. At the bottom center of the back, just above the bumper and between the two rear doors, there is a rectangular plate with a yellow background and black text\. This is the license plate whose number needs to be read\. The red bounding box should be drawn tightly around this rectangular plate at the license plate position on the back of the police van\.</think\><image\_start\> Generated Image: ![[Uncaptioned image]](https://arxiv.org/html/2607.03748v1/Figs/case/wo_nft.png) <image\_end\><think\>From the cropped view of the license plate area, I can now examine the characters more closely\. The yellow plate is visible in the left\-center portion of the image\. Reading the black characters from left to right, the first group appears to be “GI01” and the second group reads “EPD”\. The character shapes are fairly clear — the first letter has a distinct curve consistent with “G” rather than “O”, and the second character is a straight vertical stroke matching “I”\. The digits in the middle appear to be “01” based on their rounded and straight forms\. Therefore the full registration reads “GI01 EPD”\. Matching this against the available options, this corresponds to option \(A\)\.</think\><answer\>A</answer\>

##### Analysis\.

This case reveals several observations about the role of image quality in TIT\-CoT reasoning:

\(1\)BRAID produces diverse and faithful visual intermediates\.The two BRAID rollouts adopt distinct visual strategies: Rollout 1 generates a tightly cropped view of the license plate for precise character reading, while Rollout 2 produces a full\-scene image with the plate region highlighted\. Both yield high\-fidelity images where “OU68 EPD” is clearly legible, directly supporting downstream text reasoning\.

\(2\)Withoutrvisr^\{\\text\{vis\}\}, images contain hallucinated artifacts\.The w/orvisr^\{\\text\{vis\}\}variant generates an image containing a hallucinated second license plate \(“E\-EI SV”\) alongside the actual plate\. Although the model correctly identifies the relevant plate here, the fabricated visual content introduces unnecessary ambiguity that would likely mislead reasoning on harder instances\.

\(3\)WithoutℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}, images degrade and mislead reasoning\.The w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}variant generates a distorted, blurry crop where the plate characters are garbled\. The text model hallucinates “GI01 EPD” from the corrupted visual evidence and arrives at the correct answer through faulty reasoning \(incorrectly mapping “GI01 EPD” to option A\)\. The reasoning chain is fundamentally unreliable: the model cannot distinguish correct from fabricated visual evidence\.

\(4\)Correct answers can mask fragile reasoning\.All variants arrive at the correct answer, yet only BRAID produces a trajectory where the visual evidence genuinely supports the conclusion\. The ablated variants achieve correctness*despite*their visual intermediates, not*because*of them, a distinction that manifests as lower robustness under Maj@nnevaluation \(Section[5\.1](https://arxiv.org/html/2607.03748#S5.SS1)\)\.

### Appendix DText\-to\-Image Generation Quality

Beyond interleaved reasoning, we examine whether the RL training in BRAID also improves standalone text\-to\-image generation quality\. We compare generations from BAGEL, SFT, BRAID w/orvisr^\{\\text\{vis\}\}, BRAID w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}, and the full BRAID on three prompts that test different aspects of instruction following: spatial compositionality, numerical accuracy, and creative concept binding\.

Figure 8:Text\-to\-image generation comparison across prompts testing spatial compositionality \(Row 1\), numerical accuracy \(Row 2\), and creative concept binding \(Row 3\)\.##### Analysis\.

The generation results in Figure[8](https://arxiv.org/html/2607.03748#A4.F8)reveal that BRAID’s RL training yields improvements in both image quality and instruction adherence:

\(1\)Spatial compositionality\.For “A cat sitting inside a fishbowl”, BAGEL, SFT, and w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}all place the cat*next to*the fishbowl rather than inside it, failing to satisfy the spatial relation specified by the prompt\. Both w/orvisr^\{\\text\{vis\}\}and BRAID correctly render the cat inside the bowl, but BRAID additionally preserves higher visual fidelity with a fish coexisting in the scene\.

\(2\)Numerical accuracy\.For “Exactly five balloons”, BAGEL generates approximately eight, SFT produces four, w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}produces an over\-saturated chaotic scene with many balloons, and w/orvisr^\{\\text\{vis\}\}generates seven to eight\. Only BRAID correctly generates exactly five balloons, demonstrating that the joint RL objective improves precise counting ability\.

\(3\)Creative concept binding\.For “Hourglass with galaxies as sand”, BAGEL, SFT, and w/orvisr^\{\\text\{vis\}\}all produce an hourglass with regular sand and merely place it against a galaxy background, failing to bind the “galaxies as sand” concept\. The w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}variant shows a galaxy swirl inside the hourglass, and BRAID further improves this with a clearly visible galaxy serving as the flowing sand material\.

These observations suggest that the vision\-thinking process rewardrvisr^\{\\text\{vis\}\}provides optimization direction that transfers beyond the interleaved reasoning setting to general text\-to\-image generation\. The reward signal, which evaluates visual correctness and faithfulness during TIT\-CoT training, implicitly teaches the image branch to better follow compositional instructions\. Notably, w/oℒNFT\\mathcal\{L\}\_\{\\text\{NFT\}\}consistently shows the weakest instruction following among RL\-trained variants, confirming that without the DiffusionNFT gradient pathway, the image branch cannot internalize the reward\-driven improvements\.

Similar Articles

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

Hugging Face Daily Papers

AVR is an adaptive visual reasoning framework that dynamically selects optimal reasoning formats to reduce token usage by 50-90% while maintaining accuracy in visual reasoning tasks. The method addresses reasoning path redundancy by decomposing visual reasoning into three cognitive functions and using FS-GRPO training to encourage efficient format selection.