Which Source Wins? Task-Dependent Reliance in Vision-Language Models
Summary
This paper studies how vision-language models shift their reliance between image and text sources when one modality is degraded, revealing task-dependent behavior and introducing a new conflict benchmark for evaluation.
View Cached Full Text
Cached at: 08/19/26, 09:52 AM
# Which Source Wins? Task-Dependent Reliance in Vision–Language Models
Source: [https://arxiv.org/html/2608.17205](https://arxiv.org/html/2608.17205)
Rodela GhoshAviral Gupta11footnotemark:1Affiliation:University of South FloridaEmail:[aviralgupta@usf\.edu](mailto:)Guangjing WangAffiliation:University of South FloridaEmail:[guangjingwang@usf\.edu](mailto:)
###### Abstract
Vision\-language models \(VLMs\) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them\. We study this*modality reallocation*with a controlled setup: we degrade either the image or the text across four levels of legibility while keeping the other clean, and track how the model’s preference changes\. We build conflicts from GSM8K and SVAMP by pairing the rendered image of one arithmetic problem with the text of another, so the two sources support different answers\. We also introduce ChartQA\-Conflict, a manually reviewed benchmark of 229 chart\-report conflicts with matched chart and table\-image representations\. We evaluate six open\-weight VLMs using both generated answers and a length\-normalized conditional log\-likelihood margin\. On GSM8K and SVAMP, five of six models shift more strongly away from degraded text than from degraded images\. On ChartQA\-Conflict, all six likelihood\-scored models exhibit the opposite pattern, shifting more strongly away from the degraded visual source\. This reversal persists after calibrating for unimodal accuracy loss and after replacing charts with plain table images\. Two frontier API models, GPT\-5\.6\-Luna and Gemini\-3\.5\-Flash, behaviorally replicate the ChartQA\-Conflict reversal, with GPT\-5\.6\-Luna also matching the arithmetic direction\. These results show that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings\. The source code is available at[https://github\.com/Ro\-netizen004/multimodal\-arbitration\-artifact](https://github.com/Ro-netizen004/multimodal-arbitration-artifact)\.
## 1Introduction
Humans integrate multisensory cues and down\-weight noisier sources when cues conflict[8](https://arxiv.org/html/2608.17205#bib.bib15);[9](https://arxiv.org/html/2608.17205#bib.bib16)\. Vision\-language models \(VLMs\) likewise integrate visual and textual inputs, yet it remains unclear whether they appropriately recalibrate their reliance on each modality when one becomes less reliable\.
Model inputYou are given two sources describing a math problem\. Source A is the attached image\. Source B is the text below\. Solve the problem step by step and end with \#\#\#\# <answer\>\.Source A — attached imageSource B — textA robe takes 2 bolts of blue fiber and half that much white fiber\. How many bolts in total does it take?Figure 1:The VLM is prompted to solve a problem based on two independent information sources\.Source Ais an image, which renders one problem;Source Bis the text of a*different*problem\.Prior work shows that VLMs perform worse when text\-based problems are shown in an image rather than provided as plain text[35](https://arxiv.org/html/2608.17205#bib.bib4);[34](https://arxiv.org/html/2608.17205#bib.bib5);[19](https://arxiv.org/html/2608.17205#bib.bib6)\. This performance gap has been attributed in part to the additional challenge of extracting information from images before reasoning over the recovered content[30](https://arxiv.org/html/2608.17205#bib.bib7)\. However, existing evaluations typically assess visual and textual inputs in isolation, which cannot determine which modality a model relies on when both are presented but provide conflicting evidence\. In this paper, we ask: when one source becomes harder to read, does a VLM shift its reliance toward the other, and does it shift by the same amount whether we degrade the image or the text? An asymmetric shift would indicate that modality reweighting depends not only on degradation levels but also on which modality is degraded and how the model is evaluated in a specific task\.
To answer the question, we build multiple datasets for evaluation\. \(i\) We create controlled image\-text conflicts using GSM8K\([5](https://arxiv.org/html/2608.17205#bib.bib1)\)and SVAMP\([25](https://arxiv.org/html/2608.17205#bib.bib17)\)arithmetic datasets\. Specifically, we render an arithmetic problem from plain text into a picture\. Then, we pair the rendered image of one arithmetic problem with the text of another, so the two sources support different answers, as shown in Figure[1](https://arxiv.org/html/2608.17205#S1.F1)\. We define four legibility levels and then degrade either the image or the text into different legibility levels while keeping the other source clean\. To avoid implying a primary modality, we refer to the inputs as Source A and Source B and counterbalance these labels across items\. \(ii\) We further construct ChartQA\-Conflict from ChartQA[20](https://arxiv.org/html/2608.17205#bib.bib3), pairing each chart with a textual report that supports a conflicting answer\. This dataset tests whether the observed modality\-reliance patterns generalize beyond rendered\-text images\.
We evaluate six open\-weight VLMs in controlled image\-text conflicts across three benchmarks\. The experiments reveal a task\-dependent asymmetry: on the arithmetic benchmarks \(GSM8K and SVAMP\), five of six models shift more strongly away from degraded text than from degraded images\. By contrast, on ChartQA\-Conflict, all six likelihood\-scored models shift more strongly away from the degraded visual source\. This reversal persists after calibrating for unimodal accuracy loss and after replacing charts with plain table images, indicating that VLMs do not exhibit a fixed preference for text or vision\. Two frontier API models, GPT\-5\.6\-Luna and Gemini\-3\.5\-Flash, behaviorally replicate the ChartQA\-Conflict reversal, indicating the pattern is not confined to open\-weight models\. Instead, modality reallocation depends on the task, evidence structure, model, and, to a lesser extent, prompt framing\.
In summary, our contributions are:
- •A new problem formulation and conflict\-based evaluation datasets\.We introduce the problem of*modality reallocation*: how VLMs redistribute their reliance between visual and textual sources when the two conflict and one becomes harder to read\. We construct controlled conflicts from GSM8K and SVAMP and introduce ChartQA\-Conflict, a manually reviewed benchmark of 229 chart\-report conflicts with matched chart and table\-image representations\.
- •A controlled framework for measuring source reliance\.We degrade either the visual or textual source across four legibility levels while keeping the other source clean, counterbalance source labels, and quantify preference using both generated answers and a length\-normalized conditional log\-likelihood margin\. This design enables within\-item comparisons of how strongly models shift away from each degraded modality\.
- •Evidence that modality reallocation is setting\-dependent rather than fixed\.Across arithmetic conflicts, five of six models shift more strongly away from degraded text, whereas all six CLL\-scored models show the opposite pattern on ChartQA\-Conflict\. This reversal persists after calibration for unimodal accuracy loss and after replacing charts with table images, and is behaviorally replicated by two frontier API models \(GPT\-5\.6\-Luna and Gemini\-3\.5\-Flash\), showing that VLM reliance depends on the task, evidence structure, and model rather than reflecting a universal preference for text or vision\.
## 2Related Work
##### VLM modality reliance preference\.
VLMs often underperform when problems are presented as images rather than native text[35](https://arxiv.org/html/2608.17205#bib.bib4);[34](https://arxiv.org/html/2608.17205#bib.bib5);[19](https://arxiv.org/html/2608.17205#bib.bib6)\. However, these studies compare performance across input formats and do not examine which modality a model follows when visual and textual sources provide conflicting answers\. More broadly, VLMs are documented to under\-use visual information and default to textual cues[13](https://arxiv.org/html/2608.17205#bib.bib8)\. However,[12](https://arxiv.org/html/2608.17205#bib.bib9)pair images with contradictory captions and show, at a fixed conflict level, that source preference is systematic, model\-dependent, and reflected in internal representations\.[23](https://arxiv.org/html/2608.17205#bib.bib21)studies conflicts between visual evidence and a model’s stored knowledge using logit\-based analysis and causal interventions\. SimpleOCR\([26](https://arxiv.org/html/2608.17205#bib.bib20)\)also renders questions as images to study the modality\-utilization gap\. Yet, SimpleOCR uses rendering to improve visual reading rather than to induce cross\-modal conflict\. Our study extends this line of work in three ways\. First, we examine conflicts in multi\-step reasoning tasks rather than in simple object\-recognition tasks\. Second, we make either the image or the text harder to read across four legibility levels and track how the model’s preference changes for the same item\. Third, we use counterbalanced Source A/B labels so that neither source is framed as the primary input\.
##### Properties of conflicting inputs\.
Prior studies vary the properties of conflicting evidence\. Some degrade the textual information[6](https://arxiv.org/html/2608.17205#bib.bib10)\. Others change reasoning difficulty or uncertainty[27](https://arxiv.org/html/2608.17205#bib.bib11);[37](https://arxiv.org/html/2608.17205#bib.bib13)\. A further line asks whether visual evidence is internally represented even when the answer follows the text[22](https://arxiv.org/html/2608.17205#bib.bib12), and whether chain\-of\-thought traces can look visually grounded while in fact following the text[28](https://arxiv.org/html/2608.17205#bib.bib14)\. Each of these manipulates one source\. We instead degrade the image and the text in item\-matched degradation arms with nominally aligned levels, and compare how source preference changes in each\. We do not assume the two degradation levels are equally severe, and we design a secondary analysis that adjusts for severity using the measured drop in unimodal accuracy\. CMC\-Bench\([3](https://arxiv.org/html/2608.17205#bib.bib18)\)creates image\-text conflicts using ChartQA and MMMU[36](https://arxiv.org/html/2608.17205#bib.bib29), but varies the conflict type to evaluate accuracy and abstention\([21](https://arxiv.org/html/2608.17205#bib.bib19)\)\. In contrast, we keep each conflict fixed, degrade each source separately across four levels, and measure within\-item shifts in source preference using generated answers and length\-normalized likelihood margins\.
Conflict constructionDegrade one sourceOutcome measuresVisual sourceaIa\_\{I\}rendered problem or chartTextual sourceaTa\_\{T\}problem text or reportDegrade visualtextual stays clean⋅\\cdot4 levelsDegrade textualvisual stays clean⋅\\cdot4 levelsGenerated choicewhich answer it producesCLL marginlength\-normalized log\-likelihoodRole\-neutral promptvisual and textual sources labeled Source A / B, counterbalanced across itemsTwo settings: image \+ text problem \(GSM8K, SVAMP\)⋅\\cdotchart \+ report \(ChartQA\-Conflict\)
Figure 2:Experimental logic\. We pair the rendered image of one problem with the text of another so the two sources support different answers \(aIa\_\{I\}vs\.aTa\_\{T\}\)\. A role\-neutral prompt labels them Source A/B, counterbalanced across items\. From the clean pair we degrade one source at a time across four legibility levels while the other stays clean, and score every trial by both the generated source choice and the length\-normalized CLL preference margin\.
## 3Methodology
In this section, we evaluate how a model redistributes its reliance between the image and the text when one of them becomes harder to read, which is referred to as*source reallocation*\.
The pipeline is the same across two conflict settings as shown in Figure[2](https://arxiv.org/html/2608.17205#S2.F2)\. In both settings, the model receives an image and text that support different answers\. We degrade one source across four legibility levels while keeping the other clean\. We then measure source preference using both the generated answer and a length\-normalized CLL margin \(§[3\.3](https://arxiv.org/html/2608.17205#S3.SS3)\)\. For each model, we compare whether preference shifts more strongly when the image is degraded or when the text is degraded\.
The two settings differ in how the conflict is constructed\. In*rendered\-text conflict*on GSM8K and SVAMP, the image and text contain different arithmetic problems in natural language\. In*natural\-visual conflict*on ChartQA\-Conflict, both sources address the same question, but the chart and the accompanying report support different answers\. The same degradation, measurement, and comparison procedures \(§[3\.3](https://arxiv.org/html/2608.17205#S3.SS3)\) apply to both settings\.
### 3\.1Models
We evaluate six open\-weight VLMs spanning four families and 2–8B parameters: Qwen2\-VL\-2B\([31](https://arxiv.org/html/2608.17205#bib.bib22)\), Qwen2\.5\-VL\-7B\([2](https://arxiv.org/html/2608.17205#bib.bib23)\), LLaVA\-1\.6\-Mistral\-7B\([17](https://arxiv.org/html/2608.17205#bib.bib27);[18](https://arxiv.org/html/2608.17205#bib.bib28)\), LLaVA\-OneVision\-7B\([15](https://arxiv.org/html/2608.17205#bib.bib24)\), Idefics3\-8B\([14](https://arxiv.org/html/2608.17205#bib.bib25)\), and Phi\-3\.5\-Vision\-4B\([1](https://arxiv.org/html/2608.17205#bib.bib26)\)\. All six models are evaluated using conditional log\-likelihood\. Generated\-answer evaluation is also conducted for all six, although Phi\-3\.5\-Vision yields too few decidable ChartQA\-Conflict responses for reliable behavioral analysis\. A seventh model, InternVL2\-8B\([4](https://arxiv.org/html/2608.17205#bib.bib2)\), and two frontier API models, GPT\-5\.6\-Luna and Gemini\-3\.5\-Flash, lack the continuation\-scoring interface the CLL margin requires and are used for behavioral \(generated\-answer\) evidence only, reported separately \(§[4\.6](https://arxiv.org/html/2608.17205#S4.SS6), Appendix[D](https://arxiv.org/html/2608.17205#A4)\)\. Models are loaded without quantization and decoded usingbfloat16where supported\. Full software versions are in Appendix[A](https://arxiv.org/html/2608.17205#A1)\.
### 3\.2Dataset Construction
#### 3\.2\.1Rendered image conflict
We pair the rendered image of one natural language\-based math problem with the native text of a different problem\. Since the two problems support different answers, the model’s response indicates which source it followed\. Specifically, image problemiiis paired with text problem\(i\+1\)modN\(i\+1\)\\bmod N, whereNNis the number of benchmark problems\. Equivalently, text problemiiis paired with image problem\(i−1\)modN\(i\-1\)\\bmod N\. Figure[1](https://arxiv.org/html/2608.17205#S1.F1)shows a worked example, where a response ofaI=18a\_\{I\}\{=\}18is image\-following andaT=3a\_\{T\}\{=\}3is text\-following\. For odd\-indexed items, the Source A/B labels are swapped, so neither modality is always Source A\.
We exclude pairs whose two problems have the same answer to avoid confusing the follow\-up preference analysis\. Images are rendered as 900\-pixel\-wide PNGs using DejaVu Sans, black text on a white background\. Each image contains the same information as its native\-text version, so only the input format differs\. We also evaluate each problem separately in text\-only and image\-only conditions\.
#### 3\.2\.2Natural visual conflict
##### Construction\.
Each ChartQA\-Conflict item pairs a native chart supporting answeraIa\_\{I\}with an accompanying textual report giving counterfactual evidence for a different answeraTa\_\{T\}to the same question \(Figure[5](https://arxiv.org/html/2608.17205#A9.F5)\); unlike the rendered\-image setting \(§[3\.2\.1](https://arxiv.org/html/2608.17205#S3.SS2.SSS1)\), the chart carries native visual structure \(bars, axes, legends\)\. We build items from the official ChartQA test set and source tables\([20](https://arxiv.org/html/2608.17205#bib.bib3)\), choosing eachaTa\_\{T\}to preserve the original answer’s type and unit and stay semantically plausible\. Two of the authors, both senior undergraduates, independently reviewed all 230 items, confirming that the report supportsaTa\_\{T\}, thataTa\_\{T\}is a valid answer, and thataT≠aIa\_\{T\}\\neq a\_\{I\}after normalization; a post\-run audit removed one item whose report did not support its answer, leaving 229 conflicts \(excluded identically across models, arms, and levels\)\. Full details are in Appendix[I](https://arxiv.org/html/2608.17205#A9)\.
##### Prompt\.
ChartQA\-Conflict uses a role\-neutral prompt: the shared question, a statement that the model is given “two conflicting evidence sources” with “neither source privileged,” then the chart and report under counterbalanced Source A/B labels \(A first\)\. Template in Appendix I\.
##### Chart\-versus\-table control\.
To test whether the reversal depends on the chart’s graphical encoding rather than on any image\-format visual, we replace the chart with a plain\-table image of the same official source table, holding the question, report, conflict, and degradation ladder fixed\. Both representations present identical facts, so both still supportaIa\_\{I\}; only the visual form changes\. Details in Appendix I\.
Legibility degradation ladders \(same conflict item\)Figure 3:Legibility ladders for the same conflict item\.Top \(image arm\):the rendered image \(Source A\) is degraded from clean \(L0\) to heavy blur\+\+noise \(L5\) with the text clean\.Bottom \(text arm\):the conflicting text \(Source B\) is corrupted at the character level with the image clean\. L0/L2/L4/L5 mark increasing corruption within each arm, matched by level index, not absolute severity\.
### 3\.3Matched Degradation and Measurement
We test whether models shift away from a source as that source becomes harder to read\. Each experiment therefore has two parallel degradation arms\. In the image\-degradation arm, the image is degraded while the text remains clean\. In the text\-degradation arm, the text is degraded while the image remains clean\. Across arms and degradation levels, we keep the examples, source\-supported answers, models, decoding procedure, and scoring method fixed\.
##### Legibility levels\.
We evaluate four prespecified levels: clean \(L0\), light degradation \(L2\), moderate degradation \(L4\), and heavy degradation \(L5\)\. These labels come from a larger image\-corruption ladder\([10](https://arxiv.org/html/2608.17205#bib.bib30)\)used earlier in the project\. We omit L1 and L3 because they contain JPEG\-only corruptions with no direct text counterpart\. Thus, L0/L2/L4/L5 provide a shared clean\-to\-heavy progression for both modalities\.
##### Degradation\.
We evaluate four aligned degradation levels: clean \(L0\), light \(L2\), moderate \(L4\), and heavy \(L5\)\. Images are degraded using increasing Gaussian blur and pixel noise, while text is degraded by randomly deleting or replacing non\-whitespace characters, with whitespace preserved to keep word boundaries visible\. Figure[3](https://arxiv.org/html/2608.17205#S3.F3)illustrates both degradation ladders for one conflict item\. The same image operations apply to the ChartQA charts\. All corruption is deterministic and seeded\. Every model receives identical degraded inputs and the same corrupted text is used for the generation and CLL scoring\. The exact blur radii, noise levels, and corruption rates per\-level are in Appendix[B](https://arxiv.org/html/2608.17205#A2)\.
##### Behavioral source choice\.
At each level, we classify the generated response as image\-following, text\-following, neither, ambiguous, or invalid\. A trial is*decidable*when its extracted answer matches exactly one source\-supported answer\. Neither, ambiguous, and invalid responses are retained and reported but excluded from source\-preference estimates; we also report the decidable count because degradation can affect both source choice and answer attributability\.
We first extract the answer following the requested\#\#\#\#delimiter\. For ChartQA\-Conflict, some models instead return a bare answer on the first line\. When the delimiter is absent, we accept the first nonempty line if that entire line normalizes to a single valid answer\. We do not search explanations or reasoning traces for embedded numbers\. For GSM8K and SVAMP, the extracted numeric answer is rounded to the nearest integer\. For ChartQA\-Conflict, we normalize superficial formatting differences, including commas, signs, decimal notation, percentages, currency symbols, compatible units, and yes/no variants, and then require an exact match to one candidate\. Equal normalized candidates or responses matching both are ambiguous, and responses with no valid extracted answer are invalid\. We use no fuzzy matching\.
##### Continuous source\-preference margin\.
Generated choice can be insensitive when the model already prefers one source at the clean level L0\. This is common in the image arm, where several models start out favoring text even as their underlying preference keeps shifting\. We therefore add a continuous measure from the probabilities of the two candidate answers\. For a conflict trial, letaTa\_\{T\}andaIa\_\{I\}be the text\- and image\-supported answers,ccthe shared multimodal prompt, and\|aT\|,\|aI\|\|a\_\{T\}\|,\|a\_\{I\}\|their token lengths:
m\(c\)=1\|aT\|logp\(aT∣c\)−1\|aI\|logp\(aI∣c\)\.m\(c\)=\\frac\{1\}\{\|a\_\{T\}\|\}\\log p\(a\_\{T\}\\mid c\)\-\\frac\{1\}\{\|a\_\{I\}\|\}\\log p\(a\_\{I\}\\mid c\)\.\(1\)Eachlogp\(a∣c\)\\log p\(a\\mid c\)sums the teacher\-forced log\-probabilities of the candidate tokens, giving a mean log\-probability per answer token \(nats/token\); a positive margin favors the text answer, a negative one the image answer\. Length normalization reduces but does not remove the advantage of shorter answers\([11](https://arxiv.org/html/2608.17205#bib.bib31)\); Appendix[H](https://arxiv.org/html/2608.17205#A8)shows that across GSM8K, SVAMP, and ChartQA\-Conflict the exponent rescales the asymmetry’s magnitude but preserves its sign and significance for every model\. Only answer tokens after the fixed\#\#\#\#delimiter are scored, with no chain\-of\-thought, so the margin measures direct\-answer preference, not reasoning\-trace attribution\. As a validity check,sign\(m\)\\operatorname\{sign\}\(m\)agrees with the generated source choice on 75\.4% of 26,893 decidable model–item–level observations \(six models, GSM8K and SVAMP, both degradation arms\)\. More detailed analysis is in Appendix[C](https://arxiv.org/html/2608.17205#A3)\. Since the two agree and the margin is graded and available for all six models, we adopt it as the primary measure for the reallocation analysis and use generated choice as a behavioral cross\-check\.
##### Comparing the two arms\.
Letmi,L\(I\)m^\{\(I\)\}\_\{i,L\}andmi,L\(T\)m^\{\(T\)\}\_\{i,L\}denote the margins for itemiiat levelLLin the image\- and text\-degradation arms\. We orient both changes toward the source that remains clean:
RI,i=mi,L5\(I\)−mi,L0\(I\),RT,i=mi,L0\(T\)−mi,L5\(T\)\.R\_\{I,i\}=m^\{\(I\)\}\_\{i,L5\}\-m^\{\(I\)\}\_\{i,L0\},\\qquad R\_\{T,i\}=m^\{\(T\)\}\_\{i,L0\}\-m^\{\(T\)\}\_\{i,L5\}\.SoRI,i\>0R\_\{I,i\}\>0means itemiimoves toward the text\-supported answer as the image degrades, andRT,i\>0R\_\{T,i\}\>0means it moves toward the image\-supported answer as the text degrades\. The within\-item arm asymmetry is
Ai=RT,i−RI,i,A\_\{i\}=R\_\{T,i\}\-R\_\{I,i\},withAi\>0A\_\{i\}\>0indicating stronger reallocation under text degradation andAi<0A\_\{i\}<0stronger reallocation under image degradation\. For each model we report the medians ofRI,iR\_\{I,i\},RT,iR\_\{T,i\}, andAiA\_\{i\}; because the median ofAiA\_\{i\}is computed within item, it need not equal the difference of the two separately reported medians\. Statistical testing is described in §[3\.5](https://arxiv.org/html/2608.17205#S3.SS5)\.
### 3\.4Benchmarks
##### Matched\-degradation benchmarks\.
The primary rendered\-arithmetic experiments use the complete GSM8K test set\([5](https://arxiv.org/html/2608.17205#bib.bib1)\)\(1,319 problems\) and a fixed 300\-item SVAMP subset\([25](https://arxiv.org/html/2608.17205#bib.bib17)\)\. In each benchmark, image itemiiis paired with text item\(i\+1\)modN\(i\+1\)\\bmod N, with wraparound at the end, and this pairing is fixed across models, arms, levels, generation, and CLL scoring\. The two paired answers are distinct under the numerical matching rule for 1,304 GSM8K pairs and 296 SVAMP pairs; equal\-answer pairs are retained in the saved outputs but marked ambiguous and excluded from the decidable\-preference denominator\. The same pairings are used in both degradation arms\.
ChartQA\-Conflict provides a separate natural\-visual evaluation, using an original chart and an evidence\-bearing report that support different answers to the same question\. After excluding one item that failed the post\-run entailment audit, the analysis contains 229 valid conflicts; its construction and attribution are described in §[3\.2\.2](https://arxiv.org/html/2608.17205#S3.SS2.SSS2)\.
##### Data release\.
### 3\.5Statistical Analysis
##### Within\- and between\-arm changes\.
CLL results are matched by item ID\. For each item and arm we compute the L0\-to\-L5 margin change, summarize the item\-level changes by their median, and test whether the paired differences are centered at zero with a two\-sided Wilcoxon signed\-rank test\([32](https://arxiv.org/html/2608.17205#bib.bib35)\), a 95% paired\-bootstrap CI \(10,000 resamples\)\([7](https://arxiv.org/html/2608.17205#bib.bib36)\), and a paired randomization test \(10,000 draws; fixed seed\)\. The primary cross\-arm comparison is the within\-item asymmetryAiA\_\{i\}\(§[3\.3](https://arxiv.org/html/2608.17205#S3.SS3)\), tested with the same procedures;Ai\>0A\_\{i\}\>0means stronger reallocation under text degradation\.
##### Adjustment for unimodal task loss\.
Image and text degradation at the same nominal level may not remove the same amount of usable information\. We therefore estimate degradation severity using each model’s accuracy when it receives only the source being degraded\. For modelMM, channelcc, and levelLL, we calculate the proportional accuracy loss relative to the clean condition,ℓM,c,L\\ell\_\{M,c,L\}\. We then fit, separately for each model and degradation arm, a line relating this loss to the median shift in preference toward the source that remains clean\. The slopebM,cb\_\{M,c\}therefore measures reallocation per unit of lost single\-modality accuracy\. We then compare the text\- and image\-degradation slopes:
DM=bM,T−bM,I\.D\_\{M\}=b\_\{M,T\}\-b\_\{M,I\}\.A positiveDMD\_\{M\}indicates stronger reallocation when text is degraded than when the image is degraded, after accounting for the measured accuracy loss\. Confidence intervals are obtained by bootstrapping matched items\. We treat this as a secondary analysis because each slope is estimated from only three degraded levels and single\-modality accuracy is itself an imperfect estimate of source usability\. Additional pooled\-regression and character/OCR\-survival analyses\([29](https://arxiv.org/html/2608.17205#bib.bib32)\)are reported in Appendix[E](https://arxiv.org/html/2608.17205#A5)\.
## 4Results
Figure[4](https://arxiv.org/html/2608.17205#S4.F4)summarizes the findings across different settings: on the arithmetic conflicts \(GSM8K, SVAMP\), the paired asymmetry is positive for five of six models, whereas on ChartQA\-Conflict it reverses to negative for all six\. Table[1](https://arxiv.org/html/2608.17205#S4.T1)reports the full per\-modelRIR\_\{I\},RTR\_\{T\}, and paired asymmetryAAwith 95% bootstrap confidence intervals\. We first evaluate the consistency of two measures \(§[4\.1](https://arxiv.org/html/2608.17205#S4.SS1)\), then examine the arithmetic pattern \(§[4\.2](https://arxiv.org/html/2608.17205#S4.SS2)\) and the reversal \(§[4\.5](https://arxiv.org/html/2608.17205#S4.SS5)\)\.
Figure 4:Paired asymmetryA=RT−RIA=R\_\{T\}\-R\_\{I\}per model and setting; whiskers are 95% paired\-bootstrap CIs\.A\>0A\>0indicates a stronger response to text/report degradation,A<0A<0to image/chart degradation\. On GSM8K and SVAMP the asymmetry is positive for five of six models; on ChartQA\-Conflict it reverses to negative for all six\. Full per\-model values are in Table[1](https://arxiv.org/html/2608.17205#S4.T1)\.Table 1:Reliance on the clean source when the other is degraded \(neutral prompt\)\.RIR\_\{I\}/RTR\_\{T\}is the median reallocation toward the clean text/image when the image/text is degraded \(chart/report on ChartQA\-Conflict\);AAis the median within\-item contrastRT,i−RI,iR\_\{T,i\}\-R\_\{I,i\}, withA\>0A\>0\(A<0A<0\) indicating stronger reallocation under text \(image\) degradation\. As medians are not additive,AAneed not equalRT−RIR\_\{T\}\-R\_\{I\}\. Intervals are paired\-bootstrap 95% CIs \(10,000 resamples;n=229n=229for ChartQA\-Conflict\)\.### 4\.1Behavioral Choice and CLL Consistency
The two measures are moderately consistent\. Across six models, the CLL margin favors the same source as the generated answer on75\.4%75\.4\\%of26,89326\{,\}893decidable trials \(95%95\\%CI:74\.474\.4–76\.4%76\.4\\%; Appendix[C](https://arxiv.org/html/2608.17205#A3)\), and both shift in the same direction as degradation increases\. We therefore base the reallocation analysis on the CLL margin—graded and available for all six models—and use generated answers as a behavioral cross\-check, since generated choices can saturate near the ceiling\.
### 4\.2Text Degradation Produces Stronger Reallocation on Arithmetic Conflicts
Recall that a positive item\-level contrastAi=RT,i−RI,iA\_\{i\}=R\_\{T,i\}\-R\_\{I,i\}means reallocation is stronger under text than image degradation\. On GSM8K the median asymmetry is positive for all six models, and for five the paired\-bootstrap interval excludes zero \(medianAiA\_\{i\}up to\+3\.877\+3\.877nats/token\); Qwen2\.5\-VL\-7B is the exception, with a near\-zero estimate \(A=\+0\.059A=\+0\.059\) whose interval reaches zero\. The same five models show the pattern on SVAMP, while Qwen2\.5\-VL\-7B reverses \(A=−0\.731A=\-0\.731, 95% CI\[−1\.103,−0\.022\]\[\-1\.103,\-0\.022\]\)\. The arithmetic result is therefore a consistent five\-of\-six cross\-model pattern, not a universal property of every model\.
### 4\.3Calibration by Unimodal Task Loss
Calibrating each arm’s reallocation by the model’s proportional loss of unimodal accuracy reproduces the same task\-dependent pattern \(Appendix[E](https://arxiv.org/html/2608.17205#A5), Tables[7](https://arxiv.org/html/2608.17205#A5.T7)and[8](https://arxiv.org/html/2608.17205#A6.T8)\)\. On GSM8K and SVAMP, five of six models have a positive slope differenceD=bT−bID=b\_\{T\}\-b\_\{I\}with bootstrap intervals excluding zero; Qwen2\.5\-VL\-7B is again the arithmetic exception, withD<0D<0on both\. On ChartQA\-Conflict the direction reverses: all six models haveD<0D<0with intervals excluding zero\. The reversal survives adjustment for degradation severity, supporting task\- and format\-dependence rather than a fixed modality preference\. We treat this as secondary, since each slope uses only three degraded levels and conditions on the accuracy estimates\.
### 4\.4Prompt Framing Modulates Effect Size
The asymmetry remains positive across five models with matched GSM8K CLL results \(Appendix[K](https://arxiv.org/html/2608.17205#A11), Table[12](https://arxiv.org/html/2608.17205#A11.T12)\) under both the original and role\-neutral prompts, showing that framing alone does not create the arithmetic asymmetry\. Neutral framing changes its magnitude for four models \(p<10−3p<10^\{\-3\}\) but not Phi\-3\.5\-Vision \(p=\.112p=\.112\), with effects varying in direction\. The largest change is for Qwen2\.5\-VL\-7B, whose median asymmetry falls from\+0\.813\+0\.813to\+0\.060\+0\.060\. Thus, framing modulates the effect in a model\-dependent way without explaining the broader pattern\.
### 4\.5The Asymmetry Reverses for Natural Visual Evidence
Unlike the arithmetic conflicts, where both sources are essentially text \(one just shown as an image\), ChartQA\-Conflict pits a real chart against a written report, and this flips the result \(Figure[4](https://arxiv.org/html/2608.17205#S4.F4), ChartQA series; per\-model values in Table[1](https://arxiv.org/html/2608.17205#S4.T1)\)\. Across the 229 items all six models show a negative asymmetry \(−0\.280\-0\.280to−2\.175\-2\.175nats/token\) with every confidence interval below zero: each moves away from the chart more strongly when the chart is degraded than away from the report when the report is degraded\. This reverses the arithmetic pattern for all six models and persists after adjusting for how much each degradation lowers accuracy \(§[4\.3](https://arxiv.org/html/2608.17205#S4.SS3)\)\.
The reversal also holds behaviorally: generated answers are negative for all five decidable open models \(InternVL2 the exception; Phi\-3\.5\-Vision has too few decidable items\), and both frontier models replicate it \(§[4\.6](https://arxiv.org/html/2608.17205#S4.SS6)\)\. We treat the six\-model CLL results as primary and these behavioral results as converging support \(Appendices[J](https://arxiv.org/html/2608.17205#A10)and[D](https://arxiv.org/html/2608.17205#A4)\)\.
##### Graphical encoding alone does not explain the reversal\.
To see whether the reversal comes from the chart’s graphical form \(its bars, axes, and plotted marks\), we replace each chart with a table image showing the same numbers, keeping the question, report, conflict, and degradation levels unchanged\. All six models still show a negative asymmetry with the table image \(Table[9](https://arxiv.org/html/2608.17205#A7.T9), Appendix[G](https://arxiv.org/html/2608.17205#A7)\), so the chart’s graphical marks are not necessary for the reversal\. This does not identify the cause on its own, since the table image still differs from rendered text in layout and the effect is inconsistent across models, but it shows the reversal survives removing the chart’s graphical encoding\.
### 4\.6Evaluation on Frontier Larger Models
We also evaluate generated\-answer behavior from GPT\-5\.6\-Luna and Gemini\-3\.5\-Flash\. As these APIs expose no teacher\-forced scores, we use exact answer attribution only and treat it as supporting evidence, separate from the primary six\-model CLL analysis\. On the 229 ChartQA\-Conflict items both models reproduce the natural\-visual direction \(negative paired asymmetry,p<\.001p<\.001\); on 300 role\-neutral GSM8K conflicts GPT\-5\.6\-Luna instead matches the open\-model arithmetic direction \(positive,p<\.001p<\.001\)\. Full statistics, trajectories, and the clean\-endpoint robustness check are detailed in Appendix[D](https://arxiv.org/html/2608.17205#A4)\.
## 5Conclusion
We introduced a controlled framework for studying*modality reallocation*in VLMs by presenting conflicting visual and textual evidence, degrading each source in turn, and measuring source preference using both generated answers and conditional log\-likelihood margins\. Across GSM8K, SVAMP, and the newly constructed ChartQA\-Conflict benchmark, we find that VLMs do not reallocate reliance symmetrically\. Five of six models shift more strongly away from degraded text on arithmetic conflicts, whereas all six CLL\-scored models shift more strongly away from degraded visual evidence on ChartQA\-Conflict\. This reversal persists after calibration for unimodal accuracy loss and under the chart\-to\-table control, showing that modality reliance is not a fixed bias toward text or vision but a setting\-dependent behavior shaped by the task, evidence structure, model, and prompt framing\. By providing a new problem formulation, curated conflict datasets, and a reproducible evaluation methodology, this work offers the research community a systematic basis for diagnosing how multimodal models arbitrate between competing sources\.
## Limitations
First, the current conflict construction is designed for tasks with attributable numeric answers\. It does not transfer directly to multiple\-choice benchmarks such as AQuA\-RAT\([16](https://arxiv.org/html/2608.17205#bib.bib37)\), where models may output only an option label, and the selected source cannot always be identified reliably\. Extending the framework to free\-form, categorical, and multiple\-choice outputs will require alternative attribution procedures\.
Second, our model coverage is limited to open\-weight VLMs with 2B–8B parameters because the CLL analysis requires access to token\-level log\-probabilities\. Closed frontier models are therefore excluded from the primary graded analysis\. In addition, the prompt\-framing comparison covers only five models on GSM8K, and the legibility\-adjusted regression contains six model clusters\. The results should thus be interpreted as evidence for the evaluated models rather than as a model\-family\-wide generalization\.
Third, the two conflict settings do not isolate task, representation, and conflict construction independently\. In GSM8K and SVAMP, the visual source is rendered text, whereas ChartQA\-Conflict uses a chart and report that answer the same question\. The chart\-versus\-table control shows that chart\-specific graphical encoding alone does not explain the reversal, but the remaining differences between arithmetic and ChartQA\-Conflict prevent us from identifying a single causal mechanism\.
Fourth, ChartQA\-Conflict is limited to one natural\-visual benchmark and was reviewed by two of the authors, who were not blind to the counterfactual design, rather than independent annotators\. Moreover, the answer\-relevant value appears in the report as a stated figure in the text, listed alongside values for other categories or years, so recovering it requires reading and matching the text to the question but not visual chart\-reading; obtaining the corresponding value from a chart or table may additionally require visual localization and interpretation\. The reversal therefore reflects reallocation between sources with potentially different baseline directness, not a comparison between equally accessible visual and textual evidence\.
Fifth, the visual and textual corruption ladders are matched by nominal level rather than psychometric severity\. We partially address this issue by calibrating reallocation against the accuracy loss produced when each degraded source is presented alone\. However, unimodal accuracy is an approximate measure of source usability, and the calibrated slopes are estimated from three degraded levels\.
Finally, our framework measures forced source arbitration under reduced legibility rather than factual reliability or conflict awareness\. Degradation makes a source harder to read without making its content less correct, and the prompt requires a single answer\. Consequently, we do not test whether models detect the contradiction, express uncertainty, or abstain\. The CLL margin also agrees with generated source choices on 75\.4% of attributable trials, indicating that it is a useful but imperfect proxy for behavioral reliance\.
## Ethical Considerations
This work presents limited direct ethical risk, but its findings should not be interpreted as evidence that any evaluated model is safe or reliable for deployment\. Our experiments examine source reliance under controlled visual and textual degradation\. Gaussian blur and character corruption are simplified proxies for reduced legibility and do not capture the full range of accessibility barriers encountered in practice\. The results may also not generalize to other languages, scripts, or tokenization schemes because our text corruptions are restricted to English alphanumeric characters\.
## Declaration of Generative AI Assistance
The authors used generative AI coding and writing assistants, specifically Claude Code \(Anthropic\) and Codex \(OpenAI\), in a limited, supervised capacity during this work\. Their use was confined to \(i\) software\-engineering support—debugging experiment and analysis scripts, orchestrating cluster jobs, and formatting figures andLaTeXtables—and \(ii\) language editing of author\-written text, including tightening prose, improving clarity, and checking cross\-references and consistency\. The research questions, experimental design, datasets, evaluation methodology, analyses, results, and conclusions were conceived, implemented, and verified by the authors\. All AI\-assisted outputs were reviewed by the authors, who take full responsibility for the entire content of this paper, including any remaining errors\.
## Acknowledgments
We thank the CIS Lab at the University of South Florida for supporting this work through access to GPU computing resources and funding for API\-based model evaluations\. We are also grateful for the technical guidance and discussions that helped us develop and refine our experimental setup\.
## References
- Abdinet al\.\(2024\)M\. Abdin, J\. Aneja, H\. Awadalla, A\. Awadallah, A\. A\. Awan, N\. Bach, A\. Bahree, A\. Bakhtiari, J\. Bao, H\. Behl, A\. Benhaim, M\. Bilenko, J\. Bjorck, S\. Bubeck, M\. Cai, Q\. Cai, V\. Chaudhary, D\. Chen, D\. Chen, W\. Chen, Y\. Chen, Y\. Chen, H\. Cheng, P\. Chopra, X\. Dai, M\. Dixon, R\. Eldan, V\. Fragoso, J\. Gao, M\. Gao, M\. Gao, A\. Garg, A\. D\. Giorno, A\. Goswami, S\. Gunasekar, E\. Haider, J\. Hao, R\. J\. Hewett, W\. Hu, J\. Huynh, D\. Iter, S\. A\. Jacobs, M\. Javaheripi, X\. Jin, N\. Karampatziakis, P\. Kauffmann, M\. Khademi, D\. Kim, Y\. J\. Kim, L\. Kurilenko, J\. R\. Lee, Y\. T\. Lee, Y\. Li, Y\. Li, C\. Liang, L\. Liden, X\. Lin, Z\. Lin, C\. Liu, L\. Liu, M\. Liu, W\. Liu, X\. Liu, C\. Luo, P\. Madan, A\. Mahmoudzadeh, D\. Majercak, M\. Mazzola, C\. C\. T\. Mendes, A\. Mitra, H\. Modi, A\. Nguyen, B\. Norick, B\. Patra, D\. Perez\-Becker, T\. Portet, R\. Pryzant, H\. Qin, M\. Radmilac, L\. Ren, G\. de Rosa, C\. Rosset, S\. Roy, O\. Ruwase, O\. Saarikivi, A\. Saied, A\. Salim, M\. Santacroce, S\. Shah, N\. Shang, H\. Sharma, Y\. Shen, S\. Shukla, X\. Song, M\. Tanaka, A\. Tupini, P\. Vaddamanu, C\. Wang, G\. Wang, L\. Wang, S\. Wang, X\. Wang, Y\. Wang, R\. Ward, W\. Wen, P\. Witte, H\. Wu, X\. Wu, M\. Wyatt, B\. Xiao, C\. Xu, J\. Xu, W\. Xu, J\. Xue, S\. Yadav, F\. Yang, J\. Yang, Y\. Yang, Z\. Yang, D\. Yu, L\. Yuan, C\. Zhang, C\. Zhang, J\. Zhang, L\. L\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, and X\. ZhouPhi\-3 technical report: a highly capable language model locally on your phone\.External Links:2404\.14219,[Link](https://arxiv.org/abs/2404.14219)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Baiet al\.\(2025\)S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang, H\. Zhong, Y\. Zhu, M\. Yang, Z\. Li, J\. Wan, P\. Wang, W\. Ding, Z\. Fu, Y\. Xu, J\. Ye, X\. Zhang, T\. Xie, Z\. Cheng, H\. Zhang, Z\. Yang, H\. Xu, and J\. LinQwen2\.5\-VL technical report\.External Links:2502\.13923,[Link](https://arxiv.org/abs/2502.13923)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Catapang \(2026\)J\. K\. CatapangWhen image and text disagree: cross\-modal evidence conflict in multimodal retrieval\-augmented generation\.InProceedings of the 2nd Workshop on Multimodal Augmented Generation via Multimodal Retrieval \(MAGMaR 2026\),pp\. 1–10\.External Links:[Link](https://aclanthology.org/2026.magmar-main.3/)Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2024\)Z\. Chen, W\. Wang, Y\. Cao, Y\. Liu, Z\. Gao, E\. Cui, J\. Zhu, S\. Ye, H\. Tian, Z\. Liu, L\. Gu, X\. Wang, Q\. Li, Y\. Ren, Z\. Chen, J\. Luo, J\. Wang, T\. Jiang, B\. Wang, C\. He, B\. Shi, X\. Zhang, H\. Lv, Y\. Wang, W\. Shao, P\. Chu, Z\. Tu, T\. He, Z\. Wu, H\. Deng, J\. Ge, K\. Chen, K\. Zhang, L\. Wang, M\. Dou, L\. Lu, X\. Zhu, T\. Lu, D\. Lin, Y\. Qiao, J\. Dai, and W\. WangExpanding performance boundaries of open\-source multimodal models with model, data, and test\-time scaling\.arXiv preprint arXiv:2412\.05271\.Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p3.1),[§3\.4](https://arxiv.org/html/2608.17205#S3.SS4.SSS0.Px1.p1.1)\.
- Denget al\.\(2025\)A\. Deng, T\. Cao, Z\. Chen, and B\. HooiWords or vision: do vision\-language models have blind faith in text?\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Efron and Tibshirani \(1993\)B\. Efron and R\. J\. TibshiraniAn introduction to the bootstrap\.Chapman & Hall/CRC\.Cited by:[§3\.5](https://arxiv.org/html/2608.17205#S3.SS5.SSS0.Px1.p1.1)\.
- Ernst and Banks \(2002\)M\. O\. Ernst and M\. S\. BanksHumans integrate visual and haptic information in a statistically optimal fashion\.Nature415\(6870\),pp\. 429–433\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p1.1)\.
- Ernst and Bülthoff \(2004\)M\. O\. Ernst and H\. H\. BülthoffMerging the senses into a robust percept\.Trends in Cognitive Sciences8\(4\),pp\. 162–169\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p1.1)\.
- Hendrycks and Dietterich \(2019\)D\. Hendrycks and T\. DietterichBenchmarking neural network robustness to common corruptions and perturbations\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§3\.3](https://arxiv.org/html/2608.17205#S3.SS3.SSS0.Px1.p1.1)\.
- Holtzmanet al\.\(2021\)A\. Holtzman, P\. West, V\. Shwartz, Y\. Choi, and L\. ZettlemoyerSurface form competition: why the highest probability answer isn’t always right\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§3\.3](https://arxiv.org/html/2608.17205#S3.SS3.SSS0.Px4.p1.2)\.
- Huaet al\.\(2025\)T\. Hua, T\. Yun, and E\. PavlickHow do vision\-language models process conflicting information across modalities?\.arXiv preprint arXiv:2507\.01790\.Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Jainet al\.\(2025\)A\. Jain, M\. Vatsa, and R\. SinghWords over pixels? rethinking vision in multimodal large language models\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence \(IJCAI\),pp\. 10481–10489\.Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Laurençonet al\.\(2024\)H\. Laurençon, A\. Marafioti, V\. Sanh, and L\. TronchonBuilding and better understanding vision\-language models: insights and future directions\.External Links:2408\.12637,[Link](https://arxiv.org/abs/2408.12637)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Liet al\.\(2024\)B\. Li, Y\. Zhang, D\. Guo, R\. Zhang, F\. Li, H\. Zhang, K\. Zhang, P\. Zhang, Y\. Li, Z\. Liu, and C\. LiLLaVA\-OneVision: easy visual task transfer\.External Links:2408\.03326,[Link](https://arxiv.org/abs/2408.03326)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Linget al\.\(2017\)W\. Ling, D\. Yogatama, C\. Dyer, and P\. BlunsomProgram induction by rationale generation: learning to solve and explain algebraic word problems\.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 158–167\.Cited by:[Limitations](https://arxiv.org/html/2608.17205#Sx1.p1.1)\.
- Liuet al\.\(2024a\)H\. Liu, C\. Li, Y\. Li, and Y\. J\. LeeImproved baselines with visual instruction tuning\.External Links:2310\.03744,[Link](https://arxiv.org/abs/2310.03744)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Liuet al\.\(2024b\)H\. Liu, C\. Li, Y\. Li, B\. Li, Y\. Zhang, S\. Shen, and Y\. J\. LeeLLaVA\-next: improved reasoning, ocr, and world knowledge\.Note:[https://llava\-vl\.github\.io/blog/2024\-01\-30\-llava\-next/](https://llava-vl.github.io/blog/2024-01-30-llava-next/)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Liuet al\.\(2026\)Q\. Liu, J\. Feng, Y\. Wang, X\. Han, Y\. Cheng, Y\. Zhu, H\. Diao, Y\. Zhuge, and H\. LuVISTA\-Bench: do vision\-language models really understand visualized text as well as pure text?\.arXiv preprint arXiv:2602\.04802\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p2.1),[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Masryet al\.\(2022\)A\. Masry, D\. X\. Long, J\. Q\. Tan, S\. Joty, and E\. HoqueChartQA: a benchmark for question answering about charts with visual and logical reasoning\.InFindings of the Association for Computational Linguistics: ACL 2022,Dublin, Ireland,pp\. 2263–2279\.External Links:[Link](https://aclanthology.org/2022.findings-acl.177/)Cited by:[Appendix G](https://arxiv.org/html/2608.17205#A7.p1.1),[Figure 5](https://arxiv.org/html/2608.17205#A9.F5),[§1](https://arxiv.org/html/2608.17205#S1.p3.1),[§3\.2\.2](https://arxiv.org/html/2608.17205#S3.SS2.SSS2.Px1.p1.1)\.
- Moratelliet al\.\(2026\)N\. Moratelli, C\. Davis, L\. F\. R\. Ribeiro, B\. Byrne, and G\. IglesiasBenchmarking deflection and hallucination in large vision\-language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 28348–28370\.External Links:[Link](https://aclanthology.org/2026.acl-long.1307/)Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Nooralahzadehet al\.\(2026\)F\. Nooralahzadeh, O\. Rohanian, Y\. Zhang, J\. Fürst, and K\. StockingerArbitration failure, not perceptual blindness: how vision\-language models resolve visual\-linguistic conflicts\.arXiv preprint arXiv:2604\.09364\.Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Ortuet al\.\(2026\)F\. Ortu, Z\. Jin, D\. Doimo, and A\. CazzanigaWhen seeing overrides knowing: disentangling knowledge conflicts in vision\-language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14109–14130\.External Links:[Link](https://aclanthology.org/2026.acl-long.642/)Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Paszkeet al\.\(2019\)A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga, A\. Desmaison, A\. Köpf, E\. Yang, Z\. DeVito, M\. Raison, A\. Tejani, S\. Chilamkurthy, B\. Steiner, L\. Fang, J\. Bai, and S\. ChintalaPyTorch: an imperative style, high\-performance deep learning library\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Appendix A](https://arxiv.org/html/2608.17205#A1.SS0.SSS0.Px7.p1.1)\.
- Patelet al\.\(2021\)A\. Patel, S\. Bhattamishra, and N\. GoyalAre NLP models really able to solve simple math word problems?\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT\),pp\. 2080–2094\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p3.1),[§3\.4](https://arxiv.org/html/2608.17205#S3.SS4.SSS0.Px1.p1.1)\.
- Penget al\.\(2026\)Y\. Peng, P\. Xia, D\. Zhong, K\. Zeng, S\. Han, Y\. Zhou, J\. Liu, R\. Zhang, and H\. YaoSimpleOCR: rendering visual questions to teach MLLMs to read\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 10697–10710\.External Links:[Link](https://aclanthology.org/2026.findings-acl.519/)Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Pezeshkpouret al\.\(2025\)P\. Pezeshkpour, M\. Aminnaseri, and E\. HruschkaMixed signals: decoding VLMs’ reasoning and underlying bias in vision\-language conflict\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 24833–24848\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1351/)Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Sánchez Villegaset al\.\(2026\)D\. Sánchez Villegas, S\. Lewis\-Lim, N\. Aletras, and D\. ElliottReasoning dynamics and the limits of monitoring modality reliance in vision\-language models\.arXiv preprint arXiv:2604\.14888\.Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Smith \(2007\)R\. SmithAn overview of the tesseract OCR engine\.InNinth International Conference on Document Analysis and Recognition \(ICDAR\),Cited by:[§3\.5](https://arxiv.org/html/2608.17205#S3.SS5.SSS0.Px2.p1.2)\.
- Sunet al\.\(2026\)K\. Sun, X\. Yuan, H\. Liu, C\. Zhao, C\. Zhang, M\. Dredze, and F\. BaiReading, not thinking: understanding and bridging the modality gap when text becomes pixels in multimodal LLMs\.arXiv preprint arXiv:2603\.09095\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p2.1)\.
- Wanget al\.\(2024\)P\. Wang, S\. Bai, S\. Tan, S\. Wang, Z\. Fan, J\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, Y\. Fan, K\. Dang, M\. Du, X\. Ren, R\. Men, D\. Liu, C\. Zhou, J\. Zhou, and J\. LinQwen2\-VL: enhancing vision\-language model’s perception of the world at any resolution\.External Links:2409\.12191,[Link](https://arxiv.org/abs/2409.12191)Cited by:[§3\.1](https://arxiv.org/html/2608.17205#S3.SS1.p1.1)\.
- Wilcoxon \(1945\)F\. WilcoxonIndividual comparisons by ranking methods\.Biometrics Bulletin1\(6\),pp\. 80–83\.Cited by:[§3\.5](https://arxiv.org/html/2608.17205#S3.SS5.SSS0.Px1.p1.1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. M\. RushTransformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Cited by:[Appendix A](https://arxiv.org/html/2608.17205#A1.SS0.SSS0.Px7.p1.1)\.
- Xuet al\.\(2026\)Y\. Xu, Y\. Wang, Z\. Wu, K\. Song, J\. Lin, and Z\. ShenDo vision\-language models truly perform vision reasoning? a rigorous study of the modality gap\.arXiv preprint arXiv:2604\.16256\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p2.1),[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Yuanet al\.\(2025\)F\. Yuan, Y\. Yan, Y\. Jiang, H\. Zhao, T\. Feng, J\. Chen, Y\. Lou, W\. Zhang, Y\. Shen, W\. Lu, J\. Xiao, and Y\. ZhuangGSM8K\-V: can vision language models solve grade school math word problems in visual contexts\.arXiv preprint arXiv:2509\.25160\.Cited by:[§1](https://arxiv.org/html/2608.17205#S1.p2.1),[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px1.p1.1)\.
- Yueet al\.\(2024\)X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng, R\. Liu, G\. Zhang, S\. Stevens, D\. Jiang, W\. Ren, Y\. Sun, C\. Wei, B\. Yu, R\. Yuan, R\. Sun, M\. Yin, B\. Zheng, Z\. Yang, Y\. Liu, W\. Huang, H\. Sun, Y\. Su, and W\. ChenMMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert AGI\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025\)Z\. Zhang, T\. Wang, X\. Gong, Y\. Shi, H\. Wang, D\. Wang, and L\. HuWhen modalities conflict: how unimodal reasoning uncertainty governs preference dynamics in MLLMs\.arXiv preprint arXiv:2511\.02243\.Cited by:[§2](https://arxiv.org/html/2608.17205#S2.SS0.SSS0.Px2.p1.1)\.
## Appendix AImplementation Details
##### Models and checkpoints\.
We use the following Hugging Face checkpoints\. The listed revisions are the snapshots resolved in the experiment environment; InternVL2\-8B is included in generated\-answer analyses only because its model\-specific chat interface does not support the continuation\-scoring procedure used for CLL\.
Table 2:Hugging Face checkpoints and pinned commit revisions \(InternVL2\-8B is used for behavioral analysis only\)\.Models are loaded withdevice\_map="auto",torch\_dtype=bfloat16, and no quantization\. Qwen, LLaVA, Idefics3, and Phi use their checkpoint\-specific Hugging Face processors without manual image\-resolution overrides\. InternVL2 uses its model\-specific chat interface; images are converted to RGB, resized to448×448448\\times 448with bicubic interpolation, and normalized using ImageNet mean and standard deviation\. The attention backend is left at the Transformers or checkpoint default except for Phi\-3\.5\-Vision, for which the loader explicitly requestseagerattention\.
##### Prompting\.
Arithmetic prompts ask the model to solve the problem step by step and end with\#\#\#\# <answer\>\. ChartQA\-Conflict instead requests exactly one answer\-only line in that format\. For Qwen2\-VL, Qwen2\.5\-VL, LLaVA\-OneVision, and Idefics3, we render the checkpoint’s chat template usingapply\_chat\_templatewithadd\_generation\_prompt=True\. LLaVA\-1\.6 and Phi\-3\.5 use the following fixed templates, where\{prompt\}denotes the complete task prompt and line breaks are literal:
```
LLaVA-1.6:
[INST] <image>
{prompt} [/INST]
Phi-3.5:
<|user|>
<|image_1|>
{prompt}<|end|>
<|assistant|>
```
For CLL scoring, the same task prompt and model\-specific conversation scaffold are used\. The assistant context is then extended with the literal delimiter\#\#\#\#before either candidate answer is appended\.
##### Generation\.
Open\-model decoding is deterministic and greedy:do\_sample=False, with temperature, top\-pp, and top\-kkunset\. We usemax\_new\_tokens=256for the arithmetic experiments andmax\_new\_tokens=128for open\-model ChartQA\-Conflict\. Items are generated individually \(batch size one\), and generation terminates at the model’s configured stopping token or the output\-token limit\.
For arithmetic responses, answer extraction first uses the value following\#\#\#\#; if that delimiter is absent, it falls back to the final numeric value in the response\. Source attribution compares the extracted number with the two candidate answers after rounding\. ChartQA\-Conflict uses a stricter answer\-only parser: it takes the value following\#\#\#\#, or, when that delimiter is absent, the first nonempty line only if the entire line normalizes to a single valid answer\. It applies typed normalization and never extracts numbers from surrounding explanations or reasoning traces\.
##### Frontier API models\.
The two frontier models are accessed through hosted APIs: GPT\-5\.6\-Luna \(OpenAI Chat Completions,gpt\-5\.6\-luna\) and Gemini\-3\.5\-Flash \(Googlegenerate\_content,gemini\-3\.5\-flash\)\. Because neither API exposes teacher\-forced continuation scores, both are evaluated with generated answers only and never enter the CLL analysis\. We request answer\-only outputs in the same single\-line\#\#\#\# <answer\>format used for open\-model ChartQA\-Conflict and attribute responses with the identical typed, exact\-match normalization, without the numeric reasoning\-trace fallback used for the arithmetic parser\. Both models use a completion budget of1,0241\{,\}024tokens andtemperature=0\\text\{temperature\}=0\(automatically omitted for models that reject a non\-default temperature\), with low provider\-side reasoning: OpenAI reasoning effortnoneand Gemini thinking levelminimal\. We run the image\-degradation arm first and reuse its clean \(L0\) generations as the report\-degradation arm’s L0, so the shared clean endpoint is not resampled; on GSM8K the two L0 endpoints were generated independently, and we report the resulting baseline sensitivity in Appendix[D](https://arxiv.org/html/2608.17205#A4)\. Conflict items are loaded from a pinned Hugging Face dataset revision, taking the deterministic prefix of the frozen row order when a subset is used so that conflict IDs and corruption seeds match the open\-model runs\. The frontier ChartQA\-Conflict evaluation uses all229229items and the GSM8K behavioral evaluation300300role\-neutral conflicts \(GPT\-5\.6\-Luna only\)\. Before each full run we verified output plausibility on a small 5\-item generation probe—requiring non\-empty responses, valid A/B attribution, and exactly one\#\#\#\# <answer\>line—and launched the full run only after this check passed\. We persist per\-item API response metadata \(finish reason and token counts\) for post\-hoc validation\. Because hosted APIs are not fully deterministic, all frontier results are treated as supporting behavioral evidence rather than part of the primary analysis\.
##### Conditional log\-likelihood scoring\.
CLL is computed using teacher forcing\. The model\-specific scoring context ends with the literal string\#\#\#\#, including one trailing space\. Candidate strings are appended without an additional leading space\. We tokenize the context and the context–candidate concatenation withadd\_special\_tokens=False, and define the candidate span as
\|tok\(c\+a\)\|−\|tok\(c\)\|\.\\left\|\\operatorname\{tok\}\(c\+a\)\\right\|\-\\left\|\\operatorname\{tok\}\(c\)\\right\|\.Only this suffix is scored: the logit at positionp−1p\-1predicts the token at positionpp\. Letlogp\(a∣c\)\\log p\(a\\mid c\)denote the summed teacher\-forced log\-probability of the candidate’s answer tokens; dividing by the candidate token count\|a\|\|a\|gives the mean log\-probability per answer token in nats/token\. The arbitration margin is then identical to Eq\. \(1\),
m\(c\)=1\|aT\|logp\(aT∣c\)−1\|aI\|logp\(aI∣c\),m\(c\)=\\frac\{1\}\{\|a\_\{T\}\|\}\\log p\(a\_\{T\}\\mid c\)\-\\frac\{1\}\{\|a\_\{I\}\|\}\\log p\(a\_\{I\}\\mid c\),so a positive margin favors the text\-supported answer and a negative margin favors the image\-supported answer\. Both candidates are scored under the same multimodal context\. The clean \(L0\) condition is the fully clean image–text pair and is therefore identical in both degradation arms, so a single L0 CLL value is shared byRIR\_\{I\}andRTR\_\{T\}\(and counted once in the measure\-agreement analysis, Appendix[C](https://arxiv.org/html/2608.17205#A3)\)\.
##### Determinism and seeds\.
Open\-model generation uses greedy decoding, and all corruptions are deterministic\. The image corruption for the itemiiis seeded by42\+i42\+i; text corruption is seeded by the corresponding textual\-source index\. Statistical confidence intervals and permutation tests use10,00010\{,\}000resamples with fixed seeds\. The exact seeds and model\-specific offsets are provided in the analysis scripts released\.
##### Environment and hardware\.
Each open\-model inference job used one NVIDIA L40S GPU with 48 GB of memory\. The final experiments required approximately400GPU\-hours in total, excluding preliminary smoke tests and failed runs\. The experiments used Python 3\.10\.20 \(GCC 14\.3\.0\), PyTorch 2\.5\.1\([24](https://arxiv.org/html/2608.17205#bib.bib33)\)with CUDA 12\.1 \(torch==2\.5\.1\+cu121\), cuDNN 9\.1\.0, Transformers 4\.49\.0\([33](https://arxiv.org/html/2608.17205#bib.bib34)\), Torchvision 0\.20\.1 \(torchvision==0\.20\.1\+cu121\), Pillow 12\.2\.0 and Accelerate 1\.14\.0\.
## Appendix BDegradation Details
##### Image degradation\.
For GSM8K and SVAMP, we degrade the problem images rendered\. For ChartQA\-Conflict, we apply the same operations to the original charts\. L2 uses Gaussian blur with radius 1\. L4 uses Gaussian blur with radius 2 followed by Gaussian pixel noise with standard deviation 15\. L5 uses Gaussian blur with radius 3, Gaussian noise with standard deviation 25, and a contrast factor of 0\.7\. L0 leaves the image unchanged and is pixel\-identical to the clean image in the conflict condition\. Corruption is deterministic, using seed42\+i42\+ifor image itemii, so every model receives the same degraded image\.
##### Text degradation\.
In the text\-degradation arm, the image remains clean while the text is corrupted\. Each non\-whitespace character is independently selected with probability0\.080\.08at L2,0\.180\.18at L4, and0\.350\.35at L5\. A selected character is deleted with probability0\.50\.5and otherwise replaced with a random alphanumeric character\. Whitespace is preserved so that word boundaries and the overall structure remain visible\. Text corruption is deterministic, seeded by the textual source’s index—the paired item\(i\+1\)modN\(i\+1\)\\bmod Nin the arithmetic setting, or the report’s own item in ChartQA\-Conflict—and the same corrupted text is used for generation and CLL scoring\.
## Appendix CAgreement Between Generated Choice and CLL Margin
The CLL margin provides a continuous measure of direct\-answer candidate preference, whereas generated source choice records which candidate appears in the model’s final answer\. To assess how closely these measures correspond, we compare the sign of the CLL margin with the generated choice on attributable trials\. A positive margin predicts the text\-supported answer and a negative margin predicts the image\-supported answer\. We exclude generations matching neither or both candidates, invalid responses, and unavailable or exactly zero margins\.
Table[5](https://arxiv.org/html/2608.17205#A4.T5)reports agreement by model and degradation arm, pooled across GSM8K and SVAMP\. After counting the clean endpoint shared by the two arms only once, the measures agree on20,281/26,893=75\.4%20\{,\}281/26\{,\}893=75\.4\\%of observations \(item\-clustered bootstrap 95% CI:74\.4%74\.4\\%–76\.4%76\.4\\%\)\. Agreement varies across models and is lower overall in the text\-degradation arm\. Thus, the CLL margin is behaviorally grounded but is not interchangeable with generated choice; we use it as a continuous measure of candidate preference and report generated behavior separately\.
## Appendix DFrontier\-Model Behavioral Results
We evaluate GPT\-5\.6\-Luna and Gemini\-3\.5\-Flash using generated answers only, because the API interfaces used in our experiments do not expose the teacher\-forced continuation scores required for the CLL margin\. Attribution is exact: formatting and compatible unit labels are normalized, but conflicting scales, currencies, or units are rejected\. Neither model enters the primary six\-model CLL analysis\. Throughout, the subscriptIIdenotes the image channel \(the chart in ChartQA\-Conflict, the rendered image in GSM8K\) andTTthe text channel \(the report or the plain\-text problem\);RIR\_\{I\}andRTR\_\{T\}are the within\-item reallocations toward the clean channel when the other is degraded, andA=RT−RIA=R\_\{T\}\-R\_\{I\}\.
##### ChartQA\-Conflict endpoints\.
Table[3](https://arxiv.org/html/2608.17205#A4.T3)reports the paired L0\-to\-L5 contrasts\. Both frontier models have negative asymmetry, reproducing the direction of the open\-model ChartQA\-Conflict CLL analysis\. GPT\-5\.6\-Luna reallocates almost entirely under chart degradation; Gemini\-3\.5\-Flash shows a smaller but still strongly chart\-dominant contrast\.
Table 3:Generated\-answer contrasts on ChartQA\-Conflict\.RIR\_\{I\}/RTR\_\{T\}: within\-item reallocation toward the clean report/chart when the chart/report is degraded;A=RT−RIA=R\_\{T\}\-R\_\{I\}, negativeAAindicating stronger reallocation under chart degradation\. 95% CIs are paired\-bootstrap \(10,000 resamples\); both paired permutation testsp<\.001p<\.001\.
##### ChartQA\-Conflict trajectories\.
Table[4](https://arxiv.org/html/2608.17205#A4.T4)reports preference for the source that remains clean, among attributable generated answers at each level\. Both models stay largely chart\-following under light chart degradation and move sharply toward the report under moderate or heavy chart degradation, whereas report degradation produces little additional movement toward the already\-preferred chart\.
Table 4:Frontier\-model ChartQA\-Conflict trajectories: behavioral preference \(%\), among attributable generated answers, for the source that remains clean\.*Report*rows track report\-following as the chart is degraded;*Chart*rows track chart\-following as the report is degraded\. Arms coincide at L0; the report\-degradation arm reuses the image\-arm L0 generations, so the L0 values are complementary\.
##### GSM8K behavioral contrast\.
We also evaluate GPT\-5\.6\-Luna on 300 role\-neutral GSM8K conflicts at L0 and L5\. The model moves toward the clean source in both arms, but more strongly when the text is degraded:RI=0\.323R\_\{I\}=0\.323,RT=0\.614R\_\{T\}=0\.614,A=\+0\.291A=\+0\.291\. The complete\-case analysis retainsn=189n=189matched items, with a paired\-bootstrap 95% CI of\[\+0\.169,\+0\.407\]\[\+0\.169,\+0\.407\]and a two\-sided paired permutation test givingp<\.001p<\.001\. This direction agrees with the predominant open\-model arithmetic result and is opposite to GPT\-5\.6\-Luna’s ChartQA\-Conflict result\.
##### Clean\-endpoint sensitivity\.
Unlike the frontier ChartQA experiment, which reuses the image\-arm L0 generations in the report\-degradation arm, the two nominally identical GSM8K L0 endpoints were generated independently, with behavioral text preferences of65\.9%65\.9\\%and60\.2%60\.2\\%\(API nondeterminism\)\. Repeating the descriptive endpoint calculation with each L0 sample as the shared clean baseline leaves the sign unchanged:A=\+0\.317A=\+0\.317using the image\-arm L0 baseline andA=\+0\.203A=\+0\.203using the text\-arm L0 baseline\. Baseline resampling thus affects the estimated magnitude but not the direction, and we treat the frontier GSM8K experiment as supporting behavioral evidence rather than part of the primary CLL analysis\.
Table 5:Agreement between generated source choice and CLL\-margin sign \(role\-neutral arithmetic runs\)\. Cells give count agreeing / attributable observations, with the agreement percentage below\. Model rows pool GSM8K and SVAMP over all four legibility levels; the arm\-pooled row also pools across models\.
## Appendix ELegibility\-Adjusted Regression Details
The calibration uses the proportional loss of accuracy in a single channelℓM,c,L=\(AccM,c,0−AccM,c,L\)/AccM,c,0\\ell\_\{M,c,L\}=\(\\mathrm\{Acc\}\_\{M,c,0\}\-\\mathrm\{Acc\}\_\{M,c,L\}\)/\\mathrm\{Acc\}\_\{M,c,0\}, measured by presenting only the degraded channel at levelLLand scoring its source\-supported answer\.
*Per\-model slopes \(primary calibration\)\.*We summarize reallocation at each level by the median CLL\-margin shift toward the source that stays clean,
Δ~M,L\(I\)\\displaystyle\\widetilde\{\\Delta\}^\{\(I\)\}\_\{M,L\}=mediani\(mi,L\(I\)−mi,L0\(I\)\),\\displaystyle=\\operatorname\{median\}\_\{i\}\\\!\\left\(m^\{\(I\)\}\_\{i,L\}\-m^\{\(I\)\}\_\{i,L0\}\\right\),Δ~M,L\(T\)\\displaystyle\\widetilde\{\\Delta\}^\{\(T\)\}\_\{M,L\}=mediani\(mi,L0\(T\)−mi,L\(T\)\),\\displaystyle=\\operatorname\{median\}\_\{i\}\\\!\\left\(m^\{\(T\)\}\_\{i,L0\}\-m^\{\(T\)\}\_\{i,L\}\\right\),and for each model and arm fit a slope through\-originΔ~M,c,L=bM,cℓM,c,L\\widetilde\{\\Delta\}\_\{M,c,L\}=b\_\{M,c\}\\,\\ell\_\{M,c,L\}overL∈\{L0,L2,L4,L5\}L\\in\\\{\\mathrm\{L0\},\\mathrm\{L2\},\\mathrm\{L4\},\\mathrm\{L5\}\\\}\(through the origin because both quantities are zero at L0\)\. The arm contrast isDM=bM,T−bM,ID\_\{M\}=b\_\{M,T\}\-b\_\{M,I\}; we bootstrap the matched conflict items \(10,000 resamples\), recompute the median shifts, and refit both slopes with the accuracy\-loss estimates held fixed\. Per\-model slopes for the arithmetic benchmarks and for ChartQA\-Conflict are in Tables[7](https://arxiv.org/html/2608.17205#A5.T7)and[8](https://arxiv.org/html/2608.17205#A6.T8)\.
*Pooled regression \(complementary check\)\.*As a complementary item\-level analysis we also fit
Δi,M,c,L=\\displaystyle\\Delta\_\{i,M,c,L\}=\{\}β0\+β1ℓM,c,L\+β2Tc\\displaystyle\\beta\_\{0\}\+\\beta\_\{1\}\\ell\_\{M,c,L\}\+\\beta\_\{2\}T\_\{c\}\+β3\(ℓM,c,LTc\)\+γM\+εi,M,c,L,\\displaystyle\+\\beta\_\{3\}\(\\ell\_\{M,c,L\}T\_\{c\}\)\+\\gamma\_\{M\}\+\\varepsilon\_\{i,M,c,L\},whereTc=1T\_\{c\}=1for text degradation andγM\\gamma\_\{M\}are model fixed effects\. Becauseℓ\\ellis shared within a model–channel–level cell, we do not treat item rows as independent evidence for the slope; primary intervals cluster by model, and clustering by the 36 model–channel–level cells is a complementary check\. With only six model clusters, these results are evidence for the evaluated model set, not a population\-wide estimate\.
Table 6:Task\-accuracy\-adjusted channel interactions under the role\-neutral arithmetic prompt\. Positiveβ3\\beta\_\{3\}means that reallocation per unit of lost unimodal accuracy is larger under text degradation\. The item\-level regression estimates a mean shift; the cell\-median analysis is a coarser sensitivity analysis with a different estimand\.Table 7:Per\-model reallocation slopes on the rendered\-text arithmetic benchmarks after calibration by proportional unimodal task\-accuracy loss\. Slopes are fitted through the origin over L0, L2, L4, and L5\.bIb\_\{I\}is movement toward clean text per unit of lost image\-only accuracy, andbTb\_\{T\}is movement toward the clean image per unit of lost text\-only accuracy\. PositiveD=bT−bID=b\_\{T\}\-b\_\{I\}indicates greater reallocation under text degradation\. Confidence intervals condition on the observed accuracy losses and use 10,000 bootstrap resamples of the matched conflict items\.
## Appendix FThe Reversal Survives Accuracy Calibration
A natural objection to the ChartQA reversal is that degrading the chart may simply destroy more usable information than degrading the report, so that stronger reallocation away from the chart reflects a larger information loss rather than a genuine preference\. To control for this, we calibrate each arm’s reallocation by how much the same degradation lowers the model’s*unimodal*accuracy—its accuracy when only that source is available\. The slopebIb\_\{I\}\(bTb\_\{T\}\) then measures reallocation per unit of lost chart\-only \(report\-only\) accuracy, placing the two directions on equal footing: a model that reallocates only because a source has become less informative would show no difference between them\. Table[8](https://arxiv.org/html/2608.17205#A6.T8)shows the reversal persists:D=bT−bID=b\_\{T\}\-b\_\{I\}stays negative for all six models with 95% CIs excluding zero, so it is not explained by chart degradation destroying more information\. All six models exceed the prespecified0\.100\.10clean\-accuracy floor and are calibrated\.
Table 8:Calibrated slopes on ChartQA\-Conflict\.bIb\_\{I\}measures reallocation per unit of lost chart\-only accuracy, andbTb\_\{T\}measures reallocation per unit of lost report\-only accuracy\. Slopes are fitted through the origin over L0, L2, L4, and L5\. NegativeD=bT−bID=b\_\{T\}\-b\_\{I\}indicates stronger reallocation per unit of lost chart accuracy\. All six models exceed the prespecified0\.100\.10clean\-accuracy floor and are calibrated; everyDDis negative with a 95% CI excluding zero\. Intervals are paired\-bootstrap 95% confidence intervals over matched items using 10,000 resamples\.The random\-item\-intercept variance reached the boundary at zero because the response is already an L0\-referenced within\-item change\. We therefore do not use the resulting mixed\-modelpp\-values; Table[6](https://arxiv.org/html/2608.17205#A5.T6)reports the cluster\-aware fixed\-effects fits\. A separate survival\-based calibration yieldsβ3=\+5\.55\\beta\_\{3\}=\+5\.55\(p=\.013p=\.013\) on GSM8K and\+7\.97\+7\.97\(p=\.095p=\.095\) on SVAMP\.
## Appendix GChart\-versus\-Table Control
For each of the 229 conflicts we render the item’s official ChartQA source table\([20](https://arxiv.org/html/2608.17205#bib.bib3)\)as a plain table image and substitute it for the chart in the visual slot, keeping the question, report, counterfactual answeraTa\_\{T\}, Source A/B labels, and the L0/L2/L4/L5 degradation ladder identical\. We render only the rows the chart displays, so the table and chart carry the same information and neither is systematically easier to read\. Table values are taken directly from the source tables and are never edited towardaTa\_\{T\}; the table supportsaIa\_\{I\}exactly as the chart does\. Table[9](https://arxiv.org/html/2608.17205#A7.T9)reports the result\.
Table 9:Chart\-versus\-table control on ChartQA\-Conflict \(n=229n=229per model\)\.AchartA\_\{\\mathrm\{chart\}\}is the median asymmetry with the original chart \(as in Table[1](https://arxiv.org/html/2608.17205#S4.T1)\);AtableA\_\{\\mathrm\{table\}\}is the median asymmetry when the visual source is the official source table rendered as a plain image\. The last column is the within\-item paired difference \(table−\-chart\); because it is a within\-item median, it need not equalAtable−AchartA\_\{\\mathrm\{table\}\}\-A\_\{\\mathrm\{chart\}\}from the two columns\. All intervals are 95% paired\-bootstrap CIs \(10,000 resamples\);ppermp\_\{\\mathrm\{perm\}\}is a two\-sided paired permutation test\. EveryAtableA\_\{\\mathrm\{table\}\}interval excludes zero, so the reversal persists under the table representation\. The chart\-versus\-table difference is not systematic: four of six intervals include zero, and the two significant changes \(Qwen2\.5\-VL\-7B, Phi\-3\.5\-Vision\) point in opposite directions\.
## Appendix HCLL Normalization Sensitivity
The primary CLL margin normalizes each candidate’s summed token log\-probability by its answer\-token length\. To test whether the arm asymmetry depends on this choice, we define
mα\(c\)=logp\(aT∣c\)\|aT\|α−logp\(aI∣c\)\|aI\|α,m\_\{\\alpha\}\(c\)=\\frac\{\\log p\(a\_\{T\}\\mid c\)\}\{\|a\_\{T\}\|^\{\\alpha\}\}\-\\frac\{\\log p\(a\_\{I\}\\mid c\)\}\{\|a\_\{I\}\|^\{\\alpha\}\},where each log probability is summed over the candidate’s answer tokens\. Thus,α=0\\alpha=0uses the unnormalized summed log\-probability,α=0\.5\\alpha=0\.5applies partial normalization, andα=1\\alpha=1gives the primary mean log\-probability per answer token\. We recompute the item\-level arm contrast using each version of the margin for role\-neutral GSM8K \(n=1319n=1319\), role\-neutral SVAMP \(n=300n=300\), and ChartQA\-Conflict \(n=229n=229, after the single audited exclusion\)\.
Table[10](https://arxiv.org/html/2608.17205#A8.T10)shows that the normalization exponent rescales the magnitude of the asymmetry but does not change its sign or significance in any model–benchmark cell: every entry is significant \(Wilcoxon signed\-rank,p<0\.05p<0\.05; mostp≪10−10p\\ll 10^\{\-10\}\), and within each cell the estimate moves monotonically toward zero asα\\alphaincreases without crossing it\. The*direction*of the asymmetry is stable under normalization but differs by benchmark\. On both arithmetic benchmarks the median asymmetry is positive—indicating stronger reallocation under text degradation—for all models except Qwen2\.5\-VL\-7B, which is near zero on GSM8K and weakly negative on SVAMP at everyα\\alpha\. On ChartQA\-Conflict the asymmetry is negative for all six models at everyα\\alpha, indicating stronger reallocation under chart \(image\) degradation\. In no case does the choice ofα\\alphaalter these conclusions, confirming that the primaryα=1\\alpha=1margin is representative\.
Table 10:Median arm asymmetry under alternative candidate\-length normalizations \(α∈\{0,0\.5,1\}\\alpha\\in\\\{0,0\.5,1\\\}\) for role\-neutral GSM8K, role\-neutral SVAMP, and ChartQA\-Conflict\. Positive values indicate stronger reallocation under text degradation; negative values, stronger reallocation under image \(chart\) degradation\. All entries are significant \(Wilcoxon signed\-rank,p<0\.05p<0\.05\)\. Within every cell the sign and significance are preserved acrossα\\alpha; the exponent affects only the estimated magnitude\.
## Appendix IChartQA\-Conflict Construction
##### Question eligibility\.
We screened ChartQA questions for whether the chart\-supported answer could be represented and independently verified in a textual evidence report\. We retained questions with an unambiguous semantic reference to a category, series, and value \(e\.g\., a named country/year/category or a clearly specified aggregate\)\. We excluded questions that depended on visual position, color, bar height, or an ambiguous series mapping \(e\.g\., “the rightmost upper bar” or “the green graph”\), because their answer could not be reliably recovered from the source table alone\. Arithmetic, ratio, and comparison questions were retained only when every referenced value and operation was unambiguous and reproduced the ChartQA gold answer\.
##### Source items and counterfactual answers\.
We start from 449 ChartQA test questions \(430 numeric, 19 yes/no\); the frozen ChartQA\-Conflict release keeps the 230 numeric\-answer items\. For each, the original ChartQA answer is the chart\-supported answeraIa\_\{I\}, and we build the report\-supported answeraTa\_\{T\}by perturbingaIa\_\{I\}while preserving its type, unit, sign, scale, and precision, so thataTa\_\{T\}differs fromaIa\_\{I\}only after normalization\. Whenever possible we choose a value that does not appear anywhere in the chart\. This keeps attribution clean—an answer equal toaIa\_\{I\}came from the chart, one equal toaTa\_\{T\}from the report—and prevents a model from producingaTa\_\{T\}by misreading some other visible chart value\. Rather than fixing a distance bound \(e\.g\.,±X%\\pm X\\%ofaIa\_\{I\}\), we choose each perturbation to stay plausible for the question and chart, and discard any that collapses toaIa\_\{I\}after rounding, falls outside the expected answer range, or changes the unit, type, sign, or precision\.
##### Table construction\.
For the table ablation, we render each item’s official ChartQA source table as a plain table image containing the same rows, columns, labels, and values as the corresponding chart\. Tables use regular DejaVu Sans for body cells and DejaVu Sans Bold for headers\. The renderer selects the largest font size between 8 and 18 pixels that fits the complete table; 208 of the 229 tables use the maximum 18\-pixel size\. Each table is rendered on a white RGB canvas with exactly the same pixel dimensions as its corresponding chart \(190 items use800×557800\\times 557pixels\), thereby holding the image canvas fixed within item\. Headers are lightly shaded, rows alternate between white and light gray, and cells are separated by one\-pixel borders\. One unusually long table is split into two side\-by\-side panels with repeated headers so that all source values remain visible\.
##### Review\.
Two of the authors, both senior undergraduates, independently reviewed all 230 items, each confirming that \(i\) the report entailsaTa\_\{T\}, \(ii\)aTa\_\{T\}is a valid answer to the shared question, and \(iii\)aT≠aIa\_\{T\}\\neq a\_\{I\}after normalization; items flagged by either reviewer were discussed and adjudicated\. Because these reviewers are authors of this paper and were not blind to the counterfactual design, this check is not independent of the study \(see Limitations\)\.
##### Prompt template\.
The chart and report receive counterbalanced Source A/B labels, with the A source shown first:
> Question: \{question\} You are given two conflicting evidence sources\. Neither source is privileged\. Source \{chart label\} is the attached chart\. Source \{report label\} is the textual report below: \{report\} Respond with exactly one line ‘\#\#\#\# <answer\>’, giving only the answer value with no explanation\.
ChartQA\-Conflict item\(counterbalanced Source A/B prompt\)Question: What is the poverty rate in California in the year 2019?
You are given two conflicting evidence sources\. Neither source is privileged\.Source A — attached chartSource B — textual reportThe accompanying report summarizes California’s poverty rate across multiple years, reported as percentages of the population\. It records a poverty rate of14\.8%in 2019, along with 13\.3% in 2017 and 14\.3% in 2016\.Chart \(A\): 2019==11\.8%⇒aI=11\.8\\Rightarrow a\_\{I\}\{=\}11\.8*\(chart\-following\)*Report \(B\): 2019==14\.8%⇒aT=14\.8\\Rightarrow a\_\{T\}\{=\}14\.8*\(report\-following\)*Respond with exactly one line ‘\#\#\#\# <answer\>’\.Figure 5:A ChartQA\-Conflict item as given to the VLM\. The question asks for California’s 2019 poverty rate\.Source Ais the original chart, whose 2019 value supportsaI=11\.8a\_\{I\}\{=\}11\.8;Source Bis the evidence\-bearing report from our dataset, perturbed to supportaT=14\.8a\_\{T\}\{=\}14\.8\. A reply of11\.811\.8is chart\-following,14\.814\.8is report\-following; Source A/B labels are counterbalanced across items\. The answer\-identification annotations shown below the sources are explanatory overlays and were not included in the model input\. Chart from ChartQA\([20](https://arxiv.org/html/2608.17205#bib.bib3)\)\.
## Appendix JChartQA\-Conflict Behavioral Results
Generated\-answer contrasts \(Table[11](https://arxiv.org/html/2608.17205#A10.T11)\) use only items that are decidable at all four required endpoints for a given model, sonnvaries and is smaller than the 229\-item CLL sample\. The five models with usable generated answers and CLL scores corroborate the CLL reversal\. InternVL2 is behavioral\-only because its custom chat interface does not expose continuation scoring; it shows the opposite behavioral direction\. Phi\-3\.5\-Vision has valid CLL scores and, under the answer\-only extraction, produces well\-formed first\-line answers, but only nine of its items are decidable at all four endpoints—too few for a behavioral contrast—so it is calibrated \(Appendix[F](https://arxiv.org/html/2608.17205#A6)\) but not analyzed behaviorally\.
Table 11:Generated\-answer ChartQA\-Conflict contrasts after excluding the item that failed the report\-entailment audit\. Values are complete\-case mean changes in source\-choice indicators, not CLL margins\. NegativeA=RT−RIA=R\_\{T\}\-R\_\{I\}indicates stronger reallocation under chart degradation\. Intervals are paired bootstrap 95% CIs andpWp\_\{W\}is the two\-sided Wilcoxon signed\-rank test\.
## Appendix KPrompt\-Framing Results
Table[12](https://arxiv.org/html/2608.17205#A11.T12)compares the original prompt, which presents the text as the problem and the image as an attachment, with the role\-neutral Source A/Source B prompt\. The comparison uses the same GSM8K items under both framings and is available for five models with complete matched CLL results\. Positive values ofAAindicate stronger reallocation under text degradation than under image degradation\.
The asymmetry remains positive for every model under both prompts, but framing changes its magnitude in a model\-dependent direction\. Neutral framing reduces the effect for Qwen2\.5\-VL\-7B and Idefics3\-8B, increases it for Qwen2\-VL\-2B and LLaVA\-OneVision\-7B, and produces no detectable change for Phi\-3\.5\-Vision\. LLaVA\-1\.6 is omitted from the CLL comparison because matched original\-prompt CLL results are unavailable; its generated\-answer results are available under both framings\.
Table 12:Prompt\-framing sensitivity on matched GSM8K items\. PositiveAAindicates stronger reallocation under text degradation\. The neutral\-minus\-original contrast is computed within item, so it need not equal the difference between the two displayed medians\. Intervals are paired bootstrap 95% confidence intervals, andppis from a two\-sided Wilcoxon signed\-rank test\.Similar Articles
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
This paper investigates whether vision-language models can distinguish potential from established common ground in asymmetric dialogue. Experiments on MapTask data show that providing task-relevant map content (visual or textual) biases models toward over-predicting alignment, as they rely on static referential cues rather than tracking grounding through dialogue history.
Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark
This paper introduces a phrasing-controlled benchmark to measure how much vision-language models rely on textual priors versus image content. Experiments across eleven models show significant degradation when text leakage is minimized, and the authors demonstrate that in-context learning and GRPO post-training can reduce this reliance.
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
This paper challenges the 'Attention-Confidence Assumption' by demonstrating that attention map sharpness is a poor predictor of correctness in Vision-Language Models. Instead, it shows that reliability is better indicated by hidden-state geometry and self-consistency, with significant findings on architectural differences between late-fusion and early-fusion models.
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
This paper introduces CrossMath, a controlled multimodal reasoning benchmark that reveals a critical limitation in current vision-language models: they perform reasoning primarily in textual space rather than genuine vision-grounded reasoning, with visual input often degrading performance compared to text-only baselines. The authors propose fine-tuning approaches to mitigate this modality gap and improve multimodal reasoning capabilities.