GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs
Summary
GeoArbiter proposes a training-free pipeline that selectively injects image-unverifiable geographic facts into remote-sensing multimodal LLMs to reduce knowledge hallucinations while preserving retrieval accuracy gains.
View Cached Full Text
Cached at: 08/04/26, 07:43 AM
# GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs
Source: [https://arxiv.org/html/2608.00877](https://arxiv.org/html/2608.00877)
###### Abstract
Remote\-sensing multimodal large language models \(MLLMs\) often assert facts that imagery cannot establish, such as a facility’s identity or function\. Coordinate\-keyed geographic retrieval can supply this missing knowledge, improving fMoW land\-use accuracy by 12\.06–17\.19 points across three open MLLMs\. However, retrieved records can also contradict visible evidence, and we find that models frequently follow the records even when the image is decisive\. We argue that source trust should therefore depend on*cross\-modal verifiability*: geographic records are most useful for attributes the image cannot verify and most dangerous when they dispute visually verifiable attributes\. We introduce GeoArbiter, a training\-free pipeline that operationalizes this principle by injecting only image\-unverifiable geographic facts\. Unlike arbitration prompts, which leak across attribute types and bias yes/no responses, content\-level filtering preserves 84\.69–87\.15% of the full\-retrieval accuracy gain, reduces claim\-level hallucination by 9\.58–26\.34% under a source\-blinded judge, and improves robustness to conflicting records across all three models\. These results identify verifiability\-guided content selection as a simple, effective mechanism for grounding remote\-sensing MLLMs in fallible geographic knowledge\.
GeoArbiter: Verifiability\-Guided Grounding for Remote\-Sensing Multimodal LLMs
Xuechen LiUniversity of Minnesota, Twin Citiesli003487@umn\.edu
## 1Introduction
Figure 1:GeoArbiter preserves accuracy and robustness\.Full OSM injection improves land\-use QA but becomes brittle to a wrong record; verifiability\-stratified injection remains in the top\-right across all three MLLMs\.Multimodal large language models adapted to remote\-sensing \(RS\) imagery can describe satellite scenes and answer open\-ended questions fluently\(Qwen Team,[2025](https://arxiv.org/html/2608.00877#bib.bib23); Zhuet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib24); Liet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib25)\)\. Yet many consequential errors are not perceptual: a model may misname a facility, invent its institutional role, or fabricate administrative context that pixels cannot establish\. We call these*knowledge hallucinations*, distinguishing them from errors about visible content\(Liet al\.,[2023](https://arxiv.org/html/2608.00877#bib.bib26); Yinet al\.,[2023](https://arxiv.org/html/2608.00877#bib.bib27)\)\. They persist across RS model families\(Zhouet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib3); Liuet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib2)\)and undermine applications in which an incorrect facility identity can be more costly than a missed object\.
Geographic retrieval appears to offer a natural remedy\. Every RS image has coordinates, which provide an exact key into resources such as OpenStreetMap \(OSM;Haklay and Weber,[2008](https://arxiv.org/html/2608.00877#bib.bib47)\), GeoNames, and land\-cover maps\. Prior work, however, mainly uses these resources during training\(Wanget al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib16); Muhtaret al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib17); Ailuroet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib4); Baiet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib5); Andersonet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib38)\)\. Inference\-time RS retrieval instead relies on visual similarity to landmark descriptions\(Wenet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib1)\), which poorly covers ordinary schools, substations, and depots; coordinate\-keyed text systems\(Manviet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib10); Yuet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib8); Fenget al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib9)\)do not face conflicts between records and an image\.
Such conflicts make grounding a source\-selection problem\. Geographic databases are incomplete, stale, and potentially corrupted\(Hanet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib39)\); a record can supply an invisible function but can also claim that a visibly absent runway exists\. Prior conflict studies examine perception, context, and parametric memory\(Jiaet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib11); Liuet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib13); Carragheret al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib14); Ortuet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib15); Lietzowet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib40)\), typically through synthetic contradictions and instruction\- or decoding\-level remedies\(Liuet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib13); Shiet al\.,[2023](https://arxiv.org/html/2608.00877#bib.bib31)\)\. They do not ask which*attributes*each modality is qualified to resolve\.
We propose*cross\-modal verifiability*as that criterion: records should dominate for attributes that an image cannot verify \(e\.g\., function or name\), whereas the image should dominate when a record disputes visible structure\. GeoArbiter enforces this policy in the retrieved content rather than asking a frozen model to follow it\. A key\-level filter retains function and identity records and withholds physical\-structure records before prompting \(Figure[2](https://arxiv.org/html/2608.00877#S1.F2)\)\.
We make three contributions\. First, coordinate\-keyed structured retrieval improves land\-use QA by 12\.06–17\.19 points on all 28,087 functional fMoW validation images\. Second, retrieval’s benefit and risk separate by verifiability: it reduces unsupported function and name claims, whereas fabricated visible\-feature records cause 13\.50–51\.00\-point losses; none of six arbitration prompts recovers the image\-only baseline\. Third, GeoArbiter preserves 84\.69–87\.15% of the retrieval gain, cuts claim\-level hallucination by 9\.58–26\.34% under a source\-blinded judge \(19\.28–45\.19% with the standard judge\), and outperforms every prompt\-level alternative\. Historical OSM analysis further shows that apparent map–image age gaps mostly reflect mapping completion, not physical change\.
Figure 2:GeoArbiter pipeline\.Image coordinates retrieve nearby OSM features, which may be stale or incomplete\. A deterministic verifiability filter withholds physical\-structure keys that the image can check \(red\) and textualizes image\-unverifiable functional records \(blue\) into a knowledge block\. The frozen MLLM answers from the image, question, and filtered knowledge\.
## 2Related Work
#### Hallucination in general and RS MLLMs\.
Object hallucination in general\-domain MLLMs is well documented\(Liet al\.,[2023](https://arxiv.org/html/2608.00877#bib.bib26)\), and post\-hoc systems such as Woodpecker verify generated claims after decoding\(Yinet al\.,[2023](https://arxiv.org/html/2608.00877#bib.bib27)\)\. RS\-specific studies broaden the taxonomy beyond object presence: RSHallu characterizes domain\-specific knowledge errors\(Zhouet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib3)\), while RADAR separates factual and logical failures and steers attention during decoding\(Liuet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib2)\)\. These methods improve how a model uses its existing visual and parametric evidence, but cannot supply a facility name or institutional function that the model never learned\. Our setting targets precisely this missing\-knowledge regime, while also measuring the new failure mode created when external knowledge contradicts perception\.
#### Retrieval augmentation for remote sensing\.
RS\-RAG retrieves encyclopedia entries through cross\-modal embedding similarity and improves captioning and VQA for recognizable landmarks\(Wenet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib1)\)\. Other systems retrieve exemplars for SAR imagery\(Ramirezet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib43)\), image\-derived context for aerial scenes\(Xueet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib44)\), or compose Earth\-observation tools at inference time\(Zhaoet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib6)\); CrisiSense\-RAG combines reports and post\-event imagery across evidence streams\(Xiaoet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib7)\)\. These approaches establish the value of inference\-time context, but visual\-similarity retrieval favors distinctive, documented sites and may return the wrong instance of an ordinary\-looking facility\. GeoArbiter instead uses coordinates as an exact retrieval key and evaluates ordinary function\-defined locations at full fMoW scale\. More importantly, we treat retrieval as a trust problem: the retrieved block is fallible evidence, not an oracle\.
#### OSM in vision–language learning\.
OSM has been distilled into RS captions\(Wanget al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib16)\), instruction data and multimodal alignment corpora\(Muhtaret al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib17); Baiet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib5)\), map\-based pre\-training data\(Andersonet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib38)\), and rendered\-tile domain adaptation\(Ailuroet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib4)\)\. All primarily consume OSM before deployment\. SkyScript is especially related because it classifies tags by whether overhead imagery can visually ground them, retaining groundable tags to improve caption fidelity\(Wanget al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib16)\)\. We use the complementary inference\-time selection: image\-verifiable records are the dangerous ones when incorrect, so GeoArbiter withholds them and injects facts the image cannot adjudicate\. Text\-only coordinate\-grounded systems\(Manviet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib10); Yuet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib8); Fenget al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib9)\)do not observe an image and therefore do not face this source conflict\.
#### Knowledge conflict and adaptive retrieval\.
Multimodal conflict benchmarks construct contradictions among context, images, and parametric memory\(Jiaet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib11); Carragheret al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib14)\); mechanistic studies further identify modality\-specific preference patterns\(Ortuet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib15); Lietzowet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib40)\)\. Focus\-on\-vision prompting is often insufficient\(Liuet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib13)\), and injected text can override visual evidence\(Khayatanet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib42)\)\. In text RAG, incorrect or outdated passages harm generation\(Ouyang and others,[2025](https://arxiv.org/html/2608.00877#bib.bib32)\), motivating retrieval\-quality estimation, query\-complexity routing, reflection, and answer\-attribution filtering\(Asaiet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib28); Yanet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib29); Jeonget al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib30); Chenet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib37)\)\. Those methods ask whether retrieval is needed, relevant, or reliable\. Verifiability is a different axis: an OSM record may be relevant and usually correct, yet still be the wrong evidence to expose when the image can directly settle the disputed attribute\. To our knowledge, prior work has not used this property to select multimodal context or directly compared content\-level with instruction\-level arbitration\.
## 3GeoArbiter
### 3\.1Task formulation
Given an RS imageII, coordinates\(ϕ,λ\)\(\\phi,\\lambda\), footprint radiusrr, and instructionqq, a frozen MLLM predictsy^=argmaxypθ\(y∣I,q,K\)\\hat\{y\}=\\arg\\max\_\{y\}p\_\{\\theta\}\(y\\mid I,q,K\)\. The knowledge blockK=τ\(𝒮\)K=\\tau\(\\mathcal\{S\}\)deterministically textualizes retrieved records𝒮\\mathcal\{S\}\. Because no parameters are updated, the central design choice is which records to expose to the model\.
### 3\.2Coordinate\-keyed retrieval
Letc\(f\)c\(f\)be the centroid of OSM featureffandκ\(f\)\\kappa\(f\)its primary tag key\. We retrieve every feature withinr¯=clamp\(r,300,1500\)\\bar\{r\}=\\mathrm\{clamp\}\(r,300,1500\)m of\(ϕ,λ\)\(\\phi,\\lambda\), formingℛ\\mathcal\{R\}; full injection uses𝒮=ℛ\\mathcal\{S\}=\\mathcal\{R\}\. Dated Geofabrik extracts and a spatial grid index make retrieval reproducible without live API calls\. On 461 images available through both routes, local and API feature counts correlate at0\.970\.97\. Ablations add the nearest GeoNames toponym and dominant ESA WorldCover class\(Zanagaet al\.,[2022](https://arxiv.org/html/2608.00877#bib.bib46)\)\. Unlike embedding retrieval, coordinates identify ordinary as well as landmark locations and cannot return a visually similar but geographically unrelated site\.
### 3\.3Cross\-modal verifiability and stratified injection
The pivotal property of a fact is whether the image could in principle confirm or refute it\. We assign each OSM key a binary labelvv: keys expressing function or identity \(amenity,shop,tourism,military,office\) takev=0v=0, whereas keys expressing physical structure \(building,highway,aeroway,natural,landuse\) takev=1v=1; Appendix[D](https://arxiv.org/html/2608.00877#A4)gives the complete map and textualization\. Verifiability belongs to the attribute, not the object: a building may be visible while its institutional role is not\.
For attribute classaa, letG\(a\)G\(a\)be the accuracy change from injectingKKinstead of no knowledge\. We test two predictions:
𝔼\[G∣v=0\]\\displaystyle\\mathbb\{E\}\[G\\mid v\{=\}0\]\>𝔼\[G∣v=1\],\\displaystyle\>\\mathbb\{E\}\[G\\mid v\{=\}1\],\(1\)G\(a∣v=1,conflict\)\\displaystyle G\(a\\mid v\{=\}1,\\mathrm\{conflict\}\)<0\.\\displaystyle<0\.Thus, correct records may still provide a useful prior for visible attributes, but their marginal value should be smaller; when such a record conflicts with the image, its effect should be negative\. A policy conditioned onvvshould therefore outperform a fixed preference for either source\.
Verifiability is distinct from both relevance and source reliability\. A nearby runway record is relevant to a runway question, and OSM may be highly accurate on average, but a false runway record should not override an image in which the structure is visibly absent\. Conversely, a school\-function record cannot be confirmed from roof geometry even when the building itself is clear\. We definevvby whether an attribute is in principle recoverable at the image’s modality and scale, not by whether a particular model happens to answer it correctly\. This makes the policy model\-agnostic and prevents weak visual competence from being mistaken for inherent unverifiability\.
Stratified injectionenforces this policy on content, replacing𝒮\\mathcal\{S\}with the image\-unverifiable subset
𝒮−=\{f∈ℛ:v\(κ\(f\)\)=0\},\\mathcal\{S\}^\{\-\}=\\bigl\\\{f\\in\\mathcal\{R\}\\;:\\;v\(\\kappa\(f\)\)=0\\bigr\\\},\(2\)which retains 11\.1% of features\. Records are textualized into a salience\-ordered, 1,200\-character block \(named features, then tag counts\)\. The filter requires no detector, extra model call, or per\-image classifier\.
The key\-level map is intentionally conservative and auditable\. It preserves OSM keys whose values primarily encode use, ownership, service, or identity and removes keys whose values primarily encode visible geometry, surface, or transport structure\. It does not attempt to judge individual record correctness\. This separation lets us test whether controlling exposure alone can outperform asking the same frozen model to reason about source trust after seeing every record\.
Policy summary\.
Function and identity are retained; visible structure is withheld\. Boundary cases are analyzed in §[5\.5](https://arxiv.org/html/2608.00877#S5.SS5)\.
### 3\.4Instruction\-level arbitration \(contrast arm\)
The natural alternative leaves all records in the prompt and states the policy in language\. We test*image\-first*\(“if records conflict with what you see, trust the image”\),*database\-first*,*stratified*\(image for visible structure; records for functions and names\), a question\-keyed rewrite, and two semantic\-preserving paraphrases of the stratified instruction \(Appendix[B](https://arxiv.org/html/2608.00877#A2)\)\. If frozen MLLMs can apply attribute\-conditional trust, at least one of these six variants should reproduce Eq\.[2](https://arxiv.org/html/2608.00877#S3.E2); none does \(§[5\.4](https://arxiv.org/html/2608.00877#S5.SS4)\)\.
### 3\.5Implementation
Retrieval, textualization, and caching run on CPU\. Warm\-cache assembly takes 4\.9 ms per image \(p95 12\.6 ms\); injection adds 413 prompt tokens and 0\.56 s mean generation latency on Qwen2\.5\-VL\-7B\.
Table 1:Core results across three open MLLMs\.Left: land\-use QA accuracy on the full fMoW functional validation split \(n=28,087n\{=\}28\{,\}087per cell\)\. Middle: claim\-level hallucination rate \(contradictedshare\) on 3,000 scene descriptions \(≈\\approx8\.5–9 claims each\)\. Right: existence\-probe conflict\-cell accuracy, where a*fabricated*record is injected and ground truth is*no*\(n=1,069n\{=\}1\{,\}069; full per\-cell probe in Table[2](https://arxiv.org/html/2608.00877#S5.T2)\)\. Full injection \(*rag*\) is nominally best on the clean tasks but*collapses*once a record is wrong \(Qwen79\.3079\.30QA yet72\.0072\.00conflict; LLaVA down to36\.7536\.75\); stratified injection \(*rag\-filtered*\) keeps85%85\\%of the QA gain and, by withholding image\-verifiable records, is the only condition robust in conflict, exceeding even the no\-knowledge cell\.*coords*\(image plus raw latitude/longitude\) is a QA\-only baseline\. All rag−\-none and filtered−\-rag differences are significant \(p<0\.001p\{<\}0\.001paired bootstrap for QA;p<10−4p\{<\}10^\{\-4\}for hallucination and the conflict cell\)\. Per\-model and per\-type breakdowns are in Appendix[F](https://arxiv.org/html/2608.00877#A6)\.
## 4Experimental setup
#### Data\.
Our main corpus is the complete fMoW functional validation split\(Christieet al\.,[2018](https://arxiv.org/html/2608.00877#bib.bib18)\): 28,087 images in 32 function\-defined categories, with WGS84 coordinates and 2002–2017 acquisition times\. We avoid OSM\-derived QA and captions because evaluating OSM\-grounded generation against OSM\-derived labels would be circular\. fMoW and our CORINE\-derived external labels are independent of the injected knowledge\.
#### Models\.
Three open MLLMs spanning architecture families: Qwen2\.5\-VL\-7B\-Instruct\(Qwen Team,[2025](https://arxiv.org/html/2608.00877#bib.bib23)\), InternVL3\-8B\(Zhuet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib24)\), and LLaVA\-OneVision\-7B\(Liet al\.,[2024](https://arxiv.org/html/2608.00877#bib.bib25)\), all frozen, greedy decoding\.
#### Conditions\.
*none*\(image only\);*coords*\(plus raw coordinates\);*rag*\(full OSM,𝒮=ℛ\\mathcal\{S\}=\\mathcal\{R\}\); and*rag\-filtered*\(GeoArbiter,𝒮=𝒮−\\mathcal\{S\}=\\mathcal\{S\}^\{\-\}\)\. Additional comparisons include the six arbitration instructions, Wikipedia geosearch, GeoNames, WorldCover, all sources, and a reproduction of RS\-RAG retrieval\.
#### Land\-use QA and description\.
Land\-use QA is four\-way fMoW classification\. Distractors are sampled with a fixed per\-image seed, so all models and conditions receive identical options; we report accuracy with paired bootstrap tests \(10k resamples\)\. For open generation, models produce 3–5 sentence descriptions on a seeded 3,000\-image subset\. A frozen Qwen2\.5\-7B\-Instruct judge decomposes each output into atomic claims \(8\.5–9 per description on average\), assigns a type \(function,name, physicalcontext, or other\), and labels itsupported,contradicted, orunverifiableagainst the image label and retrieved reference\. Hallucination rate is thecontradictedshare; prompts and the judge template are in Appendices[C](https://arxiv.org/html/2608.00877#A3)and[E](https://arxiv.org/html/2608.00877#A5), and human validation is in §[5\.6](https://arxiv.org/html/2608.00877#S5.SS6)\.
#### Verifiable\-existence probe\.
To isolate the regime in which the image should dominate, we build 1,069 yes/no items from 400 images\. The*conflict*cell adds a fabricated record for a manually verified absent, visually distinctive feature\.*Control\-absent*asks about the same absence without fabrication, while*control\-present*asks about a manually verified present feature\. The former separates susceptibility to injected records from ordinary false positives; the latter reveals instructions that improve conflict accuracy merely by shifting the response prior toward*no*\. Historical and current OSM snapshots ensure that each selected feature’s map status is persistent, but ground truth comes from image inspection rather than OSM absence\.
#### Derived metrics\.
The*retained gain*ρ=\(afilt−a∅\)/\(arag−a∅\)\\rho=\(a\_\{\\mathrm\{filt\}\}\-a\_\{\\varnothing\}\)/\(a\_\{\\mathrm\{rag\}\}\-a\_\{\\varnothing\}\)measures how much full\-retrieval QA gain survives filtering\. Because instructions can shift the yes/no prior, we also report probe balanced accuracyBA=12\(aconf\+apres\)\\mathrm\{BA\}=\\tfrac\{1\}\{2\}\(a\_\{\\mathrm\{conf\}\}\+a\_\{\\mathrm\{pres\}\}\)rather than interpreting the conflict cell alone\.
#### External benchmark\.
We construct 2,000 four\-way land\-cover questions from a pan\-European BigEarthNet subset\(Sumbulet al\.,[2019](https://arxiv.org/html/2608.00877#bib.bib19)\)via GEO\-Bench\(Lacosteet al\.,[2023](https://arxiv.org/html/2608.00877#bib.bib45)\), preserving patch coordinates\. RSHBench\(Liuet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib2)\)and RSHalluEval\(Zhouet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib3)\)were unreleased at submission time and are used only as qualitative context\.
## 5Results
### 5\.1Structured retrieval lifts land\-use QA, and content is what matters
Coordinate\-keyed OSM injection raised accuracy by\+16\.52\+16\.52,\+17\.19\+17\.19, and\+12\.06\+12\.06points for Qwen2\.5\-VL, InternVL3, and LLaVA\-OneVision \(Table[1](https://arxiv.org/html/2608.00877#S3.T1); allp<0\.001p\{<\}0\.001\)\. Raw coordinates added at most\+1\.66\+1\.66points; retrieval’s gain was 10\.36–16\.30×\\timeslarger, ruling out location disclosure\. Three controls further localize the gain: an image\-blind OSM\-tag decoder tops out at 70\.6%, and blanking the image or pairing it with a*different*image’s knowledge drops accuracy back toward*none*, so the gain needs coordinate\-matched knowledge bound to the image, not tag leakage \(Appendix[H](https://arxiv.org/html/2608.00877#A8)\); at a matched record budget, verifiability\-selected injection also retains far more of it than random, frequency, or inverted filtering \(Appendix[G](https://arxiv.org/html/2608.00877#A7)\)\. LLaVA\-OneVision, the strongest baseline, gained least, as expected if retrieval fills a knowledge gap rather than adding a generic prior\.
### 5\.2Injection reduces claim\-level hallucination where verifiability predicts
Full injection cut claim\-level hallucination by 51\.74%, 23\.93%, and 23\.56% \(Table[1](https://arxiv.org/html/2608.00877#S3.T1)\), with stable claim counts\. Reductions concentrated on image\-unverifiable content \(Figure[3](https://arxiv.org/html/2608.00877#S5.F3); all models: Table[F\.2](https://arxiv.org/html/2608.00877#A6.T2)\)\. For Qwen2\.5\-VL,functionfell from 15\.01% to 7\.63%, andnamefrom10\.29%10\.29\\%to2\.59%2\.59\\%despite 2–5×\\timesmore naming claims; physical\-contextbegan at 1\.01% and changed little\. Theunverifiableshare also fell from 51\.76% to 32\.71%, indicating more checkable rather than merely shorter descriptions\. These rates use a judge that sees the injected records; under a source\-blinded judge the*none*→\\tofiltered reduction is 9\.58–26\.34% \(vs\. 19\.28–45\.19%\) and absolute rates roughly double, so we treat injected\-reference values as a lower bound and report the blinded figure \(Appendix[I](https://arxiv.org/html/2608.00877#A9)\)\.
Figure 3:Injection helps most on image\-unverifiable claims\.Claim\-level hallucination rates over 3,000 Qwen2\.5\-VL descriptions before retrieval \(hollow\) and after full OSM injection \(filled\)\. Arrows report relative reductions: function and name claims show the largest absolute gains, while visually verifiable context starts near the floor\.
### 5\.3Retrieval changes many answers, and the database usually wins on merit
Injection changed 27%, 28%, and 18% of model answers\. An answer change is a superset of a genuine image–database conflict—it also captures cases where the database merely supplies information the image cannot show—so we read it as a retrieval\-induced answer change rather than a verified natural conflict\. Among these changes, records corrected the image\-only answer 72–77% of the time and misled it 10–11%\. The most frequently rescued classes—places of worship, emergency services, gas stations, and schools—are defined by functions that are invisible from orbit \(Figure[4](https://arxiv.org/html/2608.00877#S5.F4)a\)\. Natural function conflicts therefore favor the database; the controlled existence probe tests the complementary, visually verifiable regime\.
### 5\.4The verifiable side: fabricated records fool every model, and instructions cannot fix it
Table 2:Existence probe, per\-cell accuracy in percent \(conf: conflict, GT*no*; abs: control\-absent, GT*no*; pres: control\-present, GT*yes*;n=1,069n\{=\}1\{,\}069\)\. Balanced accuracy is12\\tfrac\{1\}\{2\}\(conf\+\{\+\}pres\)\. Fabricated records cost the conflict cell; no instruction restores it, and stratified injection exceeds the no\-knowledge conflict cell for every model\.Fabricated records cut conflict\-cell accuracy from 88\.75% to 72\.00% for Qwen2\.5\-VL, 75\.75% to 62\.25% for InternVL3, and 87\.75% to 36\.75% for LLaVA\-OneVision \(allp<10−4p\{<\}10^\{\-4\}\): losses of 13\.50–51\.00 points\. This is the multimodal counterpart of context over\-reliance in text RAG\(Ouyang and others,[2025](https://arxiv.org/html/2608.00877#bib.bib32); Khayatanet al\.,[2026](https://arxiv.org/html/2608.00877#bib.bib42)\)\.
Arbitration prompts partially mitigated conflict but had two systematic defects\. First,*leakage*: every image\-trust instruction reduced all\-function QA by 0\.50–3\.25 points on Qwen2\.5\-VL; the bootstrap probability of a decrease exceeded 0\.999 for the worst variant\. Only database\-first helped \(\+1\.50\+1\.50,p=0\.049p\{=\}0\.049\)\. Second,*response bias*: instructions shifted the yes/no prior toward*no*, improving conflict scores while reducing control\-present accuracy\. Consequently, no instruction recovered the no\-knowledge balanced accuracy \(Table[2](https://arxiv.org/html/2608.00877#S5.T2); full grid: Table[F\.1](https://arxiv.org/html/2608.00877#A6.T1)\)\. Frozen models do not reliably execute attribute\-conditional trust from language alone\.
### 5\.5Stratified injection wins both sides
Content\-level filtering discarded 88\.9% of features yet retained 84\.69–87\.15% of QA gain and 81\.82–91\.09% of hallucination reduction, with at most a 0\.32\-point cost \(Table[1](https://arxiv.org/html/2608.00877#S3.T1)\)\. Probe conflict accuracy reached 93\.25%, 90\.50%, and 91\.25%; balanced accuracy was 80\.08%, 80\.54%, and 82\.02%, above the no\-knowledge baseline for every model \(Table[2](https://arxiv.org/html/2608.00877#S5.T2)\)\. Figure[4](https://arxiv.org/html/2608.00877#S5.F4)b–c shows GeoArbiter withholding fabricated and structural records while retaining function records\. The diagnostic exception is visibleamenity=parking, which survives the filter and fools models in 83–92% of cases\. Robustness comes from withholding records, not model\-side arbitration: at a matched record budget, keeping the image\-*verifiable*records instead \(inverted filter\) drops conflict accuracy to 20\.5–52\.8%, below full injection, whereas random selection stays robust—the direction of selection, not the reduced context, is the lever \(Appendix[G](https://arxiv.org/html/2608.00877#A7)\)\.
Figure 4:Three cases\(unedited Qwen2\.5\-VL\-7B outputs\)\.a, Function is invisible from orbit: the model guesses*stadium*; a retrieved*police*record \(‘Guardia Civil’\) settles it, and stratified injection keeps that record\.b, Existence is visible: a planted*helipad*record flips the answer to*yes*, and withholding the image\-verifiable record restores*no*\.c, The filter drops the bulk road and building geometry while injecting the two function records that identify the site\.
### 5\.6Measurement validity: human agreement
Two annotators assessed a stratified 198\-claim sample \(balanced across judge labels, models, and conditions\), blind to the judge\. Initial agreement was 78\.3% \(κ=0\.39\\kappa\{=\}0\.39\); most of the 43 disagreements concerned two underspecified cases \(hedged wrong guesses; claims restating a reference category\)\. After clarifying that hedging does not excuse asserted content and that a correct restatement is not a contradiction, the annotators resolved 28 cases and excluded 15 perceptual ones\. Against the 183 adjudicated labels the judge reached 84\.2% agreement \(κ=0\.61\\kappa\{=\}0\.61, recall 0\.88, precision 0\.60\), so we treat absolute rates as conservative upper bounds and emphasize paired differences\. The modest initial agreement remains a limitation \(§[Limitations](https://arxiv.org/html/2608.00877#Sx1)\)\.
### 5\.7Baselines and source ablations
On a seeded 3,000\-image subset \(full table in Appendix[F](https://arxiv.org/html/2608.00877#A6), Table[F\.3](https://arxiv.org/html/2608.00877#A6.T3)\), structured OSM \(ours\) reaches 80\.57/76\.50/83\.33% vs\. 69\.13/65\.40/73\.43% for Wikipedia GeoSearch\(MediaWiki,[2026](https://arxiv.org/html/2608.00877#bib.bib49)\), 65\.83/62\.77/72\.53% for GeoNames\(GeoNames,[2026](https://arxiv.org/html/2608.00877#bib.bib48)\), and 58\.50/60\.07/70\.43% for a faithful RS\-RAG reproduction\(Wenet al\.,[2025](https://arxiv.org/html/2608.00877#bib.bib1)\)over its released 14,820\-landmark base\(Haklay and Weber,[2008](https://arxiv.org/html/2608.00877#bib.bib47)\)\. RS\-RAG was neutral to harmful on ordinary scenes \(−5\.30\-5\.30points for Qwen\): retrieved landmarks looked similar \(0\.90±0\.030\.90\{\\pm\}0\.03\) but had the wrong identity\. Wikipedia GeoSearch added 2\.96–7\.17 points; OSM gains were 2\.55–4\.34×\\timeslarger\. GeoNames added 2\.03–4\.54 points and WorldCover\(Zanagaet al\.,[2022](https://arxiv.org/html/2608.00877#bib.bib46)\)1\.43–5\.30, while combining them with OSM changed accuracy by only−0\.63\-0\.63to\+0\.76\+0\.76points\.
### 5\.8External validity and cost
On BigEarthNet land\-cover QA, injection improved all three models by 7\.95, 4\.25, and 4\.10 points \(Table[F\.4](https://arxiv.org/html/2608.00877#A6.T4)\)\. These smaller gains are consistent with Eq\.[1](https://arxiv.org/html/2608.00877#S3.E1): land cover is largely image\-verifiable, leaving less missing knowledge to supply\. Deployment overhead is 4\.9 ms warm\-cache assembly, 413 prompt tokens, and 0\.56 s mean generation latency on Qwen2\.5\-VL\-7B\.
## 6Analysis: why the axis is verifiability, not time
### 6\.1Why content\-level arbitration succeeds
The two task families expose an incompatibility no global source preference can solve\. Function\-defined fMoW questions favor the records because the decisive attribute is usually invisible, so image\-first language removes useful evidence and lowers accuracy; existence questions under conflict favor the image because the disputed structure is directly observable, so database\-first behavior produces false positives\. A correct policy must switch at the attribute level, not the image, question, or database level\.
The instruction experiments show why stating this switch is insufficient: even the stratified and question\-keyed prompts leave the conflicting record in context and require the frozen model to identify the disputed attribute, classify its verifiability, and inhibit a salient textual assertion\. Their leakage on function QA shows the instruction acts partly as a global image\-trust prior; the drop on control\-present items shows an added*no*\-response bias\. Balanced accuracy exposes both, which reporting the conflict cell alone would obscure\.
GeoArbiter instead compiles the policy into the input: an image\-verifiable attribute’s record is absent and cannot compete with perception, while an unverifiable one remains available without a weakening instruction\. This explains the otherwise striking combination—retaining only 11\.1% of records preserves 84\.69–87\.15% of the land\-use gain, while conflict balanced accuracy exceeds the image\-only baseline for every model\. The filter does not make the MLLM a better arbiter; it removes the need for arbitration\.
### 6\.2What the controlled conflict establishes
The fabricated\-record probe is not meant to estimate the prevalence of OSM errors; it is a controlled intervention on the evidence channel, holding image and question fixed while the asserted record is added or removed\. The control\-absent cell estimates the model’s ordinary false\-positive tendency, and control\-present measures whether an intervention shifts the response prior\. The large conflict\-only loss under full injection thus identifies exposure to the contradictory record as the mechanism, not image difficulty or a general*yes*tendency\. Retrieval\-induced answer changes give the complementary ecological result—when disagreements concern function, records usually help—so the two analyses locate the boundary of useful grounding rather than merely showing retrieval can help or hurt\.
### 6\.3Temporal mismatch does not explain the gap
Archival imagery \(2002–2017\) paired with a current map suggests temporal staleness as a rival explanation\. Reconstructing acquisition\-time OSM, we compute the map–image divergenceΔ=1−\(\|Ft∩F0\|−\|C\|/2\)/\|Ft∪F0\|\\Delta=1\-\(\|F\_\{t\}\\cap F\_\{0\}\|\-\|C\|/2\)/\|F\_\{t\}\\cup F\_\{0\}\|\(FtF\_\{t\}/F0F\_\{0\}: historical/current features;CC: shared features with changed tags\)\. Retrieval gain was flat acrossΔ\\Deltaquantiles for both pilot models \(−0\.01\-0\.01to−0\.06\-0\.06,p≥0\.61p\{\\geq\}0\.61\), and current maps outperformed sparser time\-matched snapshots; a blinded audit explains why: 93\.9% of features mapped after acquisition were already visible in the older image, and only about 1 in 22 scenes showed new construction\. On this corpus the OSM–image delta therefore primarily measures mapping completion, not physical change; the audit is modest and we do not claim this generalizes beyond it\.
### 6\.4Where the filter still costs
Key\-level stratification drops useful evidence expressed through otherwise verifiable keys, such as hangars underaerowayand offices underbuilding, accounting for the 12\.85–15\.31% loss in retained gain; categories with strongamenityevidence lose nothing\. The natural refinement is an attribute\-level policy over tag values or question–attribute pairs—recovering functional evidence underbuilding/aerowaywhile withholding visibly checkable values such asamenity=parking—still within a deterministic, training\-free schema map\.
## 7Conclusion
Coordinate\-keyed geographic knowledge improves RS MLLMs but harms when it contradicts image\-verifiable attributes\. Cross\-modal verifiability predicts this boundary, which frozen models cannot enforce from instructions; GeoArbiter instead filters content, keeping 84\.69–87\.15% of the QA gain and 81\.82–91\.09% of the hallucination reduction while staying robust to fabricated records\. Verifiability\-guided selection is a strong, training\-free baseline for grounding VLMs in fallible structured knowledge\.
## Limitations
Our verifiability map is a key\-level approximation: a parking lot is anamenitybut plainly visible, while some useful functional evidence appears under structural keys\. The automatic hallucination judge has only moderate agreement with adjudicated human labels \(κ=0\.61\\kappa\{=\}0\.61\), and initial human agreement was also modest \(κ=0\.39\\kappa\{=\}0\.39\); absolute rates should therefore be read as conservative upper bounds, although all comparisons use the same judge\. Probe conflicts are constructed rather than naturally observed, so they measure susceptibility under controlled contradictions rather than prevalence in deployed OSM\. Results cover fMoW, BigEarthNet, English prompts, and three 7–8B open models; RS\-tuned or larger models may follow different source priors\. Two unreleased RS hallucination benchmarks limit comparison\.
## Ethical Considerations
This work uses public datasets, open\-weight models, and geographic records under their respective licenses\. Imagery comes from fMoW and BigEarthNet under their released research terms, and we use SkyScript images only \(not its OSM\-derived captions\)\. Structured knowledge comes from OpenStreetMap \(ODbL 1\.0\), GeoNames \(CC\-BY 4\.0\), and ESA WorldCover \(CC\-BY 4\.0\)\. Because ODbL is an attribution and share\-alike licence, any OSM\-derived caches we release carry the OpenStreetMap attribution and the same licence, and we redistribute derived features rather than raw third\-party imagery\. Retrieval operates only on coordinates already distributed with the benchmarks and describes public infrastructure rather than individuals\. Nevertheless, joining overhead imagery to precise geographic records can support surveillance as well as civilian land\-use analysis; deployments should therefore apply access controls and purpose\-specific review\.
Two project members performed the claim and existence checks\. Because the task labels model outputs rather than studying the annotators, it is not human\-subjects research and required no ethics approval; the annotators participated voluntarily, with informed consent, and were uncompensated\. They viewed no sensitive content\. The annotation items and guidelines accompany the submission; generated descriptions were used only for verification and are not redistributed as ground truth\.
## References
- S\. M\. Ailuro, M\. Markov, M\. Mahdi,et al\.\(2026\)OSMDA: openstreetmap\-based domain adaptation for remote sensing vlms\.External Links:2603\.11804,[Link](https://arxiv.org/abs/2603.11804)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Anderson, M\. Cha, W\. T\. Freeman,et al\.\(2025\)Measuring and mitigating hallucinations in vision\-language dataset generation for remote sensing\.External Links:2501\.14905Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Asai, Z\. Wu, Y\. Wang, A\. Sil, and H\. Hajishirzi \(2024\)Self\-rag: learning to retrieve, generate, and critique through self\-reflection\.InICLR,Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- L\. Bai, X\. Zhang, S\. Zhang,et al\.\(2025\)GeoLink: empowering remote sensing foundation model with openstreetmap data\.Note:NeurIPS 2025External Links:2509\.26016,[Link](https://arxiv.org/abs/2509.26016)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- P\. Carragher, N\. Rao, A\. Jha,et al\.\(2025\)SegSub: evaluating robustness to knowledge conflicts and hallucinations in vision\-language models\.Note:MisD Workshop 2025External Links:2502\.14908,[Link](https://arxiv.org/abs/2502.14908)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- G\. Chen, C\. Huang, Y\. Yao,et al\.\(2026\)From scenes to elements: multi\-granularity evidence retrieval for verifiable multimodal rag\.External Links:2605\.15019Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- G\. Christie, N\. Fendley, J\. Wilson, and R\. Mukherjee \(2018\)Functional map of the world\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px1.p1.1)\.
- J\. Feng, Y\. Du, T\. Liu,et al\.\(2025\)GeoRAG: a geographic retrieval augmented generation framework based on urban spatio\-temporal knowledge graph\.InSpringer \(CCIS/LNCS\),Note:doi:10\.1007/978\-981\-95\-4788\-3\_12Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- GeoNames \(2026\)GeoNames Geographical Database\.Note:Accessed 28 July 2026External Links:[Link](https://www.geonames.org/about.html)Cited by:[§5\.7](https://arxiv.org/html/2608.00877#S5.SS7.p1.5)\.
- M\. Haklay and P\. Weber \(2008\)OpenStreetMap: user\-generated street maps\.IEEE Pervasive Computing7\(4\),pp\. 12–18\.External Links:[Document](https://dx.doi.org/10.1109/MPRV.2008.80)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§5\.7](https://arxiv.org/html/2608.00877#S5.SS7.p1.5)\.
- J\. Han, C\. Li, C\. Hu,et al\.\(2026\)From clouds to hallucinations: atmospheric retrieval hijacking in remote sensing vision\-language rag\.External Links:2605\.07273Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1)\.
- S\. Jeong, J\. Baek, S\. Cho, S\. J\. Hwang, and J\. C\. Park \(2024\)Adaptive\-rag: learning to adapt retrieval\-augmented large language models through question complexity\.InNAACL,Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- Y\. Jia, K\. Jiang, Y\. Liang,et al\.\(2025\)Benchmarking multimodal knowledge conflict for large multimodal models\.External Links:2505\.19509,[Link](https://arxiv.org/abs/2505.19509)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- P\. Khayatan, J\. Parekh, A\. Dapogny,et al\.\(2026\)When prompts override vision: prompt\-induced hallucinations in lvlms\.External Links:2604\.21911Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1),[§5\.4](https://arxiv.org/html/2608.00877#S5.SS4.p1.1)\.
- A\. Lacoste, N\. Lehmann, P\. Rodriguez,et al\.\(2023\)GEO\-bench: toward foundation models for earth monitoring\.InNeurIPS Datasets and Benchmarks,Cited by:[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px7.p1.1)\.
- B\. Li, Y\. Zhang, D\. Guo,et al\.\(2024\)LLaVA\-onevision: easy visual task transfer\.External Links:2408\.03326,[Link](https://arxiv.org/abs/2408.03326)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px2.p1.1)\.
- Y\. Li, Y\. Du, K\. Zhou, J\. Wang, W\. X\. Zhao, and J\. Wen \(2023\)Evaluating object hallucination in large vision\-language models\.InEMNLP,Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px1.p1.1)\.
- N\. Lietzow, D\. Bitterman, C\. Eickhoff,et al\.\(2026\)Vision\-default, prior\-override: causal mechanisms of perception\-knowledge conflict in vision\-language models\.External Links:2606\.28273Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- X\. Liu, W\. Wang, Y\. Yuan,et al\.\(2024\)Insight over sight: exploring the vision\-knowledge conflicts in multimodal llms\.Note:ACL 2025External Links:2410\.08145,[Link](https://arxiv.org/abs/2410.08145)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- Y\. Liu, J\. Zhang, D\. Wang,et al\.\(2026\)Seeing clearly without training: mitigating hallucinations in multimodal llms for remote sensing\.External Links:2603\.02754,[Link](https://arxiv.org/abs/2603.02754)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px7.p1.1)\.
- R\. Manvi, S\. Khanna, G\. Mai, M\. Burke, D\. Lobell, and S\. Ermon \(2024\)GeoLLM: extracting geospatial knowledge from large language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- MediaWiki \(2026\)API:Geosearch\.Note:Accessed 28 July 2026External Links:[Link](https://www.mediawiki.org/wiki/API:Geosearch)Cited by:[§5\.7](https://arxiv.org/html/2608.00877#S5.SS7.p1.5)\.
- D\. Muhtar, Z\. Li, F\. Gu, X\. Zhang, and P\. Xiao \(2024\)LHRS\-bot: empowering remote sensing with vgi\-enhanced large multimodal language model\.InEuropean Conference on Computer Vision \(ECCV\),Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- F\. Ortu, Z\. Jin, D\. Doimo,et al\.\(2025\)When seeing overrides knowing: disentangling knowledge conflicts in vision\-language models\.Note:ACL 2026External Links:2507\.13868,[Link](https://arxiv.org/abs/2507.13868)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- J\. Ouyanget al\.\(2025\)HoH: a dynamic benchmark for evaluating the impact of outdated information on retrieval\-augmented generation\.InACL,External Links:2503\.04800Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1),[§5\.4](https://arxiv.org/html/2608.00877#S5.SS4.p1.1)\.
- Qwen Team \(2025\)Qwen2\.5\-vl technical report\.External Links:2502\.13923,[Link](https://arxiv.org/abs/2502.13923)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px2.p1.1)\.
- D\. F\. Ramirez, T\. Overman, K\. Jaskie,et al\.\(2026\)SAR\-rag: atr visual question answering by semantic search, retrieval, and mllm generation\.Note:SPIE DCS ATR XXXVI 2026External Links:2602\.04712Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px2.p1.1)\.
- W\. Shi, X\. Han, M\. Lewis, Y\. Tsvetkov, L\. Zettlemoyer, and S\. W\. Yih \(2023\)Trusting your evidence: hallucinate less with context\-aware decoding\.External Links:2305\.14739Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p3.1)\.
- G\. Sumbul, M\. Charfuelan, B\. Demir, and V\. Markl \(2019\)BigEarthNet: a large\-scale benchmark archive for remote sensing image understanding\.InIEEE International Geoscience and Remote Sensing Symposium \(IGARSS\),Cited by:[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px7.p1.1)\.
- Z\. Wang, R\. Prabha, T\. Huang, J\. Wu, and R\. Rajagopal \(2024\)SkyScript: a large and semantically diverse vision\-language dataset for remote sensing\.InProceedings of the AAAI Conference on Artificial Intelligence,Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Wen, Y\. Lin, X\. Qu,et al\.\(2025\)Remote sensing retrieval\-augmented generation: bridging remote sensing imagery and comprehensive knowledge with a multi\-modal dataset and retrieval\-augmented generation model\.External Links:2504\.04988,[Link](https://arxiv.org/abs/2504.04988)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px2.p1.1),[§5\.7](https://arxiv.org/html/2608.00877#S5.SS7.p1.5)\.
- Y\. Xiao, K\. Yin, and A\. Mostafavi \(2026\)CrisiSense\-rag: crisis sensing multimodal retrieval\-augmented generation for rapid disaster impact assessment\.External Links:2602\.13239,[Link](https://arxiv.org/abs/2602.13239)Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Xue, Q\. Deng, T\. Hu,et al\.\(2026\)AeroRAG: structured multimodal retrieval\-augmented llm for fine\-grained aerial visual reasoning\.External Links:2604\.17889Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Yan, J\. Gu, Y\. Zhu, and Z\. Ling \(2024\)Corrective retrieval augmented generation\.External Links:2401\.15884Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px4.p1.1)\.
- S\. Yin, C\. Fu, S\. Zhao,et al\.\(2023\)Woodpecker: hallucination correction for multimodal large language models\.External Links:2310\.16045Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px1.p1.1)\.
- D\. Yu, R\. Bao, R\. Ning,et al\.\(2025\)Spatial\-rag: spatial retrieval augmented generation for real\-world geospatial reasoning questions\.External Links:2502\.18470,[Link](https://arxiv.org/abs/2502.18470)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p2.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Zanaga, R\. Van De Kerchove, D\. Daems,et al\.\(2022\)ESA WorldCover 10 m 2021 v200\.Note:Zenododoi:10\.5281/zenodo\.7254221Cited by:[§3\.2](https://arxiv.org/html/2608.00877#S3.SS2.p1.8),[§5\.7](https://arxiv.org/html/2608.00877#S5.SS7.p1.5)\.
- S\. Zhao, F\. Liu, X\. Zhang,et al\.\(2026\)OpenEarth\-agent: from tool calling to tool creation for open\-environment earth observation\.External Links:2603\.22148,[Link](https://arxiv.org/abs/2603.22148)Cited by:[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Zhou, Y\. Feng, Y\. Chen,et al\.\(2026\)RSHallu: dual\-mode hallucination evaluation for remote\-sensing multimodal large language models with domain\-tailored mitigation\.External Links:2602\.10799,[Link](https://arxiv.org/abs/2602.10799)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§2](https://arxiv.org/html/2608.00877#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px7.p1.1)\.
- J\. Zhu, W\. Wang, Z\. Chen,et al\.\(2025\)InternVL3: exploring advanced training and test\-time recipes for open\-source multimodal models\.External Links:2504\.10479,[Link](https://arxiv.org/abs/2504.10479)Cited by:[§1](https://arxiv.org/html/2608.00877#S1.p1.1),[§4](https://arxiv.org/html/2608.00877#S4.SS0.SSS0.Px2.p1.1)\.
## Appendix AReproducibility details
All prompts \(task instructions, the six arbitration wordings of Appendix[B](https://arxiv.org/html/2608.00877#A2), judge template\), the knowledge textualization format, per\-image caches, and Slurm run scripts are released\. Local OSM extraction uses dated Geofabrik continental extracts \(2026\-07\-25\) with centroid\-in\-bbox semantics matching the ohsome centroid endpoint used in the pilot; validation against 461 API\-fetched images gives count correlation 0\.97 \(median local/API ratio 0\.92; relations are skipped\)\. The RS\-RAG reproduction and its documented deviations \(unreleased monthly\-knowledge and image\-description chunks\) are detailed in the released notes\. Table[A\.1](https://arxiv.org/html/2608.00877#A1.T1)lists the remaining settings needed to reproduce every number\.
Table A\.1:Reproducibility settings \(companion to the released code and caches\)\.
## Appendix BArbitration instruction wordings
The instruction\-level arbitration arm \(§[3\.4](https://arxiv.org/html/2608.00877#S3.SS4)\) appends one of six trust instructions to the knowledge block, listed verbatim below\. Instructions 1–2 name a global winner; instruction 3 states the verifiability rule keyed on object class, and instructions 4–5 are two paraphrases of it that control for wording; instruction 6 \(question\-keyed\) conditions on the disputed attribute rather than the object class\. The results in §[5\.4](https://arxiv.org/html/2608.00877#S5.SS4)hold across all six\.
1. 1\.Image\-first\.“If the database records conflict with what you see in the image, trust the image\.”
2. 2\.Database\-first\.“If the database records conflict with what you see in the image, trust the database records\.”
3. 3\.Stratified \(object\-keyed\)\.“If the database records conflict with what you see, decide by the type of information: for physically visible things \(buildings, roads, water, runways\), trust the image; for functions and names that cannot be verified visually \(e\.g\. whether a building is a school or an office, place names\), trust the database records\.”
4. 4\.Stratified, paraphrase 1\.“When sources disagree, apply this rule: believe the image for physically observable features such as buildings, roads, water bodies and runways; believe the database for attributes the image cannot verify, such as a building’s function or a place name\.”
5. 5\.Stratified, paraphrase 2\.“Resolve conflicts by information type: visual evidence wins for anything directly observable \(structures, roads, water, runways\); database records win for visually unverifiable facts \(facility functions, names\)\.”
6. 6\.Question\-keyed\.“If the database records conflict with what you see, decide by what is in dispute: if the question is whether something exists or how it is laid out \(a runway, a pool, a building being there\), trust the image; if the question is what a facility is for or what it is called \(school vs office, place names\), trust the database records\.”
## Appendix CPrompt templates
Braced placeholders are filled per image; the six arbitration wordings are in Appendix[B](https://arxiv.org/html/2608.00877#A2)\.
#### Land\-use QA\.
“Look at the satellite image and answer the question\. Question: What is the primary land use or function of the main facility shown in this image? Options: \{options\} Answer with the letter of the single best option\.”
#### Scene description\.
“Describe this satellite image in 3–5 sentences: the main facility or land use, notable objects, and spatial layout\. Only state what you can support; if the location or a name is uncertain, say so\.”
#### Knowledge preamble
\(prepended for*rag*\)\. “Geographic database records for this image’s location \(may be outdated or incomplete\): \{knowledge\}”
#### Coordinate preamble
\(for*coords*\)\. “This image was taken at latitude \{lat\}, longitude \{lon\}\.”
## Appendix DVerifiability keys and textualization
#### Verifiability map\.
A retrieved feature is image\-*un*verifiable \(v=0v\{=\}0, injected\) when its primary OSM key is one ofamenity,shop,tourism,military,office,healthcare,craft,religion,operator,brand\. All other keys \(building,highway,aeroway,natural,landuse,railway,man\_made,leisure,power,waterway,sport, …\) are image\-verifiable \(v=1v\{=\}1\) and withheld by stratified injection, which retains 11\.1% of retrieved features on our data\.
#### Textualization\.
Records are serialized into a character\-budgeted block \(≤1,200\\leq 1\{,\}200chars\) in salience order: an optional land\-cover line \(WorldCover\) and nearest\-toponym line \(GeoNames\), then named features, then per\-type feature counts\. Features are ordered by key:aeroway,amenity,railway,military,man\_made,leisure,tourism,shop,power,waterway,natural,landuse,highway,building,sport\.
## Appendix EHallucination judging protocol
A frozen Qwen2\.5\-7B\-Instruct judge decomposes each description into atomic claims and labels each with a type and verdict: “You are auditing a satellite\-image description for factual errors\. Reference facts: the image shows a facility of category \{category\}; \{knowledge block\}\. \[description\]\. List every distinct factual claim; for each output one lineCLAIM:…\\ldots\| TYPE: FUNCTION/NAME/CONTEXT/OTHER \| VERDICT: SUPPORTED/CONTRADICTED/UNVERIFIABLE\. Judge strictly against the reference facts; useunverifiablewhen the reference is silent\.” The hallucination rate is thecontradictedshare of all claims; human validation of the judge is in §[5\.6](https://arxiv.org/html/2608.00877#S5.SS6)\.
## Appendix FExtended results
All tables are recomputed from the released per\-image outputs\.
Table F\.1:Instruction\-level arbitration accuracy \(%\) on the all\-function QA pilot \(n=400n\{=\}400\); parenthesized differences are percentage\-point changes from no\-instruction rag\. Every image\-trust instruction hurts; only database\-first helps\. LLaVA\-OV was not run in the pilot sweep\.Table F\.2:Claim\-level hallucination rate by knowledge type \(contradictedshare of that type’s claims\)\. Injection helps most on the least verifiable types; Figure[3](https://arxiv.org/html/2608.00877#S5.F3)plots Qwen\.Table F\.3:Retrieval baselines and single\-source ablations \(3,000\-image subset; full version of the baselines in §[5\.7](https://arxiv.org/html/2608.00877#S5.SS7)\)\. WC: WorldCover\-only; GN: GeoNames\-only; OSM: coordinate\-keyed structured OSM \(ours\); all: OSM\+\{\+\}GN\+\{\+\}WC\. Structured OSM dominates; adding GeoNames and WorldCover changes little\.Table F\.4:External BigEarthNet land\-cover QA accuracy \(%;n=2,000n\{=\}2\{,\}000\)\. Gains are smaller than on fMoW, consistent with Eq\.[1](https://arxiv.org/html/2608.00877#S3.E1): land\-cover classes are largely image\-verifiable\.
## Appendix GVerifiability vs\. retrieval budget
To separate*what*is injected from*how much*, we compare the verifiability filter against three controls that keep the*same*number of records per image \(11\.4% of the retrieved set, matched to the record\):*random*keeps a random subset;*freq*keeps the commonest tag\-key records;*inverted*keeps image\-*verifiable*records \(the complement of our rule\)\. Table[G\.1](https://arxiv.org/html/2608.00877#A7.T1)scores all of them\. On QA accuracy the filter retains≈\\approx85% of the full\-injection gain while the budget\-matched controls retain only 28–43%; on the conflict probe the filter and random stay robust but*inverted collapses below full injection*, because at identical budget it keeps the fabricated image\-verifiable record \(fabrication survival: filter∼0/400\{\\sim\}0/400, random83/40083/400, inverted388/400388/400\)\. Selection by verifiability, not context size, is the lever on both sides\. A rule\-based OSM\-tag→\\rightarrowlabel decoder that never sees the image \(generous keyword matching\) tops out at 70\.6% on the same four\-way QA, below every model’s*rag*accuracy, so the gain is not the tags spelling out the label\. All confidence intervals in Tables[1](https://arxiv.org/html/2608.00877#S3.T1)and[G\.1](https://arxiv.org/html/2608.00877#A7.T1)are location\-clustered block bootstrap \(resampling wholecategory\_LOCgroups\); every headline gap survives clustering atp<10−4p\{<\}10^\{\-4\}\.
Table G\.1:Budget\-matched control filters \(Qw: Qwen2\.5\-VL, In: InternVL3, Lv: LLaVA\-OV\)\. All non\-*none*/*full*rows keep the same 11\.4% record budget\. Conflict is the existence\-probe conflict cell \(n=1,069n\{=\}1\{,\}069\);*freq*was not run on the probe\. Full injection is best on clean QA but collapses in conflict; the verifiability filter is the only budget\-respecting condition strong on both\.
## Appendix HModality isolation
Table[H\.1](https://arxiv.org/html/2608.00877#A8.T1)isolates the image’s contribution\.*knowledge only*replaces the image with a blank frame;*shuffle image*pairs each real image with a*different*image’s retrieved knowledge\. Text alone carries signal \(*knowledge only*\>\>*none*\), but the real\-image condition beats it, and*shuffle image*collapses back to the no\-knowledge baseline: the gain requires coordinate\-matched knowledge bound to the image, not free\-floating text, ruling out pure label\-leakage from the tags\. Permuting the coordinates likewise left accuracy at the*none*level \(InternVL360\.360\.3\)\.
Table H\.1:Modality\-isolation ablations on the full split \(n=28,087n\{=\}28\{,\}087\)\. Cross\-modal binding, not text leakage: mismatched knowledge \(*shuffle image*\) returns to*none*, while matched*rag*exceeds both\.
## Appendix ISource\-blinded hallucination
The judge in Table[1](https://arxiv.org/html/2608.00877#S3.T1)sees the injected records in its reference, which can mark a claimsupportedmerely because judge and model saw the same record\. Table[I\.1](https://arxiv.org/html/2608.00877#A9.T1)re\-runs the judge with the injected knowledge hidden \(category\-only reference\)\. The*none*→\\rightarrowfiltered reduction persists under blinding but shrinks: the 19\.28–45\.19% reported with the standard judge becomes 9\.58–26\.34%, and absolute rates roughly double once the judge cannot lean on the injected text\. Under blinding, full and filtered injection are within0\.30\.3points, so the filter’s advantage is in conflict robustness and retained accuracy \(Appendix[G](https://arxiv.org/html/2608.00877#A7)\) rather than a lower hallucination rate; we therefore report the blinded reduction in the main text\.
Table I\.1:Source\-blinded claim\-level hallucination \(contradictedshare, 3,000 descriptions/model\), judge reference = fMoW category only\. Reduction \(%\) is*none*→\\rightarrowfiltered under the blinded / standard judge\.Similar Articles
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
Presents an LLM-driven framework for retrieving remote sensing data from cloud-based geospatial catalogues using natural language queries, with a focus on safety and adversarial robustness. The system integrates three agents for intent interpretation, API call generation, and risk management.
Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding
This paper introduces MGAP, a training-free decoding method that reduces hallucinations in Multimodal Large Language Models by adaptively suppressing only the harmful parts of language priors while preserving the model's semantic manifold. The method outperforms prior baselines on POPE and CHAIR benchmarks.
Trust but Verify: Mitigating Medical Hallucinations via Post-Hoc Adversarial Auditing and Multi-Agent Feedback Loops
This paper proposes a multi-agent 'Trust but Verify' system to reduce medical hallucinations in LLMs. It tests three open-access models on clinical questions about banned drugs and achieves a 53% reduction in hallucination error rate.
GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs
GeoStack introduces a geometric framework to compose independently trained domain experts in Vision-Language Models without catastrophic forgetting, achieving constant-time inference and a 10x reduction in geometric error.
RemoteZero: Geospatial Reasoning with Zero Human Annotations
RemoteZero is a framework that eliminates the need for human-annotated box supervision in geospatial reasoning by leveraging the semantic verification capabilities of multimodal large language models (MLLMs) to enable self-evolving localization from unlabeled remote sensing data.