BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

arXiv cs.CL Papers

Summary

BEAR-Bench introduces a bilingual English-and-Russian benchmark for evaluating multimodal models' reasoning on text-rich professional documents, assessing 16 models and highlighting performance gaps.

arXiv:2608.17895v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:07 AM

# A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
Source: [https://arxiv.org/html/2608.17895](https://arxiv.org/html/2608.17895)
Alexandra KuleshovaAffiliation:Yandex Applied AI InstituteCorrespondence:[bazarovaai\.239@gmail\.com](mailto:email@domain)Daniil VolkovAffiliation:Yandex Applied AI InstituteCorrespondence:[bazarovaai\.239@gmail\.com](mailto:email@domain)Kirill SultanovAlexey ZaytsevAffiliation:Yandex Applied AI InstituteCorrespondence:[bazarovaai\.239@gmail\.com](mailto:email@domain)

###### Abstract

While Multimodal Large Language Models \(MLLMs\) have made significant strides in visual comprehension, their ability to reason about text\-dense, professional documents remains incompletely evaluated\. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings\. They are also largely English\- or Chinese\-centric, leaving other languages and Russian, in particular, substantially underrepresented\. To address these limitations, we introduceBEAR\-Bench\(BilingualEnterprise andAcademicReasoning\), a self\-contained, complex English\-and\-Russian benchmark comprising 1000 human\-annotated questions based on text\-rich business and scientific documents\. We evaluate 16 proprietary and open\-weight MLLMs, including Gemini 3\.1 Pro and Qwen3\.5\-397B, on BEAR\-Bench and observe clear headroom even for the strongest systems\. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR\-Bench but also how reliably those failures can be identified\.

\*\*footnotetext:Equal contribution\.## 1Introduction

Multimodal Large Language Models \(MLLMs\)[20](https://arxiv.org/html/2608.17895#bib.bib1);[27](https://arxiv.org/html/2608.17895#bib.bib2);[1](https://arxiv.org/html/2608.17895#bib.bib3)have revolutionized the machine comprehension of images\. Beyond merely solving optical character recognition \(OCR\) tasks, these models demonstrate the potential to perform complex reasoning over multimodal inputs[35](https://arxiv.org/html/2608.17895#bib.bib7);[12](https://arxiv.org/html/2608.17895#bib.bib8)\. This capability is particularly vital for text\-rich scenarios, which are central to real\-world professional domains that require reading and analyzing dense visual documents, such as scientific papers and financial reports\.

Existing benchmarks provide complementary but incomplete coverage of professional document reasoning\. Document\-oriented datasets such as DocVQA[18](https://arxiv.org/html/2608.17895#bib.bib5)primarily emphasize OCR and information extraction, whereas broad multimodal reasoning benchmarks such as MMMU[35](https://arxiv.org/html/2608.17895#bib.bib7)often require specialized factual knowledge\. OCR\-Reasoning[13](https://arxiv.org/html/2608.17895#bib.bib9)covers diverse text\-rich images, including some professional documents, but does not examine this setting in depth\. A further limitation is linguistic coverage: existing multimodal benchmarks exhibit a strong bias toward English and Chinese, leaving Slavic languages — and Russian in particular — substantially underrepresented\. Russian\-inclusive resources such as MTVQA[25](https://arxiv.org/html/2608.17895#bib.bib10), MWS Vision Bench[19](https://arxiv.org/html/2608.17895#bib.bib11), and MERA\-Multi[8](https://arxiv.org/html/2608.17895#bib.bib12)broaden this coverage, but mix documents with other image types, emphasize OCR and document processing, or focus on a narrow document category\. To our knowledge, no existing benchmark evaluates self\-contained, multi\-step reasoning over Russian\-language professional document pages whose items involve textual and graphical content such as figures, tables, charts, equations, and diagrams\.

To address these limitations, we introduceBEAR\-Bench, a complex benchmark comprising 1000 human\-annotated questions for text\-dense enterprise and scientific documents in English and Russian languages\. The science section tests the ability to interpret academic figures, transcribe mathematical formulas, and analyze data from plots\. The enterprise tasks target data cross\-referencing in financial reports and logical reasoning over business charts\. The questions follow two design principles\. First, they require multiple logical steps rather than direct extraction from a single piece of evidence\. Second, they are fully answerable from the document alone, without external expert knowledge\. Together, these principles focus the evaluation on document\-grounded reasoning rather than factual recall or shallow extraction\. The inclusion of Russian\-language tasks further broadens evaluation beyond predominantly English\- and Chinese\-language resources\.

Deploying MLLMs in professional settings requires knowing not only how often they fail, but also whether those failures can be detected reliably; yet evidence on hallucination detection for reasoning over text\-dense professional documents remains incomplete\. We address this gap by using BEAR\-Bench to compare token\-level uncertainty scores, representation\-based detectors, supervised hidden\-state probes, and MLLM\-as\-a\-judge methods across proprietary and open\-weight models, providing a practical comparison for deployment settings with and without access to model internals\.

The main contributions of this work are the following:

1. 1\.We proposeBEAR\-Bench, a bilingual benchmark for multimodal reasoning in professional scenarios\. It features human\-annotated questions targeting text\-dense business and scientific documents in English and Russian languages\.
2. 2\.We evaluate 16 MLLMs on BEAR\-Bench, including proprietary \(e\.g\., Gemini\-3\.1\-Pro, Claude Opus 4\.6\) and open\-weight \(e\.g\., Qwen3\.5, gemma\-4\) models\. Our results show that even the strongest evaluated systems leave clear headroom on BEAR\-Bench\.
3. 3\.We further use BEAR\-Bench to evaluate a diverse set of existing hallucination detection methods for OCR\-intensive professional\-document reasoning, comparing uncertainty\-, representation\-, and judge\-based methods under two deployment regimes: direct access to internal signals for open\-weight models and proxy\-based detection for proprietary ones\.

Benchmark\#Langs\#QA pairs\#RU reasoning VQAImage scopeOCR chars/imgDocVQA[18](https://arxiv.org/html/2608.17895#bib.bib5)EN5\.2Kn/aIndustry documents1,113\.0ChartQA[17](https://arxiv.org/html/2608.17895#bib.bib4)EN2\.5K\*n/aCharts231\.8CharXiv[31](https://arxiv.org/html/2608.17895#bib.bib27)EN11\.6Kn/aScientific charts165\.1OCRBench v2[10](https://arxiv.org/html/2608.17895#bib.bib6)EN, ZH10Kn/aMixed text\-rich images437\.2OCR\-Reasoning[13](https://arxiv.org/html/2608.17895#bib.bib9)EN1\.1Kn/aEveryday text\-rich scenes514\.3MWS Vision Bench[19](https://arxiv.org/html/2608.17895#bib.bib11)RU1\.3K\*\*400Business/personal documents1,126\.4LabTabVQA[8](https://arxiv.org/html/2608.17895#bib.bib12)RU349349Medical report tables637\.1CC\-OCR V2[33](https://arxiv.org/html/2608.17895#bib.bib28)RU \+ 31 langs2K\*\*\*0Finance/dashboards/blueprints1,149\.6BEAR\-BenchEN, RU1K618Science/business documents2,740\.6

Table 1:Comparison of BEAR\-Bench to existing multimodal reasoning benchmarks\.\#QA pairsrefers to the evaluation split; DocVQA and ChartQA additionally provide training data \(50K and 32\.7K items in total, respectively\), while the remaining benchmarks are evaluation\-only\. For CC\-OCR V2, the reported count is the document QA track \(2K of 7\.1K items\)\.\#RU reasoning VQA: number of evaluation questions in Russian that require answering from the image, excluding OCR, parsing, grounding, and key\-information extraction; n/a = not applicable \(no Russian split\)\.Image scopesummarizes the principal visual sources represented in each benchmark\.OCR chars/img: average number of non\-whitespace OCR characters per image \(n=200n\{=\}200randomly sampled images per benchmark\)\.\*Evenly split between human\-written and machine\-generated questions\. \*\*Publicly available\.\.
## 2Related Work

##### Multimodal benchmarks for professional documents\.

While modern LLMs achieve strong results on many established multimodal benchmarks[36](https://arxiv.org/html/2608.17895#bib.bib31);[41](https://arxiv.org/html/2608.17895#bib.bib32), their ability to operate with visually rich professional documents — which requires analyzing both textual and graphical content of an image —\- remains underexplored\. Text\-dense benchmarks oriented at professional documents are mostly OCR\-based and measure extractive skills rather than cross\-referencing of text and visuals[33](https://arxiv.org/html/2608.17895#bib.bib28);[18](https://arxiv.org/html/2608.17895#bib.bib5), or emphasize long\-page settings that conflate multimodal reasoning with long\-context handling[7](https://arxiv.org/html/2608.17895#bib.bib30);[26](https://arxiv.org/html/2608.17895#bib.bib29)\. Multimodal reasoning\-focused benchmarks, conversely, offer little signal specifically for professional documents: OCR\-Reasoning[13](https://arxiv.org/html/2608.17895#bib.bib9)and OCRBench v2[10](https://arxiv.org/html/2608.17895#bib.bib6)are deliberately broad — valuable for general\-purpose evaluation, but treating professional documents as one setting among many — while CharXiv[31](https://arxiv.org/html/2608.17895#bib.bib27)and ChartQA[17](https://arxiv.org/html/2608.17895#bib.bib4);[16](https://arxiv.org/html/2608.17895#bib.bib40)restrict evaluation to charts and thus do not test text–visual cross\-referencing\. A further issue is dependence on external knowledge: MMMU[35](https://arxiv.org/html/2608.17895#bib.bib7)and EMMA[12](https://arxiv.org/html/2608.17895#bib.bib8)pose multidisciplinary problems that presuppose domain expertise, so scores conflate multimodal reasoning failures with factual gaps; part of OCR\-Reasoning shares this confound\.

Language coverage is also narrow: most benchmarks covering multimodal reasoning in text\-dense scenarios are available only in English or Chinese, while many other languages, including Russian, remain underrepresented\. Russian\-inclusive benchmarks provide valuable but partial coverage\. MTVQA[25](https://arxiv.org/html/2608.17895#bib.bib10)and TIU\-Bench[39](https://arxiv.org/html/2608.17895#bib.bib33)include documents alongside natural scenes but offer very limited coverage of Russian\-language professional documents; TIU\-Bench, for example, contains only 10 Russian document samples in total\. MWS Vision Bench[19](https://arxiv.org/html/2608.17895#bib.bib11)provides 400 Russian reasoning\-VQA items, but its images span business scans, personal handwriting, receipts, and form\-style pages \(Figure[5](https://arxiv.org/html/2608.17895#A2.F5)\), rather than being selected specifically for reasoning over text\-dense professional documents\.LabTabVQA— the only subset of MERA\-Multi[8](https://arxiv.org/html/2608.17895#bib.bib12)built on professional documents rather than natural images or exam\-style problems — is restricted to tables from medical laboratory reports\.

##### Error detection for text\-dense multimodal inputs\.

Recent work proposes a range of hallucination detectors for MLLMs[6](https://arxiv.org/html/2608.17895#bib.bib41)\. Tool\-augmented methods rely on auxiliary models[5](https://arxiv.org/html/2608.17895#bib.bib21);[34](https://arxiv.org/html/2608.17895#bib.bib34);[24](https://arxiv.org/html/2608.17895#bib.bib38), while multi\-query methods require repeated generation or verification calls[32](https://arxiv.org/html/2608.17895#bib.bib37);[40](https://arxiv.org/html/2608.17895#bib.bib36), increasing deployment cost\. Of particular interest are lightweight white\-box methods, which detect errors using token uncertainty or internal model states collected during a single forward pass[30](https://arxiv.org/html/2608.17895#bib.bib35);[15](https://arxiv.org/html/2608.17895#bib.bib22);[14](https://arxiv.org/html/2608.17895#bib.bib39);[38](https://arxiv.org/html/2608.17895#bib.bib23)\. These methods add relatively little inference overhead when model internals are available, yet, to our knowledge, they have not been systematically compared on visually rich professional document images requiring multi\-step reasoning across dense textual and graphical evidence\.

##### Summary\.

Existing work lacks a Russian\-inclusive benchmark for self\-contained, multi\-step reasoning over text\-dense professional documents\. BEAR\-Bench is designed to fill this gap using scientific and business documents; Table[1](https://arxiv.org/html/2608.17895#S1.T1)summarizes the comparison\.

Furthermore, BEAR\-Bench enables a systematic comparison of hallucination detectors on such text\-dense professional document images requiring multi\-step reasoning, a setting not covered by prior evaluations\.

## 3BEAR\-Bench

### 3\.1Domain Scope and Taxonomy

BEAR\-Bench spans two primary domains —BusinessandScience— each subdivided into thematically coherent sub\-categories\.

##### Business Domain\.

The business subset covers three document types:financial reports\(SEC Forms 10\-K and 10\-Q\) requiring tabular reasoning and year\-over\-year calculations;investor presentations\(Form 8\-K exhibits\) combining charts, KPI tiles, and infographic maps; andflowcharts and organisational diagramsdepicting corporate ownership structures and process pipelines\.

##### Science Domain\.

The science subset covers three categories:mathematical and physical formulaefrom physics and mathematics preprints targeting symbol\-level recognition;scientific figures and plots\(line plots, scatter diagrams, heatmaps\) requiring axis and legend interpretation; andacademic layoutswith multi\-column pages testing reading\-order resolution and cross\-referential reasoning\.

The two domains are strictly disjoint: no source document appears in both subsets\.

### 3\.2Data Collection and Annotation Pipeline

#### 3\.2\.1Source Collection

##### Business Domain\.

Business documents were retrieved via targeted Google Search queries directed at publicly accessible, license\-safe sources\. English\-language documents were obtained from the U\.S\. Securities and Exchange Commission \(SEC\) EDGAR system — annual reports \(Form 10\-K\), quarterly reports \(Form 10\-Q\), and investor presentations filed as Form 8\-K exhibits — all of which constitute public records under U\.S\. federal law\. Supplementary English documents were drawn from official government portals \(\*\.gov,\*\.gov\.uk\) and intergovernmental repositories \(\*\.int\)\. Russian\-language documents were sourced from the state corporate\-disclosure platformse\-disclosure\.ruandmoex\.com, as well as from federal government domains \(\*\.gov\.ru\)\. An automated scraper retrieved candidate PDFs; each document underwent a programmatic license\-verification step examining the first and last five pages for SEC registration markers or open\-license declarations \(*“Creative Commons”*,*“CC BY”*,*“public domain”*\)\. Documents failing this check were discarded prior to further processing\.

##### Science Domain\.

English\-language papers were downloaded from arXiv via its official Python API, sampling four STEM categories:quant\-ph,cs\.AI,eess\.SP, andmath\.GM\(up to 100 papers per category\)\. Russian\-language articles were collected from CyberLeninka \(cyberleninka\.ru\) using an asynchronous Playwright\-based crawler across four subject areas: Computer Science, Mathematics, Physics, and Engineering \(up to 100 articles per category\)\.

#### 3\.2\.2Filtering and Preprocessing

Raw PDFs were rendered page\-by\-page into PNG images and processed through a two\-stage filtering pipeline\.

##### Stage 1 — Visual Content Classification\.

We obtained silver labels for a stratified sample of 3,000 images using Gemini 2\.5 Pro with a structured multi\-label prompt, producing six Boolean fields:contains\_diagrams,contains\_tables,contains\_equations,contains\_code,contains\_figures, andcontains\_handwriting\. These labels trained a lightweight classifier: SigLIP embeddings\([37](https://arxiv.org/html/2608.17895#bib.bib13)\)were L2\-normalised and passed to aMultiOutputClassifierof logistic\-regression models with balanced class weights, one per label\. After validation on a held\-out 20% split, the classifier was applied to the full corpus of∼\\sim66,000 page images, retaining only pages with at least one of \{contains\_diagrams,contains\_equations,contains\_code\} predicted positive\.

##### Stage 2 — Textual Density Filtering\.

Among content\-positive pages, we retained only those at or above the 67th percentile of OCR character count within their respective language group\. From the surviving candidates, up to 1,500 images per language were drawn via stratified random sampling \(seed=42\{=\}\\,42\), yielding the final pool submitted to human annotators\.

![Refer to caption](https://arxiv.org/html/2608.17895v1/images/figure1_benchmark_composition.png)Figure 1:Representative items from BEAR\-Bench\.Each card displays the source document image \(left\) alongside the human\-authored multi\-hop question and ground\-truth answer \(right\)\.

#### 3\.2\.3Human Annotation

##### Annotator pool\.

Annotation was conducted by 13 domain experts, each holding at minimum a Bachelor’s degree in a technical discipline\. Every annotator processed its own subset of images, authoring exactly one question–answer pair per image\.

##### Annotation task\.

For each image, annotators were required to: \(i\) re\-verify the visual\-content labels produced by the automatic classifier, correcting any erroneous predictions; and \(ii\) compose a multi\-step, multi\-hop question with a detailed ground\-truth answer\. Questions were required to elicit compositional reasoning — aggregating values across table rows, interpreting plotted trends in the context of equations, or tracing paths through flowcharts — rather than straightforward single\-step extraction\. The questions were additionally assigned areasoning depth scorerepresenting the total number of reasoning and computational steps required to arrive at the correct answer\. The text of the instruction for the annotators is reported in Figure[9](https://arxiv.org/html/2608.17895#A7.F9)\.

##### Evaluation judge\.

Model responses are scored by GPT\-4o used as an LLM\-as\-a\-judge\. The judge assesses semantic equivalence between the model answer and the ground truth, permitting surface\-level paraphrase while penalising under\-specific responses\. It returns a binary verdictv∈\{true,false\}v\\in\\\{\\texttt\{true\},\\,\\texttt\{false\}\\\}of whether the model answer is correct with a brief explanation in structured XML tags, enabling fully reproducible programmatic evaluation\. The exact prompt is provided in Figure[11](https://arxiv.org/html/2608.17895#A8.F11)\.

##### Judge reliability\.

To validate the reliability of the LLM\-as\-a\-judge protocol, we constructed a stratified audit sample of 200 judge verdicts, drawn uniformly across languages and domains \(100 English and 100 Russian items; 100 Business and 100 Science items, with 50 items per language–domain cell\)\. Human annotators independently reviewed each model response, ground\-truth answer, and judge verdict, recording agreement or disagreement\. The judge achieved an overall human\-agreement rate of 99\.0% \(198/200\), with only two disagreements in the entire sample\. Agreement remained consistently high across languages \(99% for both English and Russian\) and domains \(99% for both Business and Science\), as well as at the finer\-grained language\-domain level \(98\-100% across all four cells\), with 95% Wilson confidence intervals overlapping the overall estimate throughout\. On this stratified sample, the LLM\-as\-a\-judge protocol agrees closely with human evaluation\.

##### Quality control\.

Eight state\-of\-the\-art proprietary VLMs were queried on every item: Gemini 2\.5 Pro/Flash, Gemini 3\.1 Pro/Flash, Qwen 3\.6 Plus, Qwen3\.5 397B, Claude Sonnet 4\.6, and Claude Opus 4\.6\. Items where three or more models returned identical responses — normalised for punctuation and case — and the LLM judge assignedfalseto all answers were flagged\. A manual audit confirmed that 99% of flagged items had erroneous or ambiguous ground truth; all were excluded from the final benchmark\. A random sample of retained items was quality\-assessed along four dimensions: GT quality \(84\.4%\), judge verdict quality \(90\.6%\), question quality \(90\.9%\), and image quality \(97\.0%\)\.

##### Final dataset composition\.

After quality\-control filtering, BEAR\-Bench comprises1,000 document imagespaired with1,000 human\-authored QA instancesacross four domain–language groups\. Representative samples from BEAR\-Bench are shown in Figure[1](https://arxiv.org/html/2608.17895#S3.F1)\.

## 4Dataset Statistics and Analysis

Figure[2](https://arxiv.org/html/2608.17895#S4.F2)summarises the composition of BEAR\-Bench across three dimensions: domain–language balance, question complexity, and visual content\-type prevalence\.

##### Domain and language balance\.

The Russian business cell is the largest subset \(352 items, 35\.2%\), reflecting the higher volume of publicly accessible Russian\-language corporate disclosure documents, while the English business cell is the smallest \(180 items, 18\.0%\)\. The science cells are more evenly distributed \(266 and 202 items for Russian and English, respectively\)\.

##### Reasoning depth\.

Reasoning\-depth annotations are available for 940 of the 1,000 items in BEAR\-Bench\. Each annotation estimates the intended number of steps required to derive the correct answer from the document image\. Because such step counts depend on how annotators decompose a task, we treat them as coarse descriptive metadata rather than an objective difficulty score\. The estimates span 2–10\+ steps and peak at 4–5 steps with a moderate positive skew, indicating that the benchmark construction targeted multi\-step inference rather than simple extraction\.

##### Visual content types\.

Figures and diagrams are the most prevalent content types across all subsets, consistent with the heavy use of infographics in both corporate reports and scientific papers\. Equations appear almost exclusively in the science subsets, reflecting the mathematical nature of the arXiv and CyberLeninka source material, while code fragments are comparatively rare overall\.

![Refer to caption](https://arxiv.org/html/2608.17895v1/images/figure2_dataset_statistics_new.png)Figure 2:BEAR\-Bench dataset statistics\.\(a\)Item distribution across the four domain×\\timeslanguage cells; the central numeral indicates the total count\.\(b\)Distribution of complexity scores over all 1,000 items, where each score reflects the total number of reasoning and computational steps required to solve the corresponding question\.\(c\)Prevalence of visual content types, diagrams, equations, code, and figures, disaggregated by subset; bars show absolute counts\. Items may carry multiple content\-type labels simultaneously\.

## 5Experiments

### 5\.1Experimental Setup

##### Evaluated models\.

We evaluate a diverse set of MLLMs on BEAR\-Bench, covering both open\-weight models \(Qwen3\.5\-0\.8B/4B/9B/27B[22](https://arxiv.org/html/2608.17895#bib.bib17), Qwen3\-VL\-2B/8B\-Instruct[29](https://arxiv.org/html/2608.17895#bib.bib18), Qwen3\-VL\-4B\-Thinking[29](https://arxiv.org/html/2608.17895#bib.bib18), and Gemma\-4\-31B\-it[28](https://arxiv.org/html/2608.17895#bib.bib19)\) and proprietary systems \(Qwen3\.5\-397B\-A17B[22](https://arxiv.org/html/2608.17895#bib.bib17), Qwen 3\.6 Plus[23](https://arxiv.org/html/2608.17895#bib.bib20), Gemini 2\.5/3\.1 Pro/Flash[9](https://arxiv.org/html/2608.17895#bib.bib24), and Claude Sonnet/Opus 4\.6[3](https://arxiv.org/html/2608.17895#bib.bib25);[2](https://arxiv.org/html/2608.17895#bib.bib26)\)\. The open\-weight models were run locally on an internal GPU server equipped with NVIDIA H100 and NVIDIA L40 accelerators; the proprietary ones were queried via the OpenRouter API\.

##### Inference protocol\.

All models were evaluated zero\-shot under a fixed protocol: each instance received only the image and the raw question, with no few\-shot examples or prompt engineering, and a uniform decoding temperature of 0\.6\.

### 5\.2Results

Table 2:BEAR\-Bench leaderboard\. Results are reported as accuracy \(%\) across all evaluation subsets\.Bolddenotes the best result in each column\.##### Main results\.

Table[2](https://arxiv.org/html/2608.17895#S5.T2)shows remaining headroom on BEAR\-Bench\. Qwen3\.5\-397B\-A17B and Gemini 3\.1 Pro achieve the highest overall accuracy \(75\.4% and 75\.1%, respectively\), trading the lead across subsets — Qwen3\.5\-397B\-A17B is stronger on Science and Equations, while Gemini 3\.1 Pro edges ahead on Business and English items — indicating that no single system dominates across all domains\. All models exhibit a marked drop from English to Russian \(e\.g\., 83\.0%→\\rightarrow70\.2% for Gemini 3\.1 Pro and 81\.9%→\\rightarrow71\.4% for Qwen3\.5\-397B\-A17B\), confirming that the linguistic gap identified in prior benchmarks persists even for frontier proprietary models\. Among the evaluated systems, proprietary models generally achieve higher accuracy than their open\-weight counterparts\. Within the Qwen3\.5 and Qwen3\-VL\-Instruct families, larger models tend to perform better in both languages \(Figure[3](https://arxiv.org/html/2608.17895#S5.F3)\), although accuracy remains substantially lower on Russian items\. Overall, current models remain limited in visually grounded reasoning over text\-dense professional documents, with performance shaped jointly by model scale, language, and document type\.

Figure 3:Accuracy on BEAR\-Bench versus model size for the Qwen3\.5 and Qwen3\-VL\-Instruct families \(parameters on a log axis\), reported separately for English \(solid\) and Russian \(dashed\) items\. The English\-over\-Russian gap persists across scales\.
##### Effect of Chain\-of\-Thought prompting\.

Table 3:Overall accuracy \(%\) with and without an explicit chain\-of\-thought prompt\.Δ\\Deltais CoT minus the default \(no\-CoT\) protocol used in Table[2](https://arxiv.org/html/2608.17895#S5.T2)\.We tested whether explicit chain\-of\-thought \(CoT\) prompting improves accuracy on BEAR\-Bench\. A subset of open\-weight models was re\-evaluated with a prompt that asks the model to extract information from the image and solve the task step by step before giving a final answer \(Table[3](https://arxiv.org/html/2608.17895#S5.T3); the full prompt is given in Appendix[F](https://arxiv.org/html/2608.17895#A6)\)\. For reasoning\-oriented models — Qwen3\.5\-9B, Qwen3\.5\-4B, and Qwen3\-VL\-4B\-Thinking — accuracy is essentially unchanged or slightly lower, consistent with these models already performing intermediate reasoning under the default protocol\. By contrast, the instruct\-tuned Qwen3\-VL\-8B\-Instruct improves by 13\.7 percentage points when steered to externalize multi\-step reasoning\.

##### Effect of image resolution\.

Table 4:Overall accuracy of Qwen3\.5\-9B on BEAR\-Bench at different image downsampling factors\. A factor ofccmeans that both the width and height are divided bycc\.To measure the sensitivity of document reasoning to image resolution, we evaluated Qwen3\.5\-9B after resizing each input image to1/c1/cof its original width and height, whereccis the downsampling factor in Table[4](https://arxiv.org/html/2608.17895#S5.T4)\. All other inference settings were kept unchanged\. Accuracy decreases from 49\.3% at the original resolution to 45\.0% atc=1\.5c=1\.5and 33\.0% atc=2c=2, before falling sharply to 12\.6% atc=3c=3and 4\.7% atc=4c=4\. This pronounced degradation suggests that preserving fine\-grained visual detail is critical for reasoning over text\-dense professional documents\.

### 5\.3Error Analysis

To characterize common failure modes on BEAR\-Bench, we conducted an exploratory error analysis combining inductive taxonomy construction with manual annotation of model responses\.

#### 5\.3\.1Taxonomy construction

To derive a failure taxonomy grounded in actual model behavior, we first collected an open\-coding pilot of 150 incorrect responses produced by Gemini 3\.1 Pro on BEAR\-Bench, stratified across domains and languages\. Four members of our team manually inspected each failure case and wrote detailed, free\-form natural\-language comments explaining the underlying cause of the error, rather than assigning predefined labels\. This yielded a corpus of 150 rich failure descriptions covering a broad range of perceptual and reasoning breakdowns\. We then prompted GPT\-5\.5\-Pro, used as an advanced classification assistant, to cluster these free\-form comments into thematically coherent groups based on their underlying error mechanism\. The resulting clusters were manually reviewed and refined by the authors into five final categories: spatial misgrounding \(C1\), counting/aggregation \(C2\), OCR/visual\-attribute \(C3\), chart\-value extraction \(C4\), and semantic/reasoning \(C5\)\. Detailed descriptions of each error type are provided in Appendix[A](https://arxiv.org/html/2608.17895#A1)\.

#### 5\.3\.2Error statistics

![Refer to caption](https://arxiv.org/html/2608.17895v1/images/error_type_distribution.png)Figure 4:Error\-type distribution in a manually annotated subsample of Gemini 3\.1 Pro responses \(n=62n=62\)\. Samples may receive multiple labels\.We applied the taxonomy to a subsample of Gemini 3\.1 Pro incorrect responses, selected to cover diverse document types and complexity levels\. Figure[4](https://arxiv.org/html/2608.17895#S5.F4)shows the resulting distribution\. Visual\-perceptual errors dominate: spatial misgrounding \(C1\) and OCR/visual\-attribute errors \(C3\) together account for the majority of failures\. Our analysis suggests that the most common errors on BEAR\-Bench are perceptual, involving misread text or visual attributes and mislocalized evidence\.

### 5\.4Detecting Incorrect Answers

We next use responses generated on BEAR\-Bench to compare existing hallucination detection methods for text\-dense professional document reasoning\. Following the benchmark’s binary answer evaluation, we assess whether each method can distinguish correct from incorrect responses\.

##### Setup\.

We study two open\-weight Qwen3\.5 models using their native internal signals and eight proprietary models using proxy hidden states extracted from Qwen3\-VL\-8B[4](https://arxiv.org/html/2608.17895#bib.bib16), conditioned on each image, question, and proprietary\-model response\. We evaluate six uncertainty scores—max/mean token probability, log\-likelihood, max/mean entropy, and perplexity—plus ContextualLens[21](https://arxiv.org/html/2608.17895#bib.bib14), the supervised hidden\-state probe SUQ[15](https://arxiv.org/html/2608.17895#bib.bib22), and an MLLM\-as\-a\-judge baseline \(Qwen3\-VL\-8B\)[11](https://arxiv.org/html/2608.17895#bib.bib15)\. We report balanced accuracy \(BalAcc\), AUROC and AUC\-PR metrics\.

##### Results\.

The performance of the evaluated methods is shown in Tables[6](https://arxiv.org/html/2608.17895#A5.T6)\-[8](https://arxiv.org/html/2608.17895#A5.T8)\. No detector wins everywhere; performance depends on the access regime and on response length \(Table[5](https://arxiv.org/html/2608.17895#A4.T5)\)\. SUQ achieves the highest BalAcc on all eight proprietary models \(BalAcc0\.670\.67–0\.740\.74; median response length<150<150words\), but performs less well on the two open\-weight models \(BalAcc0\.600\.60–0\.670\.67; median length\>1,000\>1\{,\}000words\), suggesting that a last\-token embedding provides limited signal for errors occurring earlier in long reasoning chains\. Among proxy uncertainty scores, max token probability is strongest for the four models answering in two to three words \(0\.590\.59–0\.660\.66\) but at chance for the four with longer answers \(0\.490\.49–0\.510\.51\), while mean\-based score shows the reverse \(0\.480\.48–0\.560\.56versus0\.650\.65–0\.660\.66\)\. The judge improves with length, from0\.570\.57–0\.620\.62on terse answers to0\.800\.80–0\.810\.81on the verbose open\-weight models—the highest BalAcc we observe—as a detailed derivation can be checked step by step\.

## 6Conclusion

We introduced BEAR\-Bench, a bilingual benchmark of 1,000 human\-authored questions for context\-grounded, multi\-step reasoning over text\-dense business and scientific documents\. Benchmark items include dense textual and graphical page content such as figures, tables, charts, equations, and diagrams\. Across 16 proprietary and open\-weight MLLMs, the highest overall accuracy on BEAR\-Bench is 75\.4%, and every evaluated model achieves lower accuracy on Russian items\. Our error analysis suggests that common failures involve spatial grounding, OCR, and visual\-attribute perception\.

We also used the resulting model responses to compare existing hallucination detection methods\. Performance varies across target models: supervised probes perform best for proprietary outputs, while an MLLM judge achieves the highest balanced accuracy on the verbose responses from open\-weight outputs\. The best balanced accuracy is 0\.74 for proprietary outputs and 0\.81 for open\-weight models, showing that incorrect responses are not always identified reliably\. BEAR\-Bench therefore provides a common setting for tracking progress in both professional document reasoning and error detection\.

## Limitations

BEAR\-Bench is deliberately narrow in several aspects\. All questions are scoped to a single rendered page image; the benchmark therefore does not evaluate multi\-page or cross\-document reasoning\. Coverage is limited to English and Russian enterprise and academic documents drawn from public disclosure and preprint sources, so findings may not transfer to other languages, domains, or private enterprise corpora\. With 1,000 items and uneven language–domain cell sizes, subset estimates carry more variance than the overall score\. Although we validate the LLM\-as\-a\-judge protocol against humans, scoring still depends on an external model, and reasoning\-depth labels remain coarse annotator estimates rather than objective difficulty\. Finally, our hallucination\-detection study compares existing methods under two access regimes and finds that detector quality varies with response length; we do not propose a new detector, and even the best balanced accuracies leave substantial room for improvement\.

## 7Acknowledgements

The work was supported by the grant for research centers in the field of AI provided by the Ministry of Economic Development of the Russian Federation in accordance with the agreement 000000C313925P4F0002 and the agreement №139\-10\-2025\-033\.

## References

- Anthropic \(2024\)AnthropicThe Claude 3 model family: Opus, Sonnet, Haiku\.Model CardAnthropic\.Note:Accessed: 2024\-09\-07External Links:[Link](https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf)Cited by:[§1](https://arxiv.org/html/2608.17895#S1.p1.1)\.
- Anthropic \(2026a\)AnthropicClaude opus \(version 4\.6\)\.External Links:[Link](https://www.anthropic.com/news/claude-opus-4-6)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Anthropic \(2026b\)AnthropicClaude sonnet \(version 4\.6\)\.External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-6)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Baiet al\.\(2025\)S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge, W\. Ge, Z\. Guo, Q\. Huang, J\. Huang, F\. Huang, B\. Hui, S\. Jiang, Z\. Li, M\. Li, M\. Li, K\. Li, Z\. Lin, J\. Lin, X\. Liu, J\. Liu, C\. Liu, Y\. Liu, D\. Liu, S\. Liu, D\. Lu, R\. Luo, C\. Lv, R\. Men, L\. Meng, X\. Ren, X\. Ren, S\. Song, Y\. Sun, J\. Tang, J\. Tu, J\. Wan, P\. Wang, P\. Wang, Q\. Wang, Y\. Wang, T\. Xie, Y\. Xu, H\. Xu, J\. Xu, Z\. Yang, M\. Yang, J\. Yang, A\. Yang, B\. Yu, F\. Zhang, H\. Zhang, X\. Zhang, B\. Zheng, H\. Zhong, J\. Zhou, F\. Zhou, J\. Zhou, Y\. Zhu, and K\. ZhuQwen3\-vl technical report\.External Links:2511\.21631,[Link](https://arxiv.org/abs/2511.21631)Cited by:[§5\.4](https://arxiv.org/html/2608.17895#S5.SS4.SSS0.Px1.p1.1)\.
- Chenet al\.\(2024\)X\. Chen, C\. Wang, Y\. Xue, N\. Zhang, X\. Yang, Q\. Li, Y\. Shen, L\. Liang, J\. Gu, and H\. ChenUnified hallucination detection for multimodal large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3235–3252\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2026a\)Z\. Chen, Y\. Min, J\. Zhang, B\. Yan, J\. Wang, X\. Wang, and S\. ShanA survey of multimodal hallucination evaluation and detection\.International Journal of Computer Vision134\(3\),pp\. 131\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2026b\)Z\. Chen, Y\. Zhao, C\. Wang, R\. R\. Han, M\. Patwardhan, and A\. CohanSciMDR: advancing scientific multimodal document reasoning\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 44718–44742\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Chervyakovet al\.\(2026\)A\. Chervyakov, U\. Isaeva, A\. Emelyanov, A\. Safin, M\. Tikhonova, A\. Kharitonov, Y\. Lyakh, P\. Surovtsev, D\. Shevelev, V\. Saburov, V\. Konovalov, E\. Rykov, I\. Sviridov, A\. Miftakhova, I\. Alimova, A\. Panchenko, A\. Kapitanov, and A\. FenogenovaMultimodal evaluation of Russian\-language architectures\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 2114–2161\.External Links:[Link](https://aclanthology.org/2026.eacl-long.94/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.94),ISBN 979\-8\-89176\-380\-7Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.8.1),[§1](https://arxiv.org/html/2608.17895#S1.p2.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p2.1)\.
- DeepMind \(2026\)G\. DeepMindGemini model cards\.Note:Accessed: 2026\-08\-03External Links:[Link](https://deepmind.google/models/model-cards/)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Fuet al\.\(2025\)L\. Fu, Z\. Kuang, J\. Song, M\. Huang, B\. Yang, Y\. Li, L\. Zhu, Q\. Luo, X\. Wang, H\. Lu, Z\. Li, G\. Tang, B\. Shan, C\. Lin, Q\. Liu, B\. Wu, H\. Feng, H\. Liu, C\. Huang, J\. Tang, W\. Chen, L\. Jin, Y\. Liu, and X\. BaiOCRBench v2: an improved benchmark for evaluating large multimodal models on visual text localization and reasoning\.External Links:2501\.00321,[Link](https://arxiv.org/abs/2501.00321)Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.5.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Guet al\.\(2025\)J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu, S\. Wang, K\. Zhang, Y\. Wang, W\. Gao, L\. Ni, and J\. GuoA survey on llm\-as\-a\-judge\.External Links:2411\.15594,[Link](https://arxiv.org/abs/2411.15594)Cited by:[§5\.4](https://arxiv.org/html/2608.17895#S5.SS4.SSS0.Px1.p1.1)\.
- Haoet al\.\(2025\)Y\. Hao, J\. Gu, H\. W\. Wang, L\. Li, Z\. Yang, L\. Wang, and Y\. ChengCan mllms reason in multimodality? emma: an enhanced multimodal reasoning benchmark\.External Links:2501\.05444,[Link](https://arxiv.org/abs/2501.05444)Cited by:[§1](https://arxiv.org/html/2608.17895#S1.p1.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Huanget al\.\(2026\)M\. Huang, Y\. Shi, D\. Peng, S\. Lai, Z\. Xie, and L\. JinOCR\-reasoning benchmark: unveiling the true capabilities of mllms in complex text\-rich image reasoning\.External Links:2505\.17163,[Link](https://arxiv.org/abs/2505.17163)Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.6.1),[§1](https://arxiv.org/html/2608.17895#S1.p2.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Jianget al\.\(2025\)Z\. Jiang, J\. Chen, B\. Zhu, T\. Luo, Y\. Shen, and X\. YangDevils in middle layers of large vision\-language models: interpreting, detecting and mitigating object hallucinations via attention lens\.In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 25004–25014\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2024\)Q\. Li, J\. Geng, C\. Lyu, D\. Zhu, M\. Panov, and F\. KarrayReference\-free hallucination detection for large vision\-language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 4542–4551\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1),[§5\.4](https://arxiv.org/html/2608.17895#S5.SS4.SSS0.Px1.p1.1)\.
- Masryet al\.\(2025\)A\. Masry, M\. S\. Islam, M\. Ahmed, A\. Bajaj, F\. Kabir, A\. Kartha, M\. T\. R\. Laskar, M\. Rahman, S\. Rahman, M\. Shahmohammadi,et al\.Chartqapro: a more diverse and challenging benchmark for chart question answering\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 19123–19151\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Masryet al\.\(2022\)A\. Masry, D\. X\. Long, J\. Q\. Tan, S\. Joty, and E\. HoqueChartQA: a benchmark for question answering about charts with visual and logical reasoning\.InFindings of the Association for Computational Linguistics: ACL 2022,S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 2263–2279\.External Links:[Link](https://aclanthology.org/2022.findings-acl.177/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.177)Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.3.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Mathewet al\.\(2021\)M\. Mathew, D\. Karatzas, and C\. V\. JawaharDocVQA: a dataset for vqa on document images\.In2021 IEEE Winter Conference on Applications of Computer Vision \(WACV\),Vol\.,pp\. 2199–2208\.External Links:[Document](https://dx.doi.org/10.1109/WACV48630.2021.00225)Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.2.1),[§1](https://arxiv.org/html/2608.17895#S1.p2.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- MWS AI \(2025\)MWS AIMWS vision bench: a Russian\-language document benchmark for multimodal large language models\.Note:[https://github\.com/mts\-ai/MWS\-Vision\-Bench](https://github.com/mts-ai/MWS-Vision-Bench)Accessed: 2026\-07\-31Cited by:[Figure 5](https://arxiv.org/html/2608.17895#A2.F5),[Appendix B](https://arxiv.org/html/2608.17895#A2.p1.1),[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.7.1),[§1](https://arxiv.org/html/2608.17895#S1.p2.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p2.1)\.
- OpenAIet al\.\(2024\)OpenAI, J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat, R\. Avila, I\. Babuschkin, S\. Balaji, V\. Balcom, P\. Baltescu, H\. Bao, M\. Bavarian, J\. Belgum, I\. Bello, J\. Berdine, G\. Bernadett\-Shapiro, C\. Berner, L\. Bogdonoff, O\. Boiko, M\. Boyd, A\. Brakman, G\. Brockman, T\. Brooks, M\. Brundage, K\. Button, T\. Cai, R\. Campbell, A\. Cann, B\. Carey, C\. Carlson, R\. Carmichael, B\. Chan, C\. Chang, F\. Chantzis, D\. Chen, S\. Chen, R\. Chen, J\. Chen, M\. Chen, B\. Chess, C\. Cho, C\. Chu, H\. W\. Chung, D\. Cummings, J\. Currier, Y\. Dai, C\. Decareaux, T\. Degry, N\. Deutsch, D\. Deville, A\. Dhar, D\. Dohan, S\. Dowling, S\. Dunning, A\. Ecoffet, A\. Eleti, T\. Eloundou, D\. Farhi, L\. Fedus, N\. Felix, S\. P\. Fishman, J\. Forte, I\. Fulford, L\. Gao, E\. Georges, C\. Gibson, V\. Goel, T\. Gogineni, G\. Goh, R\. Gontijo\-Lopes, J\. Gordon, M\. Grafstein, S\. Gray, R\. Greene, J\. Gross, S\. S\. Gu, Y\. Guo, C\. Hallacy, J\. Han, J\. Harris, Y\. He, M\. Heaton, J\. Heidecke, C\. Hesse, A\. Hickey, W\. Hickey, P\. Hoeschele, B\. Houghton, K\. Hsu, S\. Hu, X\. Hu, J\. Huizinga, S\. Jain, S\. Jain, J\. Jang, A\. Jiang, R\. Jiang, H\. Jin, D\. Jin, S\. Jomoto, B\. Jonn, H\. Jun, T\. Kaftan, Ł\. Kaiser, A\. Kamali, I\. Kanitscheider, N\. S\. Keskar, T\. Khan, L\. Kilpatrick, J\. W\. Kim, C\. Kim, Y\. Kim, J\. H\. Kirchner, J\. Kiros, M\. Knight, D\. Kokotajlo, Ł\. Kondraciuk, A\. Kondrich, A\. Konstantinidis, K\. Kosic, G\. Krueger, V\. Kuo, M\. Lampe, I\. Lan, T\. Lee, J\. Leike, J\. Leung, D\. Levy, C\. M\. Li, R\. Lim, M\. Lin, S\. Lin, M\. Litwin, T\. Lopez, R\. Lowe, P\. Lue, A\. Makanju, K\. Malfacini, S\. Manning, T\. Markov, Y\. Markovski, B\. Martin, K\. Mayer, A\. Mayne, B\. McGrew, S\. M\. McKinney, C\. McLeavey, P\. McMillan, J\. McNeil, D\. Medina, A\. Mehta, J\. Menick, L\. Metz, A\. Mishchenko, P\. Mishkin, V\. Monaco, E\. Morikawa, D\. Mossing, T\. Mu, M\. Murati, O\. Murk, D\. Mély, A\. Nair, R\. Nakano, R\. Nayak, A\. Neelakantan, R\. Ngo, H\. Noh, L\. Ouyang, C\. O’Keefe, J\. Pachocki, A\. Paino, J\. Palermo, A\. Pantuliano, G\. Parascandolo, J\. Parish, E\. Parparita, A\. Passos, M\. Pavlov, A\. Peng, A\. Perelman, F\. de Avila Belbute Peres, M\. Petrov, H\. P\. de Oliveira Pinto, Michael, Pokorny, M\. Pokrass, V\. H\. Pong, T\. Powell, A\. Power, B\. Power, E\. Proehl, R\. Puri, A\. Radford, J\. Rae, A\. Ramesh, C\. Raymond, F\. Real, K\. Rimbach, C\. Ross, B\. Rotsted, H\. Roussez, N\. Ryder, M\. Saltarelli, T\. Sanders, S\. Santurkar, G\. Sastry, H\. Schmidt, D\. Schnurr, J\. Schulman, D\. Selsam, K\. Sheppard, T\. Sherbakov, J\. Shieh, S\. Shoker, P\. Shyam, S\. Sidor, E\. Sigler, M\. Simens, J\. Sitkin, K\. Slama, I\. Sohl, B\. Sokolowsky, Y\. Song, N\. Staudacher, F\. P\. Such, N\. Summers, I\. Sutskever, J\. Tang, N\. Tezak, M\. B\. Thompson, P\. Tillet, A\. Tootoonchian, E\. Tseng, P\. Tuggle, N\. Turley, J\. Tworek, J\. F\. C\. Uribe, A\. Vallone, A\. Vijayvergiya, C\. Voss, C\. Wainwright, J\. J\. Wang, A\. Wang, B\. Wang, J\. Ward, J\. Wei, C\. Weinmann, A\. Welihinda, P\. Welinder, J\. Weng, L\. Weng, M\. Wiethoff, D\. Willner, C\. Winter, S\. Wolrich, H\. Wong, L\. Workman, S\. Wu, J\. Wu, M\. Wu, K\. Xiao, T\. Xu, S\. Yoo, K\. Yu, Q\. Yuan, W\. Zaremba, R\. Zellers, C\. Zhang, M\. Zhang, S\. Zhao, T\. Zheng, J\. Zhuang, W\. Zhuk, and B\. ZophGPT\-4 technical report\.External Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§1](https://arxiv.org/html/2608.17895#S1.p1.1)\.
- Phukanet al\.\(2025\)A\. Phukan, Divyansh, H\. K\. Morj, Vaishnavi, A\. Saxena, and K\. GoswamiBeyond logit lens: contextual embeddings for robust hallucination detection & grounding in VLMs\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 9661–9675\.External Links:[Link](https://aclanthology.org/2025.naacl-long.488/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.488),ISBN 979\-8\-89176\-189\-6Cited by:[§5\.4](https://arxiv.org/html/2608.17895#S5.SS4.SSS0.Px1.p1.1)\.
- Qwen Team \(2026a\)Qwen TeamQwen3\.5: towards native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Qwen Team \(2026b\)Qwen TeamQwen3\.6\-Plus: towards real world agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.6)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Sahuet al\.\(2024\)P\. Sahu, K\. Sikka, and A\. DivakaranPelican: correcting hallucination in vision\-LLMs via claim decomposition and program of thought verification\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 8228–8248\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.470/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.470)Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Tanget al\.\(2025\)J\. Tang, Q\. Liu, Y\. Ye, J\. Lu, S\. Wei, C\. Lin, W\. Li, M\. F\. F\. B\. Mahmood, H\. Feng, Z\. Zhao, Y\. Wang, Y\. Liu, H\. Liu, X\. Bai, and C\. HuangMTVQA: benchmarking multilingual text\-centric visual question answering\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 7764–7794\.External Links:[Link](https://aclanthology.org/2025.findings-acl.404/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.404),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2608.17895#S1.p2.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p2.1)\.
- Tanget al\.\(2026\)Z\. Tang, E\. Haihong, R\. Li, J\. Liu, L\. Jia, Z\. Hao, Z\. Yang, Y\. Li, H\. Tian, X\. Hu,et al\.Finmmdocr: benchmarking financial multimodal reasoning with scenario awareness, document understanding, and multi\-step computation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 25858–25866\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Teamet al\.\(2025\)G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican, D\. Silver, M\. Johnson, I\. Antonoglou, J\. Schrittwieser, A\. Glaese, J\. Chen, E\. Pitler, T\. Lillicrap, A\. Lazaridou, O\. Firat, J\. Molloy, M\. Isard, P\. R\. Barham, T\. Hennigan, B\. Lee, F\. Viola, M\. Reynolds, Y\. Xu, R\. Doherty, E\. Collins, C\. Meyer, E\. Rutherford, E\. Moreira, K\. Ayoub, M\. Goel, J\. Krawczyk, C\. Du, E\. Chi, H\. Cheng, E\. Ni, P\. Shah, P\. Kane, B\. Chan, M\. Faruqui, A\. Severyn, H\. Lin, Y\. Li, Y\. Cheng, A\. Ittycheriah, M\. Mahdieh, M\. Chen, P\. Sun, D\. Tran, S\. Bagri, B\. Lakshminarayanan, J\. Liu, A\. Orban, F\. Güra, H\. Zhou, X\. Song, A\. Boffy, H\. Ganapathy, S\. Zheng, H\. Choe, Á\. Weisz, T\. Zhu, Y\. Lu, S\. Gopal, J\. Kahn, M\. Kula, J\. Pitman, R\. Shah, E\. Taropa, M\. A\. Merey, M\. Baeuml, Z\. Chen, L\. E\. Shafey, Y\. Zhang, O\. Sercinoglu, G\. Tucker, E\. Piqueras, M\. Krikun, I\. Barr, N\. Savinov, I\. Danihelka, B\. Roelofs, A\. White, A\. Andreassen, T\. von Glehn, L\. Yagati, M\. Kazemi, L\. Gonzalez, M\. Khalman, J\. Sygnowski, A\. Frechette, C\. Smith, L\. Culp, L\. Proleev, Y\. Luan, X\. Chen, J\. Lottes, N\. Schucher, F\. Lebron, A\. Rrustemi, N\. Clay, P\. Crone, T\. Kocisky, J\. Zhao, B\. Perz, D\. Yu, H\. Howard, A\. Bloniarz, J\. W\. Rae, H\. Lu, L\. Sifre, M\. Maggioni, F\. Alcober, D\. Garrette, M\. Barnes, S\. Thakoor, J\. Austin, G\. Barth\-Maron, W\. Wong, R\. Joshi, R\. Chaabouni, D\. Fatiha, A\. Ahuja, G\. S\. Tomar, E\. Senter, M\. Chadwick, I\. Kornakov, N\. Attaluri, I\. Iturrate, R\. Liu, Y\. Li, S\. Cogan, J\. Chen, C\. Jia, C\. Gu, Q\. Zhang, J\. Grimstad, A\. J\. Hartman, X\. Garcia, T\. S\. Pillai, J\. Devlin, M\. Laskin, D\. de Las Casas, D\. Valter, C\. Tao, L\. Blanco, A\. P\. Badia, D\. Reitter, M\. Chen, J\. Brennan, C\. Rivera, S\. Brin, S\. Iqbal, G\. Surita, J\. Labanowski, A\. Rao, S\. Winkler, E\. Parisotto, Y\. Gu, K\. Olszewska, R\. Addanki, A\. Miech, A\. Louis, D\. Teplyashin, G\. Brown, E\. Catt, J\. Balaguer, J\. Xiang, P\. Wang, Z\. Ashwood, A\. Briukhov, A\. Webson, S\. Ganapathy, S\. Sanghavi, A\. Kannan, M\. Chang, A\. Stjerngren, J\. Djolonga, Y\. Sun, A\. Bapna, M\. Aitchison, P\. Pejman, H\. Michalewski, T\. Yu, C\. Wang, J\. Love, J\. Ahn, D\. Bloxwich, K\. Han, P\. Humphreys, T\. Sellam, J\. Bradbury, V\. Godbole, S\. Samangooei, B\. Damoc, A\. Kaskasoli, S\. M\. R\. Arnold, V\. Vasudevan, S\. Agrawal, J\. Riesa, D\. Lepikhin, R\. Tanburn, S\. Srinivasan, H\. Lim, S\. Hodkinson, P\. Shyam, J\. Ferret, S\. Hand, A\. Garg, T\. L\. Paine, J\. Li, Y\. Li, M\. Giang, A\. Neitz, Z\. Abbas, S\. York, M\. Reid, E\. Cole, A\. Chowdhery, D\. Das, D\. Rogozińska, V\. Nikolaev, P\. Sprechmann, Z\. Nado, L\. Zilka, F\. Prost, L\. He, M\. Monteiro, G\. Mishra, C\. Welty, J\. Newlan, D\. Jia, M\. Allamanis, C\. H\. Hu, R\. de Liedekerke, J\. Gilmer, C\. Saroufim, S\. Rijhwani, S\. Hou, D\. Shrivastava, A\. Baddepudi, A\. Goldin, A\. Ozturel, A\. Cassirer, Y\. Xu, D\. Sohn, D\. Sachan, R\. K\. Amplayo, C\. Swanson, D\. Petrova, S\. Narayan, A\. Guez, S\. Brahma, J\. Landon, M\. Patel, R\. Zhao, K\. Villela, L\. Wang, W\. Jia, M\. Rahtz, M\. Giménez, L\. Yeung, J\. Keeling, P\. Georgiev, D\. Mincu, B\. Wu, S\. Haykal, R\. Saputro, K\. Vodrahalli, J\. Qin, Z\. Cankara, A\. Sharma, N\. Fernando, W\. Hawkins, B\. Neyshabur, S\. Kim, A\. Hutter, P\. Agrawal, A\. Castro\-Ros, G\. van den Driessche, T\. Wang, F\. Yang, S\. Chang, P\. Komarek, R\. McIlroy, M\. Lučić, G\. Zhang, W\. Farhan, M\. Sharman, P\. Natsev, P\. Michel, Y\. Bansal, S\. Qiao, K\. Cao, S\. Shakeri, C\. Butterfield, J\. Chung, P\. K\. Rubenstein, S\. Agrawal, A\. Mensch, K\. Soparkar, K\. Lenc, T\. Chung, A\. Pope, L\. Maggiore, J\. Kay, P\. Jhakra, S\. Wang, J\. Maynez, M\. Phuong, T\. Tobin, A\. Tacchetti, M\. Trebacz, K\. Robinson, Y\. Katariya, S\. Riedel, P\. Bailey, K\. Xiao, N\. Ghelani, L\. Aroyo, A\. Slone, N\. Houlsby, X\. Xiong, Z\. Yang, E\. Gribovskaya, J\. Adler, M\. Wirth, L\. Lee, M\. Li, T\. Kagohara, J\. Pavagadhi, S\. Bridgers, A\. Bortsova, S\. Ghemawat, Z\. Ahmed, T\. Liu, R\. Powell, V\. Bolina, M\. Iinuma, P\. Zablotskaia, J\. Besley, D\. Chung, T\. Dozat, R\. Comanescu, X\. Si, J\. Greer, G\. Su, M\. Polacek, R\. L\. Kaufman, S\. Tokumine, H\. Hu, E\. Buchatskaya, Y\. Miao, M\. Elhawaty, A\. Siddhant, N\. Tomasev, J\. Xing, C\. Greer, H\. Miller, S\. Ashraf, A\. Roy, Z\. Zhang, A\. Ma, A\. Filos, M\. Besta, R\. Blevins, T\. Klimenko, C\. Yeh, S\. Changpinyo, J\. Mu, O\. Chang, M\. Pajarskas, C\. Muir, V\. Cohen, C\. L\. Lan, K\. Haridasan, A\. Marathe, S\. Hansen, S\. Douglas, R\. Samuel, M\. Wang, S\. Austin, C\. Lan, J\. Jiang, J\. Chiu, J\. A\. Lorenzo, L\. L\. Sjösund, S\. Cevey, Z\. Gleicher, T\. Avrahami, A\. Boral, H\. Srinivasan, V\. Selo, R\. May, K\. Aisopos, L\. Hussenot, L\. B\. Soares, K\. Baumli, M\. B\. Chang, A\. Recasens, B\. Caine, A\. Pritzel, F\. Pavetic, F\. Pardo, A\. Gergely, J\. Frye, V\. Ramasesh, D\. Horgan, K\. Badola, N\. Kassner, S\. Roy, E\. Dyer, V\. C\. Campos, A\. Tomala, Y\. Tang, D\. E\. Badawy, E\. White, B\. Mustafa, O\. Lang, A\. Jindal, S\. Vikram, Z\. Gong, S\. Caelles, R\. Hemsley, G\. Thornton, F\. Feng, W\. Stokowiec, C\. Zheng, P\. Thacker, Ç\. Ünlü, Z\. Zhang, M\. Saleh, J\. Svensson, M\. Bileschi, P\. Patil, A\. Anand, R\. Ring, K\. Tsihlas, A\. Vezer, M\. Selvi, T\. Shevlane, M\. Rodriguez, T\. Kwiatkowski, S\. Daruki, K\. Rong, A\. Dafoe, N\. FitzGerald, K\. Gu\-Lemberg, M\. Khan, L\. A\. Hendricks, M\. Pellat, V\. Feinberg, J\. Cobon\-Kerr, T\. Sainath, M\. Rauh, S\. H\. Hashemi, R\. Ives, Y\. Hasson, E\. Noland, Y\. Cao, N\. Byrd, L\. Hou, Q\. Wang, T\. Sottiaux, M\. Paganini, J\. Lespiau, A\. Moufarek, S\. Hassan, K\. Shivakumar, J\. van Amersfoort, A\. Mandhane, P\. Joshi, A\. Goyal, M\. Tung, A\. Brock, H\. Sheahan, V\. Misra, C\. Li, N\. Rakićević, M\. Dehghani, F\. Liu, S\. Mittal, J\. Oh, S\. Noury, E\. Sezener, F\. Huot, M\. Lamm, N\. D\. Cao, C\. Chen, S\. Mudgal, R\. Stella, K\. Brooks, G\. Vasudevan, C\. Liu, M\. Chain, N\. Melinkeri, A\. Cohen, V\. Wang, K\. Seymore, S\. Zubkov, R\. Goel, S\. Yue, S\. Krishnakumaran, B\. Albert, N\. Hurley, M\. Sano, A\. Mohananey, J\. Joughin, E\. Filonov, T\. Kępa, Y\. Eldawy, J\. Lim, R\. Rishi, S\. Badiezadegan, T\. Bos, J\. Chang, S\. Jain, S\. G\. S\. Padmanabhan, S\. Puttagunta, K\. Krishna, L\. Baker, N\. Kalb, V\. Bedapudi, A\. Kurzrok, S\. Lei, A\. Yu, O\. Litvin, X\. Zhou, Z\. Wu, S\. Sobell, A\. Siciliano, A\. Papir, R\. Neale, J\. Bragagnolo, T\. Toor, T\. Chen, V\. Anklin, F\. Wang, R\. Feng, M\. Gholami, K\. Ling, L\. Liu, J\. Walter, H\. Moghaddam, A\. Kishore, J\. Adamek, T\. Mercado, J\. Mallinson, S\. Wandekar, S\. Cagle, E\. Ofek, G\. Garrido, C\. Lombriser, M\. Mukha, B\. Sun, H\. R\. Mohammad, J\. Matak, Y\. Qian, V\. Peswani, P\. Janus, Q\. Yuan, L\. Schelin, O\. David, A\. Garg, Y\. He, O\. Duzhyi, A\. Älgmyr, T\. Lottaz, Q\. Li, V\. Yadav, L\. Xu, A\. Chinien, R\. Shivanna, A\. Chuklin, J\. Li, C\. Spadine, T\. Wolfe, K\. Mohamed, S\. Das, Z\. Dai, K\. He, D\. von Dincklage, S\. Upadhyay, A\. Maurya, L\. Chi, S\. Krause, K\. Salama, P\. G\. Rabinovitch, P\. K\. R\. M, A\. Selvan, M\. Dektiarev, G\. Ghiasi, E\. Guven, H\. Gupta, B\. Liu, D\. Sharma, I\. H\. Shtacher, S\. Paul, O\. Akerlund, F\. Aubet, T\. Huang, C\. Zhu, E\. Zhu, E\. Teixeira, M\. Fritze, F\. Bertolini, L\. Marinescu, M\. Bölle, D\. Paulus, K\. Gupta, T\. Latkar, M\. Chang, J\. Sanders, R\. Wilson, X\. Wu, Y\. Tan, L\. N\. Thiet, T\. Doshi, S\. Lall, S\. Mishra, W\. Chen, T\. Luong, S\. Benjamin, J\. Lee, E\. Andrejczuk, D\. Rabiej, V\. Ranjan, K\. Styrc, P\. Yin, J\. Simon, M\. R\. Harriott, M\. Bansal, A\. Robsky, G\. Bacon, D\. Greene, D\. Mirylenka, C\. Zhou, O\. Sarvana, A\. Goyal, S\. Andermatt, P\. Siegler, B\. Horn, A\. Israel, F\. Pongetti, C\. "\. Chen, M\. Selvatici, P\. Silva, K\. Wang, J\. Tolins, K\. Guu, R\. Yogev, X\. Cai, A\. Agostini, M\. Shah, H\. Nguyen, N\. Ó\. Donnaile, S\. Pereira, L\. Friso, A\. Stambler, A\. Kurzrok, C\. Kuang, Y\. Romanikhin, M\. Geller, Z\. Yan, K\. Jang, C\. Lee, W\. Fica, E\. Malmi, Q\. Tan, D\. Banica, D\. Balle, R\. Pham, Y\. Huang, D\. Avram, H\. Shi, J\. Singh, C\. Hidey, N\. Ahuja, P\. Saxena, D\. Dooley, S\. P\. Potharaju, E\. O’Neill, A\. Gokulchandran, R\. Foley, K\. Zhao, M\. Dusenberry, Y\. Liu, P\. Mehta, R\. Kotikalapudi, C\. Safranek\-Shrader, A\. Goodman, J\. Kessinger, E\. Globen, P\. Kolhar, C\. Gorgolewski, A\. Ibrahim, Y\. Song, A\. Eichenbaum, T\. Brovelli, S\. Potluri, P\. Lahoti, C\. Baetu, A\. Ghorbani, C\. Chen, A\. Crawford, S\. Pal, M\. Sridhar, P\. Gurita, A\. Mujika, I\. Petrovski, P\. Cedoz, C\. Li, S\. Chen, N\. D\. Santo, S\. Goyal, J\. Punjabi, K\. Kappaganthu, C\. Kwak, P\. LV, S\. Velury, H\. Choudhury, J\. Hall, P\. Shah, R\. Figueira, M\. Thomas, M\. Lu, T\. Zhou, C\. Kumar, T\. Jurdi, S\. Chikkerur, Y\. Ma, A\. Yu, S\. Kwak, V\. Ähdel, S\. Rajayogam, T\. Choma, F\. Liu, A\. Barua, C\. Ji, J\. H\. Park, V\. Hellendoorn, A\. Bailey, T\. Bilal, H\. Zhou, M\. Khatir, C\. Sutton, W\. Rzadkowski, F\. Macintosh, R\. Vij, K\. Shagin, P\. Medina, C\. Liang, J\. Zhou, P\. Shah, Y\. Bi, A\. Dankovics, S\. Banga, S\. Lehmann, M\. Bredesen, Z\. Lin, J\. E\. Hoffmann, J\. Lai, R\. Chung, K\. Yang, N\. Balani, A\. Bražinskas, A\. Sozanschi, M\. Hayes, H\. F\. Alcalde, P\. Makarov, W\. Chen, A\. Stella, L\. Snijders, M\. Mandl, A\. Kärrman, P\. Nowak, X\. Wu, A\. Dyck, K\. Vaidyanathan, R\. R, J\. Mallet, M\. Rudominer, E\. Johnston, S\. Mittal, A\. Udathu, J\. Christensen, V\. Verma, Z\. Irving, A\. Santucci, G\. Elsayed, E\. Davoodi, M\. Georgiev, I\. Tenney, N\. Hua, G\. Cideron, E\. Leurent, M\. Alnahlawi, I\. Georgescu, N\. Wei, I\. Zheng, D\. Scandinaro, H\. Jiang, J\. Snoek, M\. Sundararajan, X\. Wang, Z\. Ontiveros, I\. Karo, J\. Cole, V\. Rajashekhar, L\. Tumeh, E\. Ben\-David, R\. Jain, J\. Uesato, R\. Datta, O\. Bunyan, S\. Wu, J\. Zhang, P\. Stanczyk, Y\. Zhang, D\. Steiner, S\. Naskar, M\. Azzam, M\. Johnson, A\. Paszke, C\. Chiu, J\. S\. Elias, A\. Mohiuddin, F\. Muhammad, J\. Miao, A\. Lee, N\. Vieillard, J\. Park, J\. Zhang, J\. Stanway, D\. Garmon, A\. Karmarkar, Z\. Dong, J\. Lee, A\. Kumar, L\. Zhou, J\. Evens, W\. Isaac, G\. Irving, E\. Loper, M\. Fink, I\. Arkatkar, N\. Chen, I\. Shafran, I\. Petrychenko, Z\. Chen, J\. Jia, A\. Levskaya, Z\. Zhu, P\. Grabowski, Y\. Mao, A\. Magni, K\. Yao, J\. Snaider, N\. Casagrande, E\. Palmer, P\. Suganthan, A\. Castaño, I\. Giannoumis, W\. Kim, M\. Rybiński, A\. Sreevatsa, J\. Prendki, D\. Soergel, A\. Goedeckemeyer, W\. Gierke, M\. Jafari, M\. Gaba, J\. Wiesner, D\. G\. Wright, Y\. Wei, H\. Vashisht, Y\. Kulizhskaya, J\. Hoover, M\. Le, L\. Li, C\. Iwuanyanwu, L\. Liu, K\. Ramirez, A\. Khorlin, A\. Cui, T\. LIN, M\. Wu, R\. Aguilar, K\. Pallo, A\. Chakladar, G\. Perng, E\. A\. Abellan, M\. Zhang, I\. Dasgupta, N\. Kushman, I\. Penchev, A\. Repina, X\. Wu, T\. van der Weide, P\. Ponnapalli, C\. Kaplan, J\. Simsa, S\. Li, O\. Dousse, F\. Yang, J\. Piper, N\. Ie, R\. Pasumarthi, N\. Lintz, A\. Vijayakumar, D\. Andor, P\. Valenzuela, M\. Lui, C\. Paduraru, D\. Peng, K\. Lee, S\. Zhang, S\. Greene, D\. D\. Nguyen, P\. Kurylowicz, C\. Hardin, L\. Dixon, L\. Janzer, K\. Choo, Z\. Feng, B\. Zhang, A\. Singhal, D\. Du, D\. McKinnon, N\. Antropova, T\. Bolukbasi, O\. Keller, D\. Reid, D\. Finchelstein, M\. A\. Raad, R\. Crocker, P\. Hawkins, R\. Dadashi, C\. Gaffney, K\. Franko, A\. Bulanova, R\. Leblond, S\. Chung, H\. Askham, L\. C\. Cobo, K\. Xu, F\. Fischer, J\. Xu, C\. Sorokin, C\. Alberti, C\. Lin, C\. Evans, A\. Dimitriev, H\. Forbes, D\. Banarse, Z\. Tung, M\. Omernick, C\. Bishop, R\. Sterneck, R\. Jain, J\. Xia, E\. Amid, F\. Piccinno, X\. Wang, P\. Banzal, D\. J\. Mankowitz, A\. Polozov, V\. Krakovna, S\. Brown, M\. Bateni, D\. Duan, V\. Firoiu, M\. Thotakuri, T\. Natan, M\. Geist, S\. tan Girgin, H\. Li, J\. Ye, O\. Roval, R\. Tojo, M\. Kwong, J\. Lee\-Thorp, C\. Yew, D\. Sinopalnikov, S\. Ramos, J\. Mellor, A\. Sharma, K\. Wu, D\. Miller, N\. Sonnerat, D\. Vnukov, R\. Greig, J\. Beattie, E\. Caveness, L\. Bai, J\. Eisenschlos, A\. Korchemniy, T\. Tsai, M\. Jasarevic, W\. Kong, P\. Dao, Z\. Zheng, F\. Liu, F\. Yang, R\. Zhu, T\. H\. Teh, J\. Sanmiya, E\. Gladchenko, N\. Trdin, D\. Toyama, E\. Rosen, S\. Tavakkol, L\. Xue, C\. Elkind, O\. Woodman, J\. Carpenter, G\. Papamakarios, R\. Kemp, S\. Kafle, T\. Grunina, R\. Sinha, A\. Talbert, D\. Wu, D\. Owusu\-Afriyie, C\. Du, C\. Thornton, J\. Pont\-Tuset, P\. Narayana, J\. Li, S\. Fatehi, J\. Wieting, O\. Ajmeri, B\. Uria, Y\. Ko, L\. Knight, A\. Héliou, N\. Niu, S\. Gu, C\. Pang, Y\. Li, N\. Levine, A\. Stolovich, R\. Santamaria\-Fernandez, S\. Goenka, W\. Yustalim, R\. Strudel, A\. Elqursh, C\. Deck, H\. Lee, Z\. Li, K\. Levin, R\. Hoffmann, D\. Holtmann\-Rice, O\. Bachem, S\. Arora, C\. Koh, S\. H\. Yeganeh, S\. Põder, M\. Tariq, Y\. Sun, L\. Ionita, M\. Seyedhosseini, P\. Tafti, Z\. Liu, A\. Gulati, J\. Liu, X\. Ye, B\. Chrzaszcz, L\. Wang, N\. Sethi, T\. Li, B\. Brown, S\. Singh, W\. Fan, A\. Parisi, J\. Stanton, V\. Koverkathu, C\. A\. Choquette\-Choo, Y\. Li, T\. Lu, A\. Ittycheriah, P\. Shroff, M\. Varadarajan, S\. Bahargam, R\. Willoughby, D\. Gaddy, G\. Desjardins, M\. Cornero, B\. Robenek, B\. Mittal, B\. Albrecht, A\. Shenoy, F\. Moiseev, H\. Jacobsson, A\. Ghaffarkhah, M\. Rivière, A\. Walton, C\. Crepy, A\. Parrish, Z\. Zhou, C\. Farabet, C\. Radebaugh, P\. Srinivasan, C\. van der Salm, A\. Fidjeland, S\. Scellato, E\. Latorre\-Chimoto, H\. Klimczak\-Plucińska, D\. Bridson, D\. de Cesare, T\. Hudson, P\. Mendolicchio, L\. Walker, A\. Morris, M\. Mauger, A\. Guseynov, A\. Reid, S\. Odoom, L\. Loher, V\. Cotruta, M\. Yenugula, D\. Grewe, A\. Petrushkina, T\. Duerig, A\. Sanchez, S\. Yadlowsky, A\. Shen, A\. Globerson, L\. Webb, S\. Dua, D\. Li, S\. Bhupatiraju, D\. Hurt, H\. Qureshi, A\. Agarwal, T\. Shani, M\. Eyal, A\. Khare, S\. R\. Belle, L\. Wang, C\. Tekur, M\. S\. Kale, J\. Wei, R\. Sang, B\. Saeta, T\. Liechty, Y\. Sun, Y\. Zhao, S\. Lee, P\. Nayak, D\. Fritz, M\. R\. Vuyyuru, J\. Aslanides, N\. Vyas, M\. Wicke, X\. Ma, E\. Eltyshev, N\. Martin, H\. Cate, J\. Manyika, K\. Amiri, Y\. Kim, X\. Xiong, K\. Kang, F\. Luisier, N\. Tripuraneni, D\. Madras, M\. Guo, A\. Waters, O\. Wang, J\. Ainslie, J\. Baldridge, H\. Zhang, G\. Pruthi, J\. Bauer, F\. Yang, R\. Mansour, J\. Gelman, Y\. Xu, G\. Polovets, J\. Liu, H\. Cai, W\. Chen, X\. Sheng, E\. Xue, S\. Ozair, C\. Angermueller, X\. Li, A\. Sinha, W\. Wang, J\. Wiesinger, E\. Koukoumidis, Y\. Tian, A\. Iyer, M\. Gurumurthy, M\. Goldenson, P\. Shah, M\. Blake, H\. Yu, A\. Urbanowicz, J\. Palomaki, C\. Fernando, K\. Durden, H\. Mehta, N\. Momchev, E\. Rahimtoroghi, M\. Georgaki, A\. Raul, S\. Ruder, M\. Redshaw, J\. Lee, D\. Zhou, K\. Jalan, D\. Li, B\. Hechtman, P\. Schuh, M\. Nasr, K\. Milan, V\. Mikulik, J\. Franco, T\. Green, N\. Nguyen, J\. Kelley, A\. Mahendru, A\. Hu, J\. Howland, B\. Vargas, J\. Hui, K\. Bansal, V\. Rao, R\. Ghiya, E\. Wang, K\. Ye, J\. M\. Sarr, M\. M\. Preston, M\. Elish, S\. Li, A\. Kaku, J\. Gupta, I\. Pasupat, D\. Juan, M\. Someswar, T\. M\., X\. Chen, A\. Amini, A\. Fabrikant, E\. Chu, X\. Dong, A\. Muthal, S\. Buthpitiya, S\. Jauhari, N\. Hua, U\. Khandelwal, A\. Hitron, J\. Ren, L\. Rinaldi, S\. Drath, A\. Dabush, N\. Jiang, H\. Godhia, U\. Sachs, A\. Chen, Y\. Fan, H\. Taitelbaum, H\. Noga, Z\. Dai, J\. Wang, C\. Liang, J\. Hamer, C\. Ferng, C\. Elkind, A\. Atias, P\. Lee, V\. Listík, M\. Carlen, J\. van de Kerkhof, M\. Pikus, K\. Zaher, P\. Müller, S\. Zykova, R\. Stefanec, V\. Gatsko, C\. Hirnschall, A\. Sethi, X\. F\. Xu, C\. Ahuja, B\. Tsai, A\. Stefanoiu, B\. Feng, K\. Dhandhania, M\. Katyal, A\. Gupta, A\. Parulekar, D\. Pitta, J\. Zhao, V\. Bhatia, Y\. Bhavnani, O\. Alhadlaq, X\. Li, P\. Danenberg, D\. Tu, A\. Pine, V\. Filippova, A\. Ghosh, B\. Limonchik, B\. Urala, C\. K\. Lanka, D\. Clive, Y\. Sun, E\. Li, H\. Wu, K\. Hongtongsak, I\. Li, K\. Thakkar, K\. Omarov, K\. Majmundar, M\. Alverson, M\. Kucharski, M\. Patel, M\. Jain, M\. Zabelin, P\. Pelagatti, R\. Kohli, S\. Kumar, J\. Kim, S\. Sankar, V\. Shah, L\. Ramachandruni, X\. Zeng, B\. Bariach, L\. Weidinger, T\. Vu, A\. Andreev, A\. He, K\. Hui, S\. Kashem, A\. Subramanya, S\. Hsiao, D\. Hassabis, K\. Kavukcuoglu, A\. Sadovsky, Q\. Le, T\. Strohman, Y\. Wu, S\. Petrov, J\. Dean, and O\. VinyalsGemini: a family of highly capable multimodal models\.External Links:2312\.11805,[Link](https://arxiv.org/abs/2312.11805)Cited by:[§1](https://arxiv.org/html/2608.17895#S1.p1.1)\.
- Team \(2026\)G\. TeamGemma 4 technical report\.External Links:2607\.02770,[Link](https://arxiv.org/abs/2607.02770)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Team \(2025\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§5\.1](https://arxiv.org/html/2608.17895#S5.SS1.SSS0.Px1.p1.1)\.
- Tonget al\.\(2026\)C\. Tong, Q\. Zhang, C\. Li, L\. Jiang, and Y\. LiuFaithSCAN: model\-driven single\-pass hallucination detection for faithful visual question answering\.arXiv preprint arXiv:2601\.00269\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2024\)Z\. Wang, M\. Xia, L\. He, H\. Chen, Y\. Liu, R\. Zhu, K\. Liang, X\. Wu, H\. Liu, S\. Malladi,et al\.Charxiv: charting gaps in realistic chart understanding in multimodal llms\.Advances in Neural Information Processing Systems37,pp\. 113569–113697\.Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.4.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2024\)J\. Wu, Q\. Liu, D\. Wang, J\. Zhang, S\. Wu, L\. Wang, and T\. TanLogical closed loop: uncovering object hallucinations in large vision\-language models\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 6944–6962\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Xuet al\.\(2026\)Z\. Xu, J\. Ji, Z\. Chen, Z\. Liu, Q\. Liu, C\. Peng, Z\. Qin, Z\. Xu, J\. Wan, J\. Tang, Z\. Yang, S\. Bai, and D\. LiuCC\-ocr v2: benchmarking large multimodal models for literacy in real\-world document processing\.arXiv preprint arXiv:2605\.03903\.Cited by:[Table 1](https://arxiv.org/html/2608.17895#S1.T1.2.1.9.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Yinet al\.\(2024\)S\. Yin, C\. Fu, S\. Zhao, T\. Xu, H\. Wang, D\. Sui, Y\. Shen, K\. Li, X\. Sun, and E\. ChenWoodpecker: hallucination correction for multimodal large language models\.Science China Information Sciences67\(12\),pp\. 220105\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Yueet al\.\(2024\)X\. Yue, Y\. Ni, T\. Zheng, K\. Zhang, R\. Liu, G\. Zhang, S\. Stevens, D\. Jiang, W\. Ren, Y\. Sun, C\. Wei, B\. Yu, R\. Yuan, R\. Sun, M\. Yin, B\. Zheng, Z\. Yang, Y\. Liu, W\. Huang, H\. Sun, Y\. Su, and W\. ChenMMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert agi\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Vol\.,pp\. 9556–9567\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.00913)Cited by:[§1](https://arxiv.org/html/2608.17895#S1.p1.1),[§1](https://arxiv.org/html/2608.17895#S1.p2.1),[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Yueet al\.\(2025\)X\. Yue, T\. Zheng, Y\. Ni, Y\. Wang, K\. Zhang, S\. Tong, Y\. Sun, B\. Yu, G\. Zhang, H\. Sun,et al\.Mmmu\-pro: a more robust multi\-discipline multimodal understanding benchmark\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15134–15186\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.
- Zhaiet al\.\(2023\)X\. Zhai, B\. Mustafa, A\. Kolesnikov, and L\. BeyerSigmoid loss for language image pre\-training\.In2023 IEEE/CVF International Conference on Computer Vision \(ICCV\),Vol\.,pp\. 11941–11952\.External Links:[Document](https://dx.doi.org/10.1109/ICCV51070.2023.01100)Cited by:[§3\.2\.2](https://arxiv.org/html/2608.17895#S3.SS2.SSS2.Px1.p1.1)\.
- Zhanget al\.\(2026\)F\. Zhang, Y\. Wu, Z\. Wang, X\. Wang, C\. Lv, X\. Huang, and X\. ZhengVib\-probe: detecting and mitigating hallucinations in vision\-language models via variational information bottleneck\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 23509–23521\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025\)K\. Zhang, L\. Niu, Z\. Cao, F\. Meng, and J\. ZhouTIU\-bench: a benchmark for evaluating large multimodal models on text\-rich image understanding\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 24286–24295\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1318/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1318),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p2.1)\.
- Zhanget al\.\(2024\)R\. Zhang, H\. Zhang, and Z\. ZhengVl\-uncertainty: detecting hallucination in large vision\-language model via uncertainty estimation\.arXiv preprint arXiv:2411\.11919\.Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px2.p1.1)\.
- Zuoet al\.\(2025\)Y\. Zuo, S\. Qu, Y\. Li, Z\. Chen, X\. Zhu, E\. Hua, K\. Zhang, N\. Ding, and B\. ZhouMedXpertQA: benchmarking expert\-level medical reasoning and understanding\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=IyVcxU0RKI)Cited by:[§2](https://arxiv.org/html/2608.17895#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix AError taxonomy

Below, we elaborate on the error taxonomy derived in our study\.

1. 1\.C1 — Spatial localization and object matching\.The model incorrectly identifies where objects are located in the image and how they relate to each other \(e\.g\., selecting a neighboring element, misreading whether a marker lies inside a region, boundary intersections, label\-to\-object matching, arrow direction, or links between blocks\)\.
2. 2\.C2 — Counting and aggregation errors of visual elements\.The model makes mistakes when counting objects or aggregating extracted elements \(e\.g\., points, circles, arrows, rows, columns, people, links, paths, labels, or table values\), including missed elements, extra elements, or incorrect summation\.
3. 3\.C3 — Errors in reading text, numbers, and visual attributes\.The model incorrectly reads text, numbers, symbols, or small labels \(OCR\-related issues\), and may also misidentify visual attributes such as color, marker shape, text style, italics/boldface, or legend encodings\.
4. 4\.C4 — Errors in extracting values from charts\.The model incorrectly reads quantitative values from plots/graphs \(e\.g\., axis scale, ticks, point coordinates, values at specificxx, peaks, minima, maxima, plateaus, trends, or ranges\)\.
5. 5\.C5 — Semantic, instruction\-following, and logical/arithmetic errors\.The model misunderstands task conditions, categories, terms, units, or filtering rules, and/or makes reasoning or arithmetic mistakes after extraction \(e\.g\., wrong interpretation, entity confusion, incorrect formulas, hallucinated assumptions, or incomplete answers\)\.

## Appendix BIllustrative items from MWS Vision Bench

Figure[5](https://arxiv.org/html/2608.17895#A2.F5)shows public validation items from MWS Vision Bench[19](https://arxiv.org/html/2608.17895#bib.bib11)\. The illustration is taken from the dataset’s Hugging Face page\.\*\*\*[https://huggingface\.co/datasets/MTSAIR/MWS\-Vision\-Bench](https://huggingface.co/datasets/MTSAIR/MWS-Vision-Bench)We include them to make the contrast with BEAR\-Bench concrete: MWS mixes business scans with personal handwriting, receipts, and form\-style pages, and a large share of its tasks are OCR, grounding, and key\-information extraction\. BEAR\-Bench instead targets multi\-step questions on text\-dense scientific and business document pages \(Figure[1](https://arxiv.org/html/2608.17895#S3.F1)\)\.

![Refer to caption](https://arxiv.org/html/2608.17895v1/images/mws_vision_bench_examples.jpg)Figure 5:Illustrative items from the public validation split of MWS Vision Bench[19](https://arxiv.org/html/2608.17895#bib.bib11)\. The mix of personal handwriting, receipts, and document\-processing tasks differs from BEAR\-Bench’s professional, text\-dense pages\.
## Appendix CAccuracy vs Reasoning Depth

Figure[6](https://arxiv.org/html/2608.17895#A3.F6)demonstrates the accuracy broken down by annotated reasoning\-step count for several proprietary models\. The trend is non\-monotonic, suggesting that on knowledge\-free tasks, frontier models are constrained more by visual perception than by the ability to execute long reasoning chains\.

We repeat this analysis for four open\-weight models \(Figure[7](https://arxiv.org/html/2608.17895#A3.F7)\), where the two model families exhibit contrasting behaviour\. The accuracy of the reasoning\-tuned Qwen3\.5 models remains stable as the annotated reasoning depth grows, whereas the instruction\-tuned Qwen3\-VL models degrade on items requiring seven or more steps\. A likely reason is that reasoning\-tuned models are trained to produce long chains of reasoning, so questions that need more steps cost them little extra accuracy, while instruction\-tuned models are not trained for this and lose accuracy as more steps are needed\.

Figure 6:Accuracy on BEAR\-Bench as a function of reasoning depth \(number of steps per item\) for three frontier models\.Figure 7:Accuracy on BEAR\-Bench as a function of reasoning depth \(number of steps per item\) for four open\-weight models\.
## Appendix DAdditional data statistics: response length

Table[5](https://arxiv.org/html/2608.17895#A4.T5)reports the median response length in words for each evaluated model\. Lengths are computed by whitespace\-splitting each model’s stored answer field\. Proprietary and open\-weight models differ sharply: several API systems return very short answers \(median 2–3 words\), while reasoning\-oriented open\-weight models often produce long intermediate chains \(median above 1,000 words\)\.

Table 5:Median response length \(words\) by model\. Words are whitespace\-split tokens from each model’s storedanswerfield\.
## Appendix EHallucination Detection Metrics

This appendix reports the full 5\-fold cross\-validation results for the hallucination detectors evaluated in Section[5\.4](https://arxiv.org/html/2608.17895#S5.SS4)\. We compare uncertainty\-based scores, ContextualLens, SUQ Probe, and an MLLM\-as\-a\-judge baseline under two regimes: proxy\-based detection for proprietary models and native\-signal detection for open\-weight models\. Tables[6](https://arxiv.org/html/2608.17895#A5.T6)and[7](https://arxiv.org/html/2608.17895#A5.T7)present AUROC, AUC\-PR, and balanced accuracy for Gemini, Qwen, and Claude outputs; Table[8](https://arxiv.org/html/2608.17895#A5.T8)reports the corresponding results for Qwen3\.5\-9B and Qwen3\.5\-27B\.

Table 6:Hallucination detection \(5\-fold CV\): Gemini models\. AUROC and AUC\-PR are omitted for VLM\-as\-judge because the judge returns binary verdicts\.Table 7:Hallucination detection \(5\-fold CV\): Qwen and Claude models\. AUROC and AUC\-PR are omitted for VLM\-as\-judge because the judge returns binary verdicts\.Table 8:Hallucination detection \(5\-fold CV\): open\-weight Qwen3\.5 models \(native signals\)\. The VLM\-as\-a\-judge results are computed over 997 valid samples for each model\. AUROC and AUC\-PR are omitted for VLM\-as\-judge because the judge returns binary verdicts\.
## Appendix FChain\-of\-Thought Prompt

Figure[8](https://arxiv.org/html/2608.17895#A6.F8)shows the system prompt prepended to each image–question pair in the CoT condition reported in Table[3](https://arxiv.org/html/2608.17895#S5.T3)\.

Figure 8:Chain\-of\-thought system prompt used in the CoT evaluation condition\.
## Appendix GHuman Annotation

### G\.1Instructions for annotators

The instructions given to annotators are shown in Figures[9](https://arxiv.org/html/2608.17895#A7.F9)and[10](https://arxiv.org/html/2608.17895#A7.F10)\(the original text and its English translation, respectively\)\. The main goal for the annotators was to design complex multi\-hop questions that can be answered solely from the image content, without requiring any external expert knowledge\. We introduced several examples of “good” and “bad” questions to elaborate on the task\. The annotators were instructed to formulate the answers to the questions as briefly as possible \(e\.g\., a single number\)\.

![Refer to caption](https://arxiv.org/html/2608.17895v1/assesors_instruction_extract_ru.png)Figure 9:Original Russian instructions given to benchmark annotators\. The excerpt states the annotation objective \(one multi\-step question per document image with a ground\-truth answer\) and Section 2, which defines question requirements: at least three reasoning steps, OCR\-grounded text, exemplar chain types, items to avoid, and answer constraints\. The translation to English is provided in Figure[10](https://arxiv.org/html/2608.17895#A7.F10)\.![Refer to caption](https://arxiv.org/html/2608.17895v1/assesors_instruction_extract_en.png)Figure 10:English translation of the annotator instructions shown in Figure[9](https://arxiv.org/html/2608.17895#A7.F9)\.
### G\.2Recruitment & payment

Annotators were recruited through an open call posted in an internal student chat channel\. Prior to the main task, each candidate completed a qualification test in which they generated 10 probe questions designed to elicit incorrect answers from Gemini 3\.1 Pro\. Candidates who successfully caused the model to fail on more than 4 out of 10 questions were selected to work on the dataset\. Annotators were compensated at 4\.4 times the Russian minimum wage\.

### G\.3Ethics & Consent

No formal ethics review was required for this non\-invasive annotation task\. All participants provided informed consent\.

### G\.4Demographics

All annotators were aged 22–25 years and held at least a bachelor’s degree in a technical field, with self\-reported English proficiency at CEFR B2 or higher\. The sample included 60% male and 40% female participants\.

## Appendix HBroader Impact, Data Use, and Compute Details

##### Potential risks\.

BEAR\-Bench is intended for research evaluation rather than autonomous decision making\. Errors on its document\-reasoning tasks can arise from misreading text, tables, figures, or equations and can consequently produce incorrect calculations or unsupported conclusions\. If similar systems are used without human verification in financial, scientific, or other high\-stakes workflows, such errors could lead to incorrect analyses or decisions\. Performance also varies across languages and document types; therefore, aggregate benchmark scores should not be interpreted as evidence of reliable performance for every user population or document genre\. We recommend using the benchmark for comparative evaluation and retaining human oversight in consequential settings\.

##### Intended use\.

BEAR\-Bench is designed primarily as an evaluation benchmark for multimodal reasoning and hallucination detection\. It does not provide step\-by\-step rationale annotations, and therefore does not directly support training or evaluating explicit chain\-of\-thought reasoning\. The benchmark should not be interpreted as a resource for certifying models for deployment in high\-stakes financial, legal, or scientific decision\-making settings\.

##### Privacy and content\.

All source pages were obtained from publicly accessible official disclosure, government, intergovernmental, or academic sources, as described in Section[3](https://arxiv.org/html/2608.17895#S3)\. We did not collect personal data directly from individuals and did not conduct additional content screening beyond selecting documents from these sources and verifying their licensing status\. Public source documents may nevertheless contain names, affiliations, or other information present in the original records\. We retain source attribution and recommend that users treat the benchmark as a research resource rather than as a source of information about individuals\.

##### Compute infrastructure\.

Open\-weight models were evaluated on an internal server with two NVIDIA H100 GPUs and five NVIDIA L40 GPUs\. Proprietary models were accessed through the OpenRouter API; their underlying hardware configuration and parameter counts are not publicly available for all evaluated models\.

![Refer to caption](https://arxiv.org/html/2608.17895v1/llm_as_judge_prompt_appendix.png)Figure 11:LLM\-as\-a\-judge prompt used to score model responses against the ground truth\.

Similar Articles