DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
Summary
Introduces DrawingVQA, the first benchmark for evaluating multimodal large language models on real-world construction drawings, with 33 drawings and 92 QA pairs across three reasoning depths, revealing a gap between model and expert performance.
View Cached Full Text
Cached at: 07/20/26, 09:21 AM
# DrawingVQA: A Real-World Benchmark for Multi-Depth Visual–Textual Reasoning on Construction Drawings
Source: [https://arxiv.org/html/2607.15418](https://arxiv.org/html/2607.15418)
Yoonhwa Jung Louisiana State University yoonhwa\.jung@lsu\.eduEqual contribution†Work done during graduate studies at the University of Illinois Urbana\-Champaign\.Mani Golparvar\-Fard University of Illinois Urbana\-Champaign mgolpar@illinois\.edu
###### Abstract
We introduceDrawingVQA, the first benchmark designed to evaluate multimodal large language models \(MLLMs\) on real\-world construction drawings—a core media in architecture, civil, and many other engineering practices\. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain\-specific text, forming a uniquely complex visual–textual domain core to engineering workflows\.DrawingVQAbridges this gap with 33 “Issued for Construction” drawings and 92 expertly curated question–answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain\-expert reasoning\. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction\-engineering and four MLLM capability dimensions– the first to explicitly map engineering workflows to AI reasoning competencies\. Evaluations of state\-of\-the\-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths\. This benchmark lays a foundation for domain\-specialized multimodal reasoning to allow for advancement on integration of AI\-driven understanding and real\-world engineering workflows\.DrawingVQAis available here:[https://joonv2\.github\.io/DrawingVQA/](https://joonv2.github.io/DrawingVQA/)
## 1Introduction
Figure 1:The Multimodal Challenges of IFC Drawings andDrawingVQA’s Engineering\-Level VQA\.\(A\) Deconstructing IFC Drawing Complexity:Real\-world “Issued for Construction” \(IFC\) drawings are information\-dense artifacts that present unique multimodal challenges \(See Section[3\.2](https://arxiv.org/html/2607.15418#S3.SS2)\)\.\(B\)DrawingVQAVQA Examples:We present challenging question\-answer pairs fromDrawingVQAthat demand industry\-professional engineering reasoning, categorized by our three Reasoning Depths \(indicated by stars:⋆\\starPerception,⋆⋆\\star\\starContextual,⋆⋆⋆\\star\\star\\starExpert\-level Knowledge and Reasoning\)\. Each example maps to specificMLLM capabilities\(e\.g\., OCR, Visual Perception, Knowledge, Reasoning\) andConstruction Engineering \(CE\) tasks\(e\.g\., QTO, Design Understanding, Element Identification, Compliance\)\. These tasks extend far beyond academic benchmarks \(See Section[3\.4](https://arxiv.org/html/2607.15418#S3.SS4)\)\.The rapid advancement of Multimodal Large Language Models \(MLLMs\) has revealed a significant milestone toward artificial general intelligence \(AGI\), demonstrating remarkable capabilities that often match or even surpass human performance on perception, reasoning, and cross\-modal understanding\[[2](https://arxiv.org/html/2607.15418#bib.bib12)\]\. Recent multimodal benchmarks have shown that frontier models can approach or even surpass human performance in many generalist\[[19](https://arxiv.org/html/2607.15418#bib.bib36),[32](https://arxiv.org/html/2607.15418#bib.bib3),[27](https://arxiv.org/html/2607.15418#bib.bib4),[14](https://arxiv.org/html/2607.15418#bib.bib5)\], academic\[[9](https://arxiv.org/html/2607.15418#biba.bib9),[21](https://arxiv.org/html/2607.15418#bib.bib2),[35](https://arxiv.org/html/2607.15418#bib.bib13),[10](https://arxiv.org/html/2607.15418#bib.bib15)\], or professional fields such as finance\[[9](https://arxiv.org/html/2607.15418#bib.bib6),[22](https://arxiv.org/html/2607.15418#bib.bib7)\], medical\[[1](https://arxiv.org/html/2607.15418#bib.bib8),[11](https://arxiv.org/html/2607.15418#bib.bib9),[17](https://arxiv.org/html/2607.15418#bib.bib10)\], and biology\[[8](https://arxiv.org/html/2607.15418#biba.bib8),[20](https://arxiv.org/html/2607.15418#bib.bib14)\]\. However, deep engineering disciplines, particularly civil and construction engineering, remain largely underexplored\.
Civil and construction engineering are among the most established yet least digitized engineering domains, where professional reasoning heavily depends on construction drawings—dense, multi\-layered visual artifacts that combine geometric entities, annotations, tables, notes, engineering symbols, and cross\-sheet references to communicate design intent and engineering problems\. Existing research in the domain has mostly addressed low\-level perceptual tasks, such as object detection or segmentation of simple floor plans\[[8](https://arxiv.org/html/2607.15418#bib.bib22),[34](https://arxiv.org/html/2607.15418#bib.bib23),[24](https://arxiv.org/html/2607.15418#bib.bib24)\]\. The few existing QA datasets for engineering disciplines tend to focus on academic knowledge, such as “FE exam” style questions\[[1](https://arxiv.org/html/2607.15418#biba.bib1),[30](https://arxiv.org/html/2607.15418#bib.bib17)\]\. This raises a critical question:Do MLLMs that excel at academic exams possess the practical reasoning skills of industry professionals?And more broadly,To what extent do high scores on general benchmarks translate to competence in a complex, high\-stakes domain like construction?
To address this gap, we introduceDrawingVQA, the first benchmark designed to evaluate MLLMs on real\-world construction drawings\. Unlike previous datasets, DrawingVQA uses “Issued for Construction” \(IFC\)–grade drawings, reflecting the standards, complexity, and semantics encountered by practicing civil, structural, and construction engineers\. The dataset challenging multi\-image and interleaved text\-image VQA questions comprises 33 drawings paired with 92 expertly curated question–answer pairs, reflecting the kinds of reasoning tasks engineers perform in real\-world workflows\. Each question is classified into three reasoning levels—easy \(perceptual recognition\), moderate \(contextual interpretation\), and difficult \(expert reasoning\)—capturing how humans progressively reason from perception to professional judgment\.
Beyond conventional accuracy metrics, DrawingVQA introduces a dual\-categorization framework to analyze model performance jointly across seven construction\-engineering dimensions \(e\.g\., Quantity Take\-Off, Design Intent, Code and Specification Compliance\) and four core MLLM capability dimensions \(Visual Perception, Knowledge, Reasoning, and OCR/Text Understanding\)\. This multi\-axis evaluation is the first to explicitly map engineering workflows to AI reasoning competencies, revealing where current models succeed, where they fail, and why\.
It is the first benchmarking to evaluate open\-source and proprietary MLLMs on IFC drawings with engineering professional levels\. We summarize our contributions as follows:
- •We introduceDrawingVQA, the first VQA benchmark built on complex, real\-world “Issued for Construction” \(IFC\) structural drawings, with carefully curated QA pairs reflecting professional reasoning depth\.
- •We propose a dual categorization framework that explicitly links professional engineering workflows to core AI reasoning competencies, enabling multidimensional analysis of model strengths and failures\.
- •We provide a comprehensive analysis of MLLMs, highlighting a critical gap in high\-level engineering reasoning and setting a new, challenging baseline for future research\.
## 2Related Work
MLLM reasoning benchmarksThe evaluation of MLLMs has evolved from general\-perception VQA on natural images \(e\.g\., VQA\[[4](https://arxiv.org/html/2607.15418#bib.bib25)\], GQA\[[12](https://arxiv.org/html/2607.15418#bib.bib26)\], A\-OKVQA\[[27](https://arxiv.org/html/2607.15418#bib.bib4)\]\) to benchmarks testing deep, domain\-specific expertise\. Thisexpert shiftrequires deep disciplinary understanding, by testing college\-level academic knowledge \(e\.g\., MMMU\[[9](https://arxiv.org/html/2607.15418#biba.bib9)\], SceMQA\[[16](https://arxiv.org/html/2607.15418#bib.bib18)\]\) and targeting research\-level reasoning in professional fields such as medicine\[[1](https://arxiv.org/html/2607.15418#bib.bib8),[11](https://arxiv.org/html/2607.15418#bib.bib9),[36](https://arxiv.org/html/2607.15418#bib.bib19)\]and biology\[[8](https://arxiv.org/html/2607.15418#biba.bib8),[20](https://arxiv.org/html/2607.15418#bib.bib14)\]\. While this trend toward deep reasoning is critical, it has not yet addressed the unique, practical reasoning required for engineering, where visual and textual information are deeply intertwined through symbolic diagrams, geometric and spatial relationships, and domain\-specific conventions\.
MLLM in Engineering\.Existing MLLM benchmarks related to engineering predominantly evaluate academic or declarative knowledge, not practical skills\. For example, the Engineering category of MMMU\[[9](https://arxiv.org/html/2607.15418#biba.bib9)\]consists of college\-level exam\-style questions using natural images, tables, or schematic diagrams\. Similarly, R\-Bench\[[10](https://arxiv.org/html/2607.15418#bib.bib15)\]features graduate\-level exam\-style questions in engineering disciplines\. ScienceQA\[[21](https://arxiv.org/html/2607.15418#bib.bib2)\]focuses on scientific experiment questions on natural and general images \(i\.e\., photos\) without domain\-level visual grounding\. As we argue, these datasets, while valuable, do not test comprehensive multimodal reasoning—such as cross\-and multi\-view spatial understanding, detail referencing, or interpreting technical annotations—skills essential for professional engineering tasks\.
Closer to our domain, other benchmarks have partially addressed design\-phase documents\. MM\-Vet\[[32](https://arxiv.org/html/2607.15418#bib.bib3)\]features VQA on simplified floor plans, analyzing OCR, spatial awareness, and math capabilities of MLLMs\. DesignQA\[[2](https://arxiv.org/html/2607.15418#biba.bib2)\]performs VQA over mechanical CAD drawings along with additional PDF documents, while DrafterBench\[[15](https://arxiv.org/html/2607.15418#bib.bib21)\]introduces LLM\-based editing on vectorized CAD data\. While valuable, these datasets focus on schematic, low\-density design representations and fall short of evaluating the multimodal reasoning required for Issued for Construction \(IFC\) drawings—documents that integrate geometric complexity, symbolic notation, tabular data, annotations, and discipline\-specific text\.
Construction domain benchmarksWithin the AEC domain, research has historically bypassed high\-level reasoning on IFC drawings, focusing instead on other data types\. A significant body of work targets low\-level perceptual tasks on schematic floor plans, such as object detection, segmentation, 3D reconstruction\[[8](https://arxiv.org/html/2607.15418#bib.bib22),[34](https://arxiv.org/html/2607.15418#bib.bib23),[24](https://arxiv.org/html/2607.15418#bib.bib24),[23](https://arxiv.org/html/2607.15418#bib.bib29)\]or image captioning\[[13](https://arxiv.org/html/2607.15418#bib.bib28),[6](https://arxiv.org/html/2607.15418#bib.bib27)\]\. Other benchmarks address text\-only question answering based on licensing exams or professional guidelines\[[1](https://arxiv.org/html/2607.15418#biba.bib1),[30](https://arxiv.org/html/2607.15418#bib.bib17)\]\. Such approaches completely bypass the multimodal nature of engineering communication\. Consequently, there remains no benchmark that captures how engineers reason across visual, textual, and spatial modalities in construction professional workflows\.DrawingVQAis the first to address this gap with an expert\-level VQA dataset grounded in real\-world IFC drawings\. Section[3\.5](https://arxiv.org/html/2607.15418#S3.SS5)shows howDrawingVQAdiffer from other standard MLLM benchmarks in engineering disciplines\.
## 3DrawingVQA
### 3\.1Overview ofDrawingVQA
The primary goal ofDrawingVQAis to address the significant gap between existing MLLM benchmarks and the practical, real\-world visual reasoning tasks performed by engineering professionals\. Current benchmarks often focus on academic knowledge or simplified schematics, failing to capture the complexity of industry\-standard documents\.DrawingVQAis designed to identify and measure these performance gaps in frontier models, providing a crucial tool to drive industry adoption\.
To achieve this, our benchmark is built on two core principles:
- •Authentic Data Source:We use real\-world, “Issued for Construction” \(IFC\) structural drawings, which are the dense, multi\-modal, and legally\-binding documents used in the field\.
- •Practical Questioning:The question\-answer \(QA\) pairs are crafted by three domain experts to resemble the day\-to\-day queries and reasoning workflows professionals perform, not just surface\-level academic questions\.
The dataset spans three distinct reasoning depths: \(1\) simple perceptual understanding, \(2\) contextual interpretation, and \(3\) deep domain\-expert reasoning\. To analyze model performance with high granularity, we introduce a novel dual categorization framework that jointly maps each QA pair across seven construction\-engineering dimensions and four foundational MLLM capability dimensions\. This is the first framework, to our knowledge, to explicitly map engineering workflows to core AI reasoning competencies, offering diagnostic insights into why and where models fail in real\-world construction engineering contexts\.
Table 1:DrawingVQAstatisticsFigure 2:Comparison of drawing complexity\. \(a\) A low\-density, schematic architectural floor plan from\[[8](https://arxiv.org/html/2607.15418#bib.bib22)\]\. \(b\) A mechanical design from\[[2](https://arxiv.org/html/2607.15418#biba.bib2)\]\. \(c\) A high\-density, “Issued for Construction” \(IFC\) level drawing from ourDrawingVQAcorpus\.
### 3\.2Construction Drawings \(Issued for Construction\)
The visual foundation ofDrawingVQAconsists of 33 structural drawing sets from six real\-world educational building projects\. Our corpus is composed of high\-entropy “Issued for Construction” \(IFC\) documents\. It is crucial to differentiate the visual and semantic complexity of IFC drawings from the schematic architectural floor plans or mechanical design used in other datasets \(see Figure[2](https://arxiv.org/html/2607.15418#S3.F2)\)\. Compared to other 2D drawings, an IFC drawing spans the entire project lifecycle as “source of truth” that integrates precise geometry, engineering annotations, material specifications, revision histories, and regulatory compliance notes—all of which form a dense visual–textual interface that professionals interpret daily\.
In details, IFC drawings represent a unique multimodal challenge precisely because they are dense, information\-rich artifacts that layer:
- •Complex Geometries:Precise 2D representations \(plans, sections, elevations\) of structural elements \(Figure[1](https://arxiv.org/html/2607.15418#S1.F1)\)\.
- •Symbolic Language:Symbols and domain\-specific vocabulary \(e\.g\., weld symbols, material hatches, discipline specific language and acronyms\)
- •Dense Annotations:Textual and symbolic callouts, dimensions, and annotations with arrows that require high\-fidelity OCR\.
- •Data\-Rich Tables and Notes:Complex tabular schedules for columns, beams, footings, and revisions and general notes\.
- •Cross\-References:Pointers that require a model to align information and understand different dimensional views across different details and multiple sheets with callouts\.
### 3\.3Data Curation Process
Data CollectionWe curated our dataset from a collection of 33 structural drawing sets from 6 campus\-building projects, ensuring each document meets professional Issued for Construction \(IFC\) standards\. The visual artifacts include plans, sections, details, annotations, symbols, and general notes covering diverse structural components \(e\.g\., slabs, columns, beams, and foundations\)\. All drawings used in the dataset conform to a common industry standard for publishing IFC drawings called the U\.S\. National CAD Standard \(NCS\): core elements such as symbols, callouts, sheet IDs, and table formats remain consistent across all construction and facility design projects in the United States\. The NCS standard ensures consistency with professional drafting conventions and serves as the foundational rationale for interpreting symbols, annotations, and callouts within the dataset, aligning our benchmark with established industry documentation practices\. All QA pairs were generated by a team of three domain experts, including licensed civil engineers and construction professionals with 2 to 10\+ years of field experience\. The process was designed to emulate the authentic, information\-seeking behaviors of construction professionals\.
Reasoning DepthExperts were first instructed to formulate questions spanning three progressive reasoning depths\. The first level, Perceptual Understanding, includes simple, direct questions requiring visual search or basic OCR \(e\.g\.,“What is the sheet title?”\)\. The second, Contextual Interpretation, requires models to connect and reason about multiple pieces of information on a single sheet \(e\.g\.,“What is the top of the footing elevation for footing ‘F7x13\.83’ along horizontal grid line 24?”\)\. The final and most complex level, Domain\-Expert Reasoning, demands compositional reasoning, implicit domain knowledge, or aligning information across different drawings \(e\.g\.,“What type of vertical reinforcements would I need for the walls shown? Use the other drawing if needed\.”\)\. All answers were required to be concise, factual, and directly verifiable from the provided drawing set\.
VQA FormatEach question was formulated as either a multiple\-choice or a short open\-answer question\. For the multiple\-choice questions, our experts manually created all distractors to avoid low\-quality, easily guessable, or LLM\-generated options\. Distractors often included realistic but subtly incorrect alternatives such as picking up nearby texts on the drawings that are actually irrelevant or challenging options such as “None of the above” or “All of the above” This expert\-driven design ensures the distractors reflect authentic misinterpretations of IFC drawings, serving as a rigorous proxy for professional ambiguity\. Finally, comparative testing between the MCQ and open\-ended formats confirms that the closed\-ended structure does not artificially reduce task difficulty \(further details are provided in the Supplementary\)\. This multi\-stage, expert\-driven process produced 92 question\-answer pairs grounded in real construction drawings—capturing not only the visual and textual complexity of engineering documents but also the depth of reasoning required in professional practice\.
Quality Control and ValidationTo ensure dataset integrity, we employed a rigorous, multi\-pass validation protocol\. First, each QA pair generated by one expert was independently reviewed by a second expert for correctness, ambiguity, and practical relevance\. Any disagreements were escalated to a senior engineer for final adjudication\. Second, we explicitly filtered out any questions that could be answered without the image or were unanswerable with the provided context\. Finally, after validation, every QA pair was annotated according to our dual categorization framework, which jointly captures \(a\) the construction\-engineering domain aspect of the inquiry and \(b\) the MLLM capability dimension required to solve the question\. These annotations underwent the same multi\-pass review cycle to maintain labeling reliability and conceptual clarity\.
### 3\.4Dual Categorization Framework
The core ofDrawingVQAevaluation comes from our dual categorization framework\. Each QA pair is tagged along two orthogonal axes, allowing for a multifaceted analysis of model capabilities\.
1\. MLLM Capability Dimension \(The “AI Task”\)This axis categorizes the foundational AI competency required to answer the question\. It is divided into four primary dimensions:Visual Perception\(simple identification, attribute recognition, and counting\);Reasoning\(higher\-order cognition, which is further subdivided intoVisual Reasoning,Alignment & Grounding, andSpatial Reasoning\);Knowledge\(accessing implicit, domain\-specific information not explicitly written on the sheet\); andOCR\(text recognition and understanding, insideTabularfor schedules andTextualfor callouts and notes\)\. When a task requires only reading text, its MLLM capability dimension is OCR alone\. If the question necessitates aligning that text from tabular with geometric entities or layouts, it is explicitly mapped to both OCR and Reasoning category \(Visual, Spatial, Alignment\)\.
2\. Construction\-Engineering Dimension \(The “Practical Task”\)This axis categorizes the real\-world engineering workflow the question belongs to\. The seven dimensions are:General / Administrative\(reading metadata, tables, and general notes\);Element Identification\(interpreting structural symbols\);Language Understanding\(interpreting domain\-specific languages\);Dimensional Understanding\(cross\-referencing on multi\-view drawing parts and relationships\);Design Semantic Understanding\(inferring drawing type or design intent\);Quantity Take\-Off \(QTO\)\(tasks related to counting elements\); andCompliance\(verifying components against a code or specification\)\.
We provide a detailed breakdown of the dataset statistics and the distribution of these categories in Table[1](https://arxiv.org/html/2607.15418#S3.T1)\. This dual framework allows us to move beyond a single accuracy score and precisely diagnose why a model fails for engineering workflows, providing a clear roadmap for future improvements\.
Figure 3:Visualization of the dual categorization mapping MLLM Capabilities \(left\) to their corresponding Construction\-Engineering Dimensions \(right\) inDrawingVQA\. The flow diagram highlights how complex engineering tasks, such as “Dimensional Understanding”, draw upon multiple MLLM Reasoning and OCR capabilities\.
Figure 4:Comparison ofDrawingVQAwith other VQA benchmarks, which contain any of Architecture, Civil, Structural, or Construction Engineering disciplines, or drawing\-related questions\. Q\. means Question, Expln\. means Explanations, Profl\. means Professional, Eng\. means Engineering, MC means multiple\-choice, and Open means open\-ended\. DesignQA is not a pure VQA dataset, with document\-reference dependence\. Compared to others,DrawingVQAhas the highest visual complexity \(also featured in Figure[2](https://arxiv.org/html/2607.15418#S3.F2)\) and the highest professional reasoning level in engineering practices\.
### 3\.5Comparisons with Existing Benchmarks
DrawingVQAoccupies a unique position at the intersection of multimodal reasoning and engineering document packed with domain specific knowledge\. As summarized in Figure[4](https://arxiv.org/html/2607.15418#S3.F4), existing benchmarks that include AEC content address only narrow, academic aspects of the discipline\. MMMU\[[9](https://arxiv.org/html/2607.15418#biba.bib9)\]contains a small subset of 135 civil and structural engineering\-related questions framed as textbook problems\. Similarly, RBench\-M\[[10](https://arxiv.org/html/2607.15418#bib.bib15)\]includes 54 questions from structural engineering textbook exams\. ScienceQA\[[21](https://arxiv.org/html/2607.15418#bib.bib2)\]offers 24 relevant engineering practice questions based on natural images, and MM\-Vet\[[32](https://arxiv.org/html/2607.15418#bib.bib3)\]includes only four schematic floor plans for basic spatial reasoning\. These datasets evaluate declarative knowledge rather than assessing context specific multimodal reasoning required in real construction workflows\.
One notable VQA benchark, DesignQA\[[2](https://arxiv.org/html/2607.15418#biba.bib2)\], focuses on the mechanical and industrial engineering domain, which possess parallels with DrawingVQA on 2D drawing and compliance pdf document on compliance check\. However, the QA images utilized were cleaned of any confusing annotations, and other textual or tabular, and symbolic visuals that hinder perception performance\. Additionally, the majority of their QA dataset relies heavily on simple, text\-based measurement checks against PDF documents\. These “hindrances” were kept true to its color in theDrawingVQA, just as anyone sees everyday at real construction projects, requiring authentic, real\-world engineering knowledge rather than basic text extraction\.
DrawingVQAis the first dataset built entirely on real IFC\-level structural drawings, reflecting how professionals interpret, cross\-reference, and reason across sheets in daily practice\. To avoid artificially inflating model performance with repetitive, low\-level queries \(e\.g\., repeatedly extracting sheet IDs\), we curated ahigh\-entropy datasetthat maximizes diagnostic coverage\. Each question targets a specific failure mode and distinct reasoning pathway, ensuring the task demands the multi\-step inferential workflows required at jobsites rather than reducing to simple visual grounding\. Finally, by incorporating a dual categorization framework spanning both construction\-engineering and MLLM capability dimensions, it performs fine\-grained diagnosis of model reasoning gaps, establishing a foundation for advancing AI toward domain\-grounded engineering intelligence\.
## 4Experiments
### 4\.1Baselines
MLLMsTo evaluateDrawingVQA, we benchmarked a range of state\-of\-the\-art and recent MLLMs\. This includes leading proprietary models such as OpenAI’s o3 and GPT\-4o, Google’s Gemini 2\.5 Pro and 2\.5 Flash, and Anthropic’s Claude\-4\.5 Sonnet and open\-source models, including LLaVA\-OneVision\-1\.5\-8B\-Instruct\[[10](https://arxiv.org/html/2607.15418#biba.bib10)\], LLaVA\-v1\.6\-Mistral\-7B\[[5](https://arxiv.org/html/2607.15418#biba.bib5)\], Llama\-3\.2\-11B\-Vision\-Instruct\[[4](https://arxiv.org/html/2607.15418#biba.bib4)\], Qwen3\-VL\-8B\-Instruct\[[11](https://arxiv.org/html/2607.15418#biba.bib11)\], InternVL3\.5\-30B\-A3B\[[12](https://arxiv.org/html/2607.15418#biba.bib12)\], and Phi\-4\-multimodal\-instruct\[[7](https://arxiv.org/html/2607.15418#biba.bib7)\], which are known for their strong performance on general vision\-language tasks\. All models were tested using standard chain\-of\-thought prompting in a zero\-shot setting and answer parsing, following the protocol established in MicroVQA\[[8](https://arxiv.org/html/2607.15418#biba.bib8)\]\. All experiments are conducted with an NVIDIA A100 GPU\.
Human Benchmarking and EvaluationTo benchmark MLLM performance against true domain expertise, we established a robust human baseline\. We randomly selected a subset of20QA pairs from our test set, stratified across the three reasoning depths \(4 Perceptual, 8 Contextual, 8 Expert\)\. 52 Responses were collected from three groups representing increasing levels of domain expertise: \(i\) undergraduate civil engineering students \(not in their freshman year, mainly junior and seniors\), \(ii\) graduate civil engineering students and/or early\-career professionals \(less than three years of industry experience\), and \(iii\) experienced professionals \(three or more years of experience\)\. This allows us to not only compare models to peak human performance but also to understand how model reasoning capabilities align with different stages of professional development\. To maintain a fair experiment, all the QA answering were timed \(20 minutes\) to ensure all participants had equal amount of time to answer the same number of questions\. In addition, we added Random Choice baseline for reference\.
EvaluationWe report overall mean accuracy for both multiple\-choice and open\-ended questions\. In addition, we provide fine\-grained analysis across both MLLM capability dimensions \(visual perception, reasoning, knowledge, OCR\) and construction\-engineering dimensions \(general, element identification, language understanding, dimensional reasoning, semantic understanding, quantity takeoff, code/specification compliance\)\. More details on the dimensions can be found in the Supplementary\. We also evaluate model performance under no\-image conditions and using randomized multiple\-choice label prefixes \(ABCD alphabetical order\), reported in the Supplementary\.
Table 2:Model Performance with Reasoning, MLLM, and Construction Metrics\.MLLM Dimensions\(VP: Visual Perception, K: Knowledge, R: Reasoning, OCR: Optical Character Recognition\);Construction Eng\. Domain Aspects\(Adm\.: Administration, EI: Element Identification, LU: Domain Language Understanding, DR: Dimensional Recognition, SU: Design Semantic Understanding, QTO: Quantity Take Off, Comp\.: Compliance\)\. Underscored values indicate the highest score for each model within the MLLM Dimensions and CE Domain Aspects categories\.Reasoning DepthMLLM DimensionsConstruction Eng\. Domain AspectsModelOverallR1R2R3VPKROCRAdm\.EILUDRSUQTOCompGPT\-4o48\.961\.344\.140\.747\.159\.444\.348\.756\.332\.470\.841\.950\.025\.042\.9o358\.764\.561\.848\.258\.865\.654\.159\.262\.556\.854\.248\.461\.525\.042\.9Gemini\-2\.5\-pro71\.780\.676\.555\.667\.678\.167\.272\.487\.556\.879\.267\.775\.041\.757\.1Gemini\-2\.5\-flash66\.377\.473\.544\.464\.781\.257\.463\.287\.551\.479\.251\.667\.333\.371\.4Claude\-4\.5\-Sonnet57\.664\.564\.740\.758\.859\.457\.954\.168\.848\.766\.741\.957\.741\.757\.1Claude\-4\.5\-Haiku54\.351\.673\.533\.352\.965\.652\.554\.062\.543\.250\.048\.451\.933\.342\.9LLaVA\-OneVision\-1\.5\-8B\-Instruct\[[10](https://arxiv.org/html/2607.15418#biba.bib10)\]40\.261\.341\.214\.838\.246\.931\.139\.543\.824\.362\.525\.842\.30\.00\.0LLaVA\-v1\.6\-Mistral\-7B\[[5](https://arxiv.org/html/2607.15418#biba.bib5)\]28\.325\.829\.429\.629\.437\.527\.923\.731\.224\.316\.732\.330\.825\.014\.3Llama\-3\.2\-11B\-Vision\-Instruct\[[4](https://arxiv.org/html/2607.15418#biba.bib4)\]39\.145\.247\.122\.241\.240\.634\.438\.250\.029\.745\.832\.342\.38\.328\.6Qwen3\-VL\-8B\-Instruct\[[11](https://arxiv.org/html/2607.15418#biba.bib11)\]53\.351\.661\.844\.450\.053\.145\.955\.381\.240\.566\.745\.246\.233\.342\.9InternVL3\.5\-30B\-A3B\[[12](https://arxiv.org/html/2607.15418#biba.bib12)\]41\.351\.635\.337\.026\.543\.842\.642\.150\.040\.545\.845\.146\.18\.314\.3Phi\-4\[[7](https://arxiv.org/html/2607.15418#biba.bib7)\]40\.241\.947\.129\.629\.440\.642\.640\.837\.540\.545\.835\.540\.425\.028\.6Random27\.225\.817\.740\.723\.537\.527\.929\.037\.518\.929\.229\.023\.016\.642\.8Human68\.475\.672\.659\.975\.881\.484\.975\.046\.871\.982\.782\.199\.283\.366\.2\* Undergraduate62\.871\.870\.253\.864\.366\.774\.953\.043\.649\.770\.775\.497\.469\.446\.2\* Graduate & Young Professionals78\.285\.476\.668\.873\.981\.983\.876\.050\.070\.081\.975\.0100\.083\.762\.5\* Professionals94\.990\.085\.093\.389\.195\.696\.096\.046\.796\.095\.696\.0100\.096\.690\.0
### 4\.2Main Results
Overall Performance: SOTA vs\. Human Experts\.A primary finding is the significant performance gap that still exists between most MLLMs and human\-level expertise\. Only one model,Gemini\-2\.5\-pro, achieved an overall score of 71\.7, surpassing the average human score \(68\.4\) and the undergraduate student cohort \(62\.8\)\. However, this still falls substantially short of the performance of experiencedProfessionals\(94\.9\), indicating that while SOTA models can match entry\-level performance, they do not yet possess deep domain expertise\. The next\-best models,Gemini\-2\.5\-flash\(66\.3\) ando3 \(58\.7\), also demonstrate strong capabilities but, like most other models \(e\.g\.,GPT\-4oat 48\.9,Claude\-4\.5\-Sonnetat 57\.6\), still struggle to match the baseline performance of an undergraduate student in civil engineering major\.
Figure 5:Comparison of multimodal reasoning performance across human cohorts and Gemini\-2\.5\-pro onDrawingVQA\. Left: CV\-derived competencies, Right: CE\-derived competencies\.#### Reasoning Depth Analysis\.
Our Reasoning Depth category evaluates a model’s ability to progress from basic perception to deep expertise\.R1 \(Visual Perception\)tests foundational skills, such as recognizing elements, where models likeGemini\-2\.5\-pro\(80\.6\) andGemini\-2\.5\-flash\(77\.4\) score highest\. This is followed byR2 \(Contextual Understanding\), which requires connecting multiple pieces of information\. Performance here shows a mixed pattern: while top\-performing models \(Gemini\-2\.5\-pro, dropping to 76\.5\) and all human cohorts show a slight decrease, indicating the inherent difficulty of synthesis, several other models \(e\.g\.,Qwen3\-VL\-8Bfrom 51\.6 to 61\.8\) actually improve, suggesting they gain better insights as more context is provided\. Finally,R3 \(Expert Reasoning\), which demands deep domain\-specific inference, is clearly the most significant bottleneck\. Scores for all models fall dramatically\. This contrasts sharply with humanProfessionals, whose R3 score \(93\.3\) is their highest, highlighting the primary gap between current AI and true domain expertise\.
General MLLM aspects\.Analyzing the “MLLM Dimensions” reveals a distinctspecialistprofile for models compared to humans\. Models consistently excel in theKnowledge \(K\)dimension, withGemini\-2\.5\-flash\(81\.2\) andGemini\-2\.5\-pro\(78\.1\) achieving the highest scores\. This suggests models are highly effective at retrieving and applying their vast store of accumulated knowledge\. Conversely, humans demonstrate a clear advantage in theReasoning \(R\)dimension, with the average human scoring 84\.9 and professionals reaching 96\.0\. As illustrated in Figure[5](https://arxiv.org/html/2607.15418#S4.F5), even undergraduate students are better than the state\-of\-the\-art MLLM on reasoning capability\. It highlights that humans retain a superior ability to perform flexible and abstract reasoning on complex visual information\.
Fine\-Grained Construction Engineering Domain Aspects\.Unlike the well\-rounded human experts, Figure[5](https://arxiv.org/html/2607.15418#S4.F5)reveals that MMLM models exhibit a different performance pattern \(i\.e\., “spiky” specialist profile\)\. MMLM strengths lie inAdministration \(Adm\.\)andLanguage Understanding \(LU\), withGemini\-2\.5\-proachieving 87\.5\. This indicates a strong grasp of domain\-specific terminology and the ability to process administrative information from documents\. However, the most significant weakness for all MLLMs isQuantity Take Off \(QTO\)\. With the top model \(Gemini\-2\.5\-pro\) scoring only 41\.7 and several models near 0\.0, this task remains a major challenge\. This difficulty is likely attributable to MLLMs’ limitations in precise spatial grounding and the inability to accurately decompose and count objects in dense, large\-scale engineering drawings\. Domain\-specific tasks such as quantity takeoff require multiple MLLM competencies simultaneously, including visual reasoning, OCR, and embedded domain knowledge, as illustrated in Figure[3](https://arxiv.org/html/2607.15418#S3.F3)\. Among open\-source models,Qwen3\-VL\-8B\-Instructshows a very strong and competitive performance at 53\.3, notably outperforming several proprietary models and showing exceptional strength inGeneral Administrationaspect \(81\.2\)\. More details are analyzed in Supplementary\.
Human EvaluationThe “Construction Eng\. Domain Aspects” category provides the most detailed insights\. Human performance clearly illustrates the value of experience: scores show a progressive trend fromUndergraduate\(62\.8\) toGraduate & Young Professionals\(78\.2\) toProfessionals\(94\.9\)\. This trend is especially pronounced in theCompliance \(Comp\)task, where the professional score \(90\.0\) is nearly double that of undergraduates \(46\.2\)\. This suggests that understanding complex specifications and requirements is a skill acquired through practical experience beyond academic achievement\. In contrast,Semantic Understanding \(SU\)appears to be a more fundamental skill, with all human cohorts performing exceptionally well \(97\.4\-100\.0\)\.
Content Realism\.We also asked participants to rate question realism on a 1\-5 scale \(5=very realistic\)\. The feedback confirmed our dataset’s authenticity: a combined71\.2%of participants rated the questions as highly realistic \(32\.7% rated ’4’; 38\.5% rated ’5’\)\. Critically, 0% of participants gave a ’1’ or ’2’ rating\. This feedback validatesDrawingVQAas an authentic benchmark for real\-world AEC challenges\.
Validating Multimodal Dependency \(Text\-Only Ablation\)To validate thatDrawingVQArequires genuine multimodal reasoning and is not solvable through linguistic priors or context leakage, we performed a text\-only ablation study\. As shown in Table[3](https://arxiv.org/html/2607.15418#S4.T3), this ablation is expected to cause a significant performance collapse, dropping scores to near or even lower than the random\-guess baseline\. This confirms that the visual information in the drawings is not just supplementary butindispensablefor solving the tasks, validating our dataset’s reliance on grounded visual understanding\.
Table 3:Text\-Only Ablation Study
## 5Error Analysis
Overall\.As part of the analysis to understand why models made certain decisions for questions that they got wrong, their explanation or “chain of thoughts” were investigated to break apart the steps it took, and where it may have gone wrong\. A more detailed breakdown of these analysis will be provided in the Supplementary material\.
Visual Grounding and Perception Failures\.A common failure mode occurs when the model’s reasoning plan is logically sound, but its visual perception fails\. This was most pronounced inQTO \(Quantity Take Off\)tasks\. For instance, when asked to count columns along a specific grid line \(e\.g\., ”grid line Q”\), the model’s CoT would correctly state the plan:“Step 1: Locate Vertical Grid Line Q… Step 3: Trace Grid Line Q… Step 4: Identify and Count the Columns”\. However, the final count would be incorrect because the model visually “derailed” while tracing, misidentifying the line or “hallucinating” intersections that were not present\.
This decoupling of logic and perception was also seen in detail\-finding tasks\. When asked to find a callout for a footingnext toanother, the model’s CoT would flawlessly describe finding a callout symbol \(e\.g\., “9 / S3\.00”\), but it would be thewrongcallout, having visually latched onto a different, nearby footing\. This shows a fundamental inability to maintain precise spatial grounding in dense, symbol\-rich IFC drawings illustrated in Figure[1](https://arxiv.org/html/2607.15418#S1.F1)\.
Failure to Interpret Core Symbolic Language for Cross\-Referencing\.The most consequential breakdowns occur in tasks that require multi\-step visual reasoning—particularly those involving callouts, cross\-sheet links, or geometric dependencies\. These tasks form the backbone of dimensional and compliance verification\. When a model must search across multiple details for a governing reference, it often stops after identifying the first plausible link and fails to perform a complete visual scan\. Once anchored to that partial discovery, the model shows a tendency to force\-fit its reasoning to the available answer choices rather than reassessing the drawing\. For engineering workflows—where reliability depends on accurately tracing references, elevations, sections, governing geometry, or scope boundaries—this behavior makes current models fundamentally untrustworthy for tasks requiring disciplined cross\-referencing or multi\-step interpretation\.
Lack of Expert Knowledge\.A recurring source of failure is the MLLM’s limited domain knowledge in core construction and engineering disciplines\. Models frequently misinterpret basic drafting conventions—such as weld symbols, standard structural symbols, and annotations with domain acronyms—because they lack the expert familiarity required to decode these symbols reliably\. When MLLM cannot recognize what a triangular weld symbol means, or cannot distinguish a footing from other structural elements, it begins reasoning from a fundamentally incorrect premise\. This knowledge gap does not merely produce isolated mistakes; it accelerates cascading errors, where misidentified elements lead to incorrect dimensional logic, false assumptions about load paths, or fabricated interpretations of the drawing\. Without this baseline professional literacy, the model cannot properly read or interpret construction documents, and its reasoning degrades before the analytical process even begins\.
## 6Conclusion
We introduceDrawingVQA, the first VQA benchmark built on complex, real\-world IFC drawings, designed to assess the multimodal reasoning capabilities of MLLMs on tasks that reflect authentic engineering practice\. We employ a dual categorization analysis that spans both MLLM capabilities and construction\-engineering aspects\. We hope this benchmark provides a foundation for evaluating AI systems in expert\-driven domains\. Future direction includes exploration of domain\-specific steering, the role of model scale, the text and vision embedding capacity, and expansion into Civil, and Mechanical, Electrical, and Plumbing disciplines\.
## References
- \[1\]A\. B\. Abacha, S\. A\. Hasan, V\. V\. Datla, J\. Liu, D\. Demner\-Fushman, and H\. Müller\(2019\)VQA\-med: overview of the medical visual question answering task at imageclef 2019\.\.CLEF \(working notes\)2\(6\),pp\. 1–11\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[2\]J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.\(2023\)Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[3\]X\. An, Y\. Xie, K\. Yang, W\. Zhang, X\. Zhao, Z\. Cheng, Y\. Wang, S\. Xu, C\. Chen, C\. Wu, H\. Tan, C\. Li, J\. Yang, J\. Yu, X\. Wang, B\. Qin, Y\. Wang, Z\. Yan, Z\. Feng, Z\. Liu, B\. Li, and J\. Deng\(2025\)LLaVA\-onevision\-1\.5: fully open framework for democratized multimodal training\.External Links:2509\.23661,[Link](https://arxiv.org/abs/2509.23661)Cited by:[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.9.9.1),[Table 11](https://arxiv.org/html/2607.15418#A6.T11.8.8.8.8.8.2),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.8.7.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.15418#S4.T2.6.1.1.1.9.9.1)\.
- \[4\]S\. Antol, A\. Agrawal, J\. Lu, M\. Mitchell, D\. Batra, C\. L\. Zitnick, and D\. Parikh\(2015\)Vqa: visual question answering\.pp\. 2425–2433\.Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[5\]J\. Burgess, J\. J\. Nirschl, L\. Bravo\-Sánchez, A\. Lozano, S\. R\. Gupte, J\. G\. Galaz\-Montoya, Y\. Zhang, Y\. Su, D\. Bhowmik, Z\. Coman, S\. M\. Hasan, A\. Johannesson, W\. D\. Leineweber, M\. G\. Nair, R\. Yarlagadda, C\. Zuraski, W\. Chiu, S\. Cohen, J\. N\. Hansen, M\. D\. Leonetti, C\. Liu, E\. Lundberg, and S\. Yeung\-Levy\(2025\-06\)MicroVQA: a multimodal reasoning benchmark for microscopy\-based scientific research\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 19552–19564\.Cited by:[§F\.5](https://arxiv.org/html/2607.15418#A6.SS5.p1.1),[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p1.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1)\.
- \[6\]X\. Chen and Z\. Zou\(2025\)Are large pre\-trained vision language models effective construction safety inspectors?\.External Links:2508\.11011,[Link](https://arxiv.org/abs/2508.11011)Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[7\]A\. C\. Doris, D\. Grandi, R\. Tomich, M\. F\. Alam, M\. Ataei, H\. Cheong, and F\. Ahmed\(2024\)DesignQA: a multimodal benchmark for evaluating large language models’ understanding of engineering documentation\.External Links:2404\.07917,[Link](https://arxiv.org/abs/2404.07917)Cited by:[1st item](https://arxiv.org/html/2607.15418#A3.I1.i1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p3.1),[Figure 2](https://arxiv.org/html/2607.15418#S3.F2),[Figure 2](https://arxiv.org/html/2607.15418#S3.F2.4.2),[Figure 4](https://arxiv.org/html/2607.15418#S3.F4.1.1.m1.2.2.2.2.2.2.2.2.2.1),[§3\.5](https://arxiv.org/html/2607.15418#S3.SS5.p2.1)\.
- \[8\]Z\. Fan, L\. Zhu, H\. Li, X\. Chen, S\. Zhu, and P\. Tan\(2021\)Floorplancad: a large\-scale cad drawing dataset for panoptic symbol spotting\.pp\. 10128–10137\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p2.1),[§2](https://arxiv.org/html/2607.15418#S2.p4.1),[Figure 2](https://arxiv.org/html/2607.15418#S3.F2),[Figure 2](https://arxiv.org/html/2607.15418#S3.F2.4.2)\.
- \[9\]Z\. Gan, D\. Zhang, H\. Li, Y\. Wu, X\. Lin, J\. Liu, H\. Wu, C\. Fu, Z\. Xu, R\. Zhang,et al\.\(2025\)Mme\-finance: a multimodal finance benchmark for expert\-level understanding and reasoning\.InProceedings of the 33rd ACM International Conference on Multimedia,pp\. 12867–12874\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[10\]M\. Guo, J\. Xu, Y\. Zhang, J\. Song, H\. Peng, Y\. Deng, X\. Dong, K\. Nakayama, Z\. Geng, C\. Wang, B\. Ni, G\. Yang, Y\. Rao, H\. Peng, H\. Hu, G\. Wetzstein, and S\. Hu\(2025\-13–19 Jul\)RBench: graduate\-level multi\-disciplinary benchmarks for LLM &; MLLM complex reasoning evaluation\.InProceedings of the 42nd International Conference on Machine LearningProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)Proceedings of the IEEE/CVF international conference on computer visionProceedings of the IEEE/CVF International Conference on Computer VisionProceedings of the IEEE/CVF conference on computer vision and pattern recognitionProceedings of the IEEE international conference on computer visionProceedings of the IEEE/CVF conference on computer vision and pattern recognitionThe Thirty\-ninth Annual Conference on Neural Information Processing SystemsComputer Vision – ECCV 2024,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, J\. Zhu, L\. Ku, A\. Martins, V\. Srikumar, A\. Leonardis, E\. Ricci, S\. Roth, O\. Russakovsky, T\. Sattler, and G\. Varol \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 21184–21201\.External Links:[Link](https://proceedings.mlr.press/v267/guo25r.html)Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p2.1),[Figure 4](https://arxiv.org/html/2607.15418#S3.F4.1.1.m1.2.2.2.2.2.2.2.2.7.4.1),[§3\.5](https://arxiv.org/html/2607.15418#S3.SS5.p1.1)\.
- \[11\]Y\. Hu, T\. Li, Q\. Lu, W\. Shao, J\. He, Y\. Qiao, and P\. Luo\(2024\)Omnimedvqa: a new large\-scale comprehensive evaluation benchmark for medical lvlm\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 22170–22183\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[12\]D\. A\. Hudson and C\. D\. Manning\(2019\)Gqa: a new dataset for real\-world visual reasoning and compositional question answering\.pp\. 6700–6709\.Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[13\]Y\. Jung, I\. Cho, S\. Hsu, and M\. Golparvar\-Fard\(2024\)VisualSiteDiary: a detector\-free vision\-language transformer model for captioning photologs for daily construction reporting and image retrievals\.Automation in Construction165,pp\. 105483\.Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[14\]B\. Li, Y\. Ge, Y\. Ge, G\. Wang, R\. Wang, R\. Zhang, and Y\. Shan\(2024\)Seed\-bench: benchmarking multimodal large language models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13299–13308\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[15\]Y\. Li, Z\. Dong, and Y\. Shao\(2025\)DrafterBench: benchmarking large language models for tasks automation in civil engineering\.External Links:2507\.11527,[Link](https://arxiv.org/abs/2507.11527)Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p3.1)\.
- \[16\]Z\. Liang, K\. Guo, G\. Liu, T\. Guo, Y\. Zhou, T\. Yang, J\. Jiao, R\. Pi, J\. Zhang, and X\. Zhang\(2024\-08\)SceMQA: a scientific college entrance level multimodal question answering benchmark\.Bangkok, Thailand,pp\. 109–119\.External Links:[Link](https://aclanthology.org/2024.acl-short.11/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-short.11)Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[17\]B\. Liu, K\. Zou, L\. Zhan, Z\. Lu, X\. Dong, Y\. Chen, C\. Xie, J\. Cao, X\. Wu, and H\. Fu\(2025\)Gemex: a large\-scale, groundable, and explainable medical vqa benchmark for chest x\-ray diagnosis\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 21310–21320\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[18\]H\. Liu, C\. Li, Y\. Li, and Y\. J\. Lee\(2023\)Improved baselines with visual instruction tuning\.External Links:2310\.03744Cited by:[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.10.10.1),[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.11.11.1),[Table 10](https://arxiv.org/html/2607.15418#A6.T10.4.2.14.12.1),[Table 11](https://arxiv.org/html/2607.15418#A6.T11.9.9.9.9.9.2),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.9.8.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.15418#S4.T2.6.1.1.1.10.10.1)\.
- \[19\]Y\. Liu, H\. Duan, Y\. Zhang, B\. Li, S\. Zhang, W\. Zhao, Y\. Yuan, J\. Wang, C\. He, Z\. Liu, K\. Chen, and D\. Lin\(2025\)MMBench: is your multi\-modal model an all\-around player?\.Cham,pp\. 216–233\.External Links:ISBN 978\-3\-031\-72658\-3Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[20\]A\. Lozano, J\. Nirschl, J\. Burgess, S\. R\. Gupte, Y\. Zhang, A\. Unell, and S\. Yeung\(2024\)Micro\-bench: a microscopy benchmark for vision\-language understanding\.Advances in Neural Information Processing Systems37,pp\. 30670–30685\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[21\]P\. Lu, S\. Mishra, T\. Xia, L\. Qiu, K\. Chang, S\. Zhu, O\. Tafjord, P\. Clark, and A\. Kalyan\(2022\)Learn to explain: multimodal reasoning via thought chains for science question answering\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 2507–2521\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/11332b6b6cf4485b84afadb1352d3a9a-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p2.1),[Figure 4](https://arxiv.org/html/2607.15418#S3.F4.1.1.m1.2.2.2.2.2.2.2.2.5.2.1),[§3\.5](https://arxiv.org/html/2607.15418#S3.SS5.p1.1)\.
- \[22\]J\. Luo, Z\. Kou, L\. Yang, X\. Luo, J\. Huang, Z\. Xiao, J\. Peng, C\. Liu, J\. Ji, X\. Liu, S\. Han, M\. Zhang, and Y\. Guo\(2025\-07\)FinMME: benchmark dataset for financial multi\-modal reasoning evaluation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 29465–29489\.External Links:[Link](https://aclanthology.org/2025.acl-long.1426/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1426),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[23\]R\. Luo, Z\. Liu, T\. Cheng, J\. Wang, T\. Wang, F\. Cheng, F\. Chai, Y\. Li, X\. Wei, H\. Wang,et al\.ArchCAD\-400k: a large\-scale cad drawings dataset and new baseline for panoptic symbol spotting\.Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[24\]X\. Lv, S\. Zhao, X\. Yu, and B\. Zhao\(2021\)Residential floor plan recognition and reconstruction\.pp\. 16717–16726\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p2.1),[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[25\]Meta\(2024\)Llama\-3\.2\-11b\-vision\-instruct\.Technical reportExternal Links:[Link](https://huggingface.co/meta-llama/%20Llama-3.2-11B-Vision-Instruct)Cited by:[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.12.12.1),[Table 10](https://arxiv.org/html/2607.15418#A6.T10.4.2.11.9.1),[Table 11](https://arxiv.org/html/2607.15418#A6.T11.10.10.10.10.10.2),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.10.9.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.15418#S4.T2.6.1.1.1.11.11.1)\.
- \[26\]Microsoft, :, A\. Abouelenin, A\. Ashfaq, A\. Atkinson, H\. Awadalla, N\. Bach, J\. Bao, A\. Benhaim, M\. Cai, V\. Chaudhary, C\. Chen, D\. Chen, D\. Chen, J\. Chen, W\. Chen, Y\. Chen, Y\. Chen, Q\. Dai, X\. Dai, R\. Fan, M\. Gao, M\. Gao, A\. Garg, A\. Goswami, J\. Hao, A\. Hendy, Y\. Hu, X\. Jin, M\. Khademi, D\. Kim, Y\. J\. Kim, G\. Lee, J\. Li, Y\. Li, C\. Liang, X\. Lin, Z\. Lin, M\. Liu, Y\. Liu, G\. Lopez, C\. Luo, P\. Madan, V\. Mazalov, A\. Mitra, A\. Mousavi, A\. Nguyen, J\. Pan, D\. Perez\-Becker, J\. Platin, T\. Portet, K\. Qiu, B\. Ren, L\. Ren, S\. Roy, N\. Shang, Y\. Shen, S\. Singhal, S\. Som, X\. Song, T\. Sych, P\. Vaddamanu, S\. Wang, Y\. Wang, Z\. Wang, H\. Wu, H\. Xu, W\. Xu, Y\. Yang, Z\. Yang, D\. Yu, I\. Zabir, J\. Zhang, L\. L\. Zhang, Y\. Zhang, and X\. Zhou\(2025\)Phi\-4\-mini technical report: compact yet powerful multimodal language models via mixture\-of\-loras\.External Links:2503\.01743,[Link](https://arxiv.org/abs/2503.01743)Cited by:[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.20.20.1),[Table 11](https://arxiv.org/html/2607.15418#A6.T11.13.13.13.13.13.2),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.15.14.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.15418#S4.T2.6.1.1.1.14.14.1)\.
- \[27\]D\. Schwenk, A\. Khandelwal, C\. Clark, K\. Marino, and R\. Mottaghi\(2022\)A\-okvqa: a benchmark for visual question answering using world knowledge\.InEuropean conference on computer vision,pp\. 146–162\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
- \[28\]Q\. Team\(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.14.14.1),[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.15.15.1),[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.16.16.1),[Table 10](https://arxiv.org/html/2607.15418#A6.T10.4.2.3.1.1),[Table 11](https://arxiv.org/html/2607.15418#A6.T11.11.11.11.11.11.2),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.11.10.1),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.12.11.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.15418#S4.T2.6.1.1.1.12.12.1)\.
- \[29\]W\. Wang, Z\. Gao, L\. Gu, H\. Pu, L\. Cui, X\. Wei, Z\. Liu, L\. Jing, S\. Ye, J\. Shao,et al\.\(2025\)InternVL3\.5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv preprint arXiv:2508\.18265\.Cited by:[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.17.17.1),[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.18.18.1),[Table 4](https://arxiv.org/html/2607.15418#A1.T4.6.1.1.1.19.19.1),[Table 10](https://arxiv.org/html/2607.15418#A6.T10.4.2.7.5.1),[Table 11](https://arxiv.org/html/2607.15418#A6.T11.12.12.12.12.12.2),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.13.12.1),[Table 14](https://arxiv.org/html/2607.15418#A6.T14.1.1.1.1.14.13.1),[§4\.1](https://arxiv.org/html/2607.15418#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2607.15418#S4.T2.6.1.1.1.13.13.1)\.
- \[30\]Y\. Wu, L\. Wang, and R\. Liu\(2025\)CEQuest: benchmarking large language models for construction estimation\.arXiv preprint arXiv:2508\.16081\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p2.1),[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[31\]R\. Xiong, Y\. Wang, S\. Gunhan, Y\. Zhu, and C\. Berryman\(2025\)Can ai master construction management \(cm\)? benchmarking state\-of\-the\-art large language models on cm certification exams\.External Links:2504\.08779,[Link](https://arxiv.org/abs/2504.08779)Cited by:[1st item](https://arxiv.org/html/2607.15418#A1.I1.i1.p1.1),[§1](https://arxiv.org/html/2607.15418#S1.p2.1),[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[32\]W\. Yu, Z\. Yang, L\. Li, J\. Wang, K\. Lin, Z\. Liu, X\. Wang, and L\. Wang\(2024\)MM\-vet: evaluating large multimodal models for integrated capabilities\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p3.1),[Figure 4](https://arxiv.org/html/2607.15418#S3.F4.1.1.m1.2.2.2.2.2.2.2.2.6.3.1),[§3\.5](https://arxiv.org/html/2607.15418#S3.SS5.p1.1)\.
- \[33\]X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng, R\. Liu, G\. Zhang, S\. Stevens, D\. Jiang, W\. Ren, Y\. Sun, C\. Wei, B\. Yu, R\. Yuan, R\. Sun, M\. Yin, B\. Zheng, Z\. Yang, Y\. Liu, W\. Huang, H\. Sun, Y\. Su, and W\. Chen\(2024\-06\)MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert agi\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 9556–9567\.Cited by:[§F\.5](https://arxiv.org/html/2607.15418#A6.SS5.p1.1),[§1](https://arxiv.org/html/2607.15418#S1.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p1.1),[§2](https://arxiv.org/html/2607.15418#S2.p2.1),[Figure 4](https://arxiv.org/html/2607.15418#S3.F4.1.1.m1.1.1.1.1.1.1.1.1.1.2),[§3\.5](https://arxiv.org/html/2607.15418#S3.SS5.p1.1)\.
- \[34\]Z\. Zeng, X\. Li, Y\. K\. Yu, and C\. Fu\(2019\)Deep floor plan recognition using a multi\-task network with room\-boundary\-guided attention\.pp\. 9096–9104\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p2.1),[§2](https://arxiv.org/html/2607.15418#S2.p4.1)\.
- \[35\]W\. Zhang, M\. Aljunied, C\. Gao, Y\. K\. Chia, and L\. Bing\(2023\)M3exam: a multilingual, multimodal, multilevel benchmark for examining large language models\.Advances in Neural Information Processing Systems36,pp\. 5484–5505\.Cited by:[§1](https://arxiv.org/html/2607.15418#S1.p1.1)\.
- \[36\]X\. Zhang, C\. Wu, Z\. Zhao, W\. Lin, Y\. Zhang, Y\. Wang, and W\. Xie\(2024\)PMC\-vqa: visual instruction tuning for medical visual question answering\.External Links:2305\.10415,[Link](https://arxiv.org/abs/2305.10415)Cited by:[§2](https://arxiv.org/html/2607.15418#S2.p1.1)\.
DrawingVQA: A Real\-World Benchmark for Multi\-Depth Visual–Textual Reasoning on Construction Drawings Supplementary Material
###### Contents
1. [1Introduction](https://arxiv.org/html/2607.15418#S1)
2. [2Related Work](https://arxiv.org/html/2607.15418#S2)
3. [3DrawingVQA](https://arxiv.org/html/2607.15418#S3)1. [3\.1Overview ofDrawingVQA](https://arxiv.org/html/2607.15418#S3.SS1) 2. [3\.2Construction Drawings \(Issued for Construction\)](https://arxiv.org/html/2607.15418#S3.SS2) 3. [3\.3Data Curation Process](https://arxiv.org/html/2607.15418#S3.SS3) 4. [3\.4Dual Categorization Framework](https://arxiv.org/html/2607.15418#S3.SS4) 5. [3\.5Comparisons with Existing Benchmarks](https://arxiv.org/html/2607.15418#S3.SS5)
4. [4Experiments](https://arxiv.org/html/2607.15418#S4)1. [4\.1Baselines](https://arxiv.org/html/2607.15418#S4.SS1) 2. [4\.2Main Results](https://arxiv.org/html/2607.15418#S4.SS2)
5. [5Error Analysis](https://arxiv.org/html/2607.15418#S5)
6. [6Conclusion](https://arxiv.org/html/2607.15418#S6)
7. [References](https://arxiv.org/html/2607.15418#bib)
8. [AFull Main Results](https://arxiv.org/html/2607.15418#A1)
9. [BAuthor contributions](https://arxiv.org/html/2607.15418#A2)
10. [CLimitations and future work](https://arxiv.org/html/2607.15418#A3)
11. [DBenchmark details](https://arxiv.org/html/2607.15418#A4)1. [D\.1AccessingDrawingVQA](https://arxiv.org/html/2607.15418#A4.SS1) 2. [D\.2General Guidelines](https://arxiv.org/html/2607.15418#A4.SS2) 3. [D\.3Question curation guidelines](https://arxiv.org/html/2607.15418#A4.SS3) 4. [D\.4Dataset VQA Structure](https://arxiv.org/html/2607.15418#A4.SS4) 5. [D\.5Training contamination mitigation](https://arxiv.org/html/2607.15418#A4.SS5) 6. [D\.6Question Format: Multiple\-Choice vs\. Open\-Ended](https://arxiv.org/html/2607.15418#A4.SS6) 7. [D\.7Generalization to Unseen Drawing Styles](https://arxiv.org/html/2607.15418#A4.SS7) 8. [D\.8Ethical Consideration](https://arxiv.org/html/2607.15418#A4.SS8)
12. [EDual Categorization Taxonomy](https://arxiv.org/html/2607.15418#A5)
13. [FExperiments details](https://arxiv.org/html/2607.15418#A6)1. [F\.1Evaluation Prompts and Parsing](https://arxiv.org/html/2607.15418#A6.SS1) 2. [F\.2Model details](https://arxiv.org/html/2607.15418#A6.SS2) 3. [F\.3Human baseline](https://arxiv.org/html/2607.15418#A6.SS3) 4. [F\.4Dense and MoE models](https://arxiv.org/html/2607.15418#A6.SS4) 5. [F\.5No\-Image Ablations](https://arxiv.org/html/2607.15418#A6.SS5) 6. [F\.6Breakdown results on Dual Category Mapping](https://arxiv.org/html/2607.15418#A6.SS6) 7. [F\.7More experiments on Image resolution](https://arxiv.org/html/2607.15418#A6.SS7) 8. [F\.8More experiments on on Option Permutation and Prefix Sensitivity](https://arxiv.org/html/2607.15418#A6.SS8) 9. [F\.9Ablation Study on Input Modality: The PDF Hybrid Setting](https://arxiv.org/html/2607.15418#A6.SS9)
14. [GError Analysis](https://arxiv.org/html/2607.15418#A7)1. [G\.1Visual Perception errors](https://arxiv.org/html/2607.15418#A7.SS1) 2. [G\.2Visual and Text Perception errors](https://arxiv.org/html/2607.15418#A7.SS2) 3. [G\.3Text Perception \(OCR\) errors](https://arxiv.org/html/2607.15418#A7.SS3) 4. [G\.4Knowledge errors](https://arxiv.org/html/2607.15418#A7.SS4) 5. [G\.5Reasoning errors](https://arxiv.org/html/2607.15418#A7.SS5) 6. [G\.6Instruction Adherence and Output Formatting](https://arxiv.org/html/2607.15418#A7.SS6) 7. [G\.7Parametric Bias vs\. Visual Grounding](https://arxiv.org/html/2607.15418#A7.SS7) 8. [G\.8Visual Attention](https://arxiv.org/html/2607.15418#A7.SS8)
15. [Supplementary References](https://arxiv.org/html/2607.15418#biba)
## Appendix AFull Main Results
Table[4](https://arxiv.org/html/2607.15418#A1.T4)presents the comprehensive evaluation of all models across three reasoning depths \(R1–R3\), four MLLM dimensions, and seven construction\-specific domains\.
- •Gaps Between Theory and Practice: Previous studies have demonstrated that models like GPT\-4o pass construction certification exams with nearly 90% accuracy\[[1](https://arxiv.org/html/2607.15418#biba.bib1)\]\. However, our results show a significant performance drop when these models are tasked with deciphering engineering drawings\. For instance, GPT\-4o achieves 48\.9% onDrawingVQA\. This highlights that while current MLLMs possess strong domain knowledge retrieval capabilities, they fundamentally lack the spatial and semantic visual reasoning required to execute real\-world engineering workflows\. This discrepancy aligns common notion that MLLMs excel atKnowledge\(their strongest dimension\), rather than the visual reasoning required to execute real\-world engineering workflows\.
- •Multimodal Architecture: Among the evaluated architectures, Gemini and Qwen family achieve state\-of\-the\-art performance\. This suggests that the native multimodal architecture of the family may have better visual grounding and model architecture for dense technical documents compared to others\. Especially,Gemini\-3\-pro\-preview111This was released on 11/19/2025 at the time of writing supplementary materials\.shows improved knowledge and OCR capability in their modality, leading to the best result \(77\.2%\), which now moves much closer to the graduate and young professionals’ benchmark score\.
- •Benchmarking Against Human Expertise: Most MLLMs significantly lag behind Graduate & Young professionals \(78\.2%\) and the Professional Expert baseline \(94\.9%\)\. This gap between the best AI and industry experts indicates that while MLLMs are approaching the competency of an entry\-level engineer, they are not yet reliable enough for autonomous professional practice\.
Detailed model size impacts and qualitative error analysis are provided in the subsequent sections\.
Table 4:Full Main Results\.MLLM Dimensions\(VP: Visual Perception, K: Knowledge, R: Reasoning, OCR: Optical Character Recognition\);Construction Eng\. Domain Aspects\(Adm\.: Administration, EI: Element Identification, LU: Domain Language Understanding, DR: Dimensional Recognition, SU: Design Semantic Understanding, QTO: Quantity Take Off, Comp\.: Compliance\)\. Underscored values indicate the highest score for each model within the MLLM Dimensions and CE Domain Aspects categories\.Reasoning DepthMLLM DimensionsConstruction Eng\. Domain AspectsModelOverallR1R2R3VPKROCRAdm\.EILUDRSUQTOCompGPT\-4o48\.961\.344\.140\.747\.159\.444\.348\.756\.332\.470\.841\.950\.025\.042\.9o358\.764\.561\.848\.258\.865\.654\.159\.262\.556\.854\.248\.461\.525\.042\.9Gemini\-2\.5\-pro71\.780\.676\.555\.667\.678\.167\.272\.487\.556\.879\.267\.775\.041\.757\.1Gemini\-2\.5\-flash66\.377\.473\.544\.464\.781\.257\.463\.287\.551\.479\.251\.667\.333\.371\.4Claude\-4\.5\-Sonnet57\.664\.564\.740\.758\.859\.457\.954\.168\.848\.766\.741\.957\.741\.757\.1Claude\-4\.5\-Haiku54\.351\.673\.533\.352\.965\.652\.554\.062\.543\.250\.048\.451\.933\.342\.9LLaVA\-OneVision\-1\.5\-8B\-Instruct\[[10](https://arxiv.org/html/2607.15418#biba.bib10)\]40\.261\.341\.214\.838\.246\.931\.139\.543\.824\.362\.525\.842\.30\.00\.0LLaVA\-v1\.6\-Mistral\-7B\[[5](https://arxiv.org/html/2607.15418#biba.bib5)\]28\.325\.829\.429\.629\.437\.527\.923\.731\.224\.316\.732\.330\.825\.014\.3LLaVA\-v1\.6\-34B\[[5](https://arxiv.org/html/2607.15418#biba.bib5)\]42\.438\.744\.144\.450\.053\.142\.640\.843\.837\.841\.735\.540\.433\.342\.9Llama\-3\.2\-11B\-Vision\-Instruct\[[4](https://arxiv.org/html/2607.15418#biba.bib4)\]39\.145\.247\.122\.241\.240\.634\.438\.250\.029\.745\.832\.342\.38\.328\.6Llama\-4\-Scout\-17B\-16E\-Instruct\[[6](https://arxiv.org/html/2607.15418#biba.bib6)\]53\.358\.158\.840\.752\.962\.552\.555\.362\.537\.858\.341\.948\.150\.042\.9Qwen3\-VL\-8B\-Instruct\[[11](https://arxiv.org/html/2607.15418#biba.bib11)\]53\.351\.661\.844\.450\.053\.145\.955\.381\.240\.566\.745\.246\.233\.342\.9Qwen3\-VL\-32B\-Instruct\[[11](https://arxiv.org/html/2607.15418#biba.bib11)\]58\.748\.479\.444\.455\.956\.355\.763\.275\.048\.766\.767\.757\.733\.342\.9Qwen3\-VL\-30B\-A3B\-Instruct\[[11](https://arxiv.org/html/2607.15418#biba.bib11)\]47\.851\.658\.829\.638\.253\.141\.048\.762\.529\.762\.545\.248\.116\.728\.6InternVL3\_5\-8B\[[12](https://arxiv.org/html/2607.15418#biba.bib12)\]50\.058\.158\.829\.644\.153\.149\.250\.056\.337\.858\.341\.938\.558\.328\.6InternVL3\.5\-30B\-A3B\[[12](https://arxiv.org/html/2607.15418#biba.bib12)\]41\.351\.635\.337\.026\.543\.842\.642\.150\.040\.545\.845\.146\.18\.314\.3InternVL3\_5\-38B\[[12](https://arxiv.org/html/2607.15418#biba.bib12)\]50\.048\.464\.733\.341\.259\.447\.554\.062\.535\.162\.554\.844\.216\.714\.3Phi\-4\-multimodal\-instruct\[[7](https://arxiv.org/html/2607.15418#biba.bib7)\]40\.241\.947\.129\.629\.440\.642\.640\.837\.540\.545\.835\.540\.425\.028\.6Random27\.225\.817\.740\.723\.537\.527\.929\.037\.518\.929\.229\.023\.016\.642\.8Human68\.475\.672\.659\.975\.881\.484\.975\.046\.871\.982\.782\.199\.283\.366\.2\* Undergraduate62\.871\.870\.253\.864\.366\.774\.953\.043\.649\.770\.775\.497\.469\.446\.2\* Graduate & Young Professionals78\.285\.476\.668\.873\.981\.983\.876\.050\.070\.081\.975\.0100\.083\.762\.5\* Professionals94\.990\.085\.093\.389\.195\.696\.096\.046\.796\.095\.696\.0100\.096\.690\.0Gemini\-3\-pro\-preview77\.280\.679\.470\.476\.581\.372\.180\.393\.875\.783\.377\.478\.833\.385\.7
## Appendix BAuthor contributions
- •Project Conception:Y\.J\., J\.F\., M\.G\.
- •Task Definition & Benchmark Design:Y\.J\., J\.F\., M\.G\.
- •Data Curation & Management:Y\.J\., J\.F\., M\.G\.
- •Benchmark Question Generation:Y\.J\., J\.F\., M\.G\.
- •Model Evaluation:Y\.J\., J\.F\.
- •Qualitative & Quantitative Analysis:Y\.J\., J\.F\.
- •Writing & Visualization:Y\.J\., J\.F\.
- •Supervision:Y\.J\., J\.F\., M\.G\.
## Appendix CLimitations and future work
- •Dataset Scale, Density, and the Expert BottleneckOur final dataset consists of 92 high\-quality, professionally annotated samples\. We acknowledge that this scale is smaller than general\-domain VQA benchmarks or DesignQA\[[2](https://arxiv.org/html/2607.15418#biba.bib2)\]\. However, as highlighted in comparisons \(in the main paper\), there is a distinct trade\-off between dataset size and reasoning depth\. High\-level cognitive tasks especially toward real\-world engineering workflows require annotations from licensed experts, creating an “expert bottleneck” that limits rapid scaling\. Crucially, despite the modest total count,DrawingVQAis larger and more diverse than the construction\-engineering subsets of existing generalist benchmarks\. As shown in Figure 4 in the main paper, the relevant subsets of MM\-Vet \(4 samples\), ScienceQA \(24 samples\), and RBench\-M \(54 samples\) are significantly smaller\. Furthermore, where other datasets typically focus on a single question type \(e\.g\., from textbook, exam\),DrawingVQAspans 7 distinct construction\-engineering workflow aspects\. While smaller in volume, our dataset offers a higher “information density” and a more comprehensive evaluation of professional versatility than existing alternatives\.
- •Closed vs\. Open\-EndedTo ensure robust and reproducible comparisons, the majority of our quantitative evaluation relies on Multiple\-Choice Questions\. We recognize, however, that real\-world engineering workflows operate in an open\-ended setting—practitioners do not select from a list of options but must generate solutions derived from compliance checks to spatial and quantity calculations\. While we included an open\-ended subset in our benchmark, it remains a challenge\. Future iterations ofDrawingVQAaim to establish a rigorous question set for open\-ended engineering VQA\.
- •Scope of Engineering DisciplinesDrawingVQAcurrently emphasizes structural engineering and construction management, utilizing real\-world IFC\-level structural drawings\. While structural discipline is the backbone of any construction, the domain is multidisciplinary\. Our current scope excludes other disciplines such as Architectural, Civil, and other critical system disciplines such as Mechanical, Electrical, and Plumbing \(MEP\)\. While the way to read IFC drawings and the main rationale are similar across disciplines, visual perception and the reasoning process differ\. MEP drawings rely heavily on schematic symbols, deciphering complex relations and connections, rather than the physical scale and dimensioning critical to structural drawings\. Future work must expand the dataset to cover these modalities\. This would allow for a more holistic assessment of an MLLM’s ability to understand construction engineering drawings in a meaningful way\.
## Appendix DBenchmark details
### D\.1AccessingDrawingVQA
DrawingVQAis an expertly curated benchmark datasets that utilizes Issued for Construction \(IFC\) level drawings\. In this iteration ofDrawingVQA, the focus was made on structural discipline drawings, as structural drawings include the backbone of any building structures that become the basis for coordination with other disciplines\.
The QA sets within the drawings focus on practical questions that would be asked by construction industry professionals to better understand design intent and design requirements\. Contrary to many dataset that may require ‘intrinsic’ or ‘latent’ knowledge,DrawingVQAfocuses on reasoning within the given context window of a drawing or a portion of the drawing, to truly test MLLMs capability in understanding local contexts without hallucinating\.
TheDrawingVQAis shared on[https://huggingface\.co/datasets/S2\-MIND/DrawingVQA](https://huggingface.co/datasets/S2-MIND/DrawingVQA)\. It is distributed under the Creative Commons Attribution\-NonCommercial\-ShareAlike 4\.0 International \(CC BY\-NC\-SA 4\.0\) license\. Please note that most of the drawings do contain private information such as the architects, engineers, and general contractors\. Any information that indicates private information were redacted\.
### D\.2General Guidelines
There are two types of questions asked within this QA set: Multiple\-Choice Questions, and Open\-Ended Questions\. The multiple choice questions included True/False, and choices of more than 4 options with one correct option\. The closed ended questions encompassed short descriptive responses, however, the questions were asked in a way to ensure that long answers were avoided\. Each questions had at least one input of a construction drawing image, where it was either a full drawing image, or a proportion of the full drawing image\.
Construction drawings utilized in the industry are typically saved as PDFs for portability\. For the purpose of benchmarking, these drawings were converted into PNG images\. As part of a process to export from PDF to images, there are various resolutions that can be selected in the unit of pixels per inch \(ppi\)\. For theDrawingVQA, range of resolution from 25 to 125 were extracted for testing\. Practically, the resolutions can defer based on the usage, where for lightweight purpose, the ppi can be as low as 75 ppi, but in other cases such as for printing applications, this can be up until 300 ppi\.
The table below shows the average width and height of the images used, as it contains full drawing images to portion of the drawing image, to a very small patch for symbol detection\.
### D\.3Question curation guidelines
To curate the dataset, 33 Issued for Construction level structural discipline drawings were collected\. These drawings became the basis for brainstorming potential questions that can be asked, specifically in the sense that if this was a real project, what sort of questions would professionals ask first when seeing the drawings\. These questions and answer pairs were recorded first to ensure that the complexity of the questions were realistic, and matched what was needed for a meaningful benchmark\.
Upon completion for the first round of Question\-Answer creation, a more scrutinized evaluation was made to come up with plausible options that may arise as a result of misreading or misunderstanding the drawings\. These options were added as part of the Question\-Answer pair\. Finally, these questions were evaluated and designated for its dual categorization to possibly create meaningful insights upon completion of the benchmark testing\.
Table 5:Drawing image size variations in pixels
### D\.4Dataset VQA Structure
The code block below indicates the dataset schema and structure used to conduct the test in various models this work has benchmarked\.
\{
"id":Integer,
"image\_name":String,
"image2\_name":String,
"question":String,
"options":\{
"A":String,
"B":String,
"C":String,
"D":String
\},
"answer":String,
"explanation":String,
"cv\_field":List\[String\]
"cv\_subfield":List\[String\]
"ce\_field":List\[String\]
"ce\_subfield":List\[String\],
"topic\_difficulty":String,
"question\_type":String
\},
### D\.5Training contamination mitigation
A critical challenge in benchmarking MLLMs is the risk of data contamination, where test samples inadvertently appear in the model’s pre\-training corpus\. This leakage allows models to solve problems through memorization rather than reasoning\.
To guarantee the integrity of our evaluation,DrawingVQAis designed to be strictly contamination\-free through two key mechanisms:
- •Proprietary and Offline Image Sources:Unlike benchmarks derived from public internet crawls, the IFC construction engineering drawings in our dataset were sourced from private construction projects and have not been visible to public search engines\.
- •Novel Annotation and Temporal Separation:The Question\-Answer \(QA\) pairs were generatedde novoby our research team in late 2025, following industry practices\. This creation date post\-dates the knowledge cutoff of all models tested\.
Consequently, we can confirm thatDrawingVQArepresents a true zero\-shot evaluation environment\. The performance metrics reported in this study reflect the models’ genuine capability to contextualize novel visual engineering context window of a drawing, rather than their ability to recall training data\.
### D\.6Question Format: Multiple\-Choice vs\. Open\-Ended
To ensure rigorous and reproducible evaluation,DrawingVQAemploys a Multiple\-Choice Question \(MCQ\) format\. The distractors within these MCQs are meticulously expert\-designed to reflect realistic cognitive and procedural errors commonly made by practitioners when interpreting IFC \(Issued for Construction\) drawings\. This design ensures the benchmark serves as a reliable proxy for real\-world engineering challenges\.
To verify that the closed\-ended format does not artificially alter the inherent difficulty of the tasks, we conducted an ablation study by converting a subset of 20 MCQs into open\-ended questions\. An evaluation of MLLM performance across these two formats yielded no statistically significant difference in accuracy \(paired t\-test,p=0\.09p=0\.09\)\. This confirms that the MCQ format inDrawingVQAeffectively evaluates reasoning capabilities without introducing a format\-based advantage or bias\.
### D\.7Generalization to Unseen Drawing Styles
We prioritize reasoning density over repetitive volume\. To further validate the generalization capabilities of our dataset, we conducted an additional evaluation using unseen data\. Specifically, we compiled a supplementary set of 49 new VQA pairs derived from 7 newly introduced drawings, distributing 7 QA sets evenly across 7 distinct Construction Engineering \(CE\) dimensions\.
Comparing model performance on this unseen set against the original benchmark, a paired t\-test revealed no statistically significant difference in overall MLLM accuracy \(p=0\.07p=0\.07\)\. This confirms that the performance benchmarked by our dataset remains consistent and generalizes effectively when models are confronted with different drawing styles, due to our drawing QA is careful engineering reasoning required not just extracting information\.
### D\.8Ethical Consideration
- •Copyright and Licensing:We maintain strict adherence to all applicable copyright and licensing regulations\.
- •Data Privacy and Anonymity:All project\-specific identifiers \(e\.g\., client names, specific site addresses\) were redacted prior to inclusion\.
## Appendix EDual Categorization Taxonomy
To move beyond aggregate accuracy metrics and diagnose specific model bottlenecks, we developed a dual\-category taxonomy\. This framework cross\-references theDomain\-Specific Dimension\(the engineering intent of the query\) with theCognitive Capability\(the underlying mechanism required by the MLLM to solve it\)\.
The taxonomy classifies queries into seven domain dimensions, ranging fromGeneral Administrativetasks to complexCode and Specification Compliance verification\. As detailed in Table[6](https://arxiv.org/html/2607.15418#A5.T6), each dimension is associated with the corresponding primary MLLM capabilities necessary for resolution:
- •OCR \(Text Regocnition/Understanding\):Essential for interpreting dense technical annotations, tables, and callouts \(i\.e\., other cross\-references\)\.
- •Visual Perception:Low\-level identification of graphical entities \(symbols, lines, shapes\), delineation of major spaces within drawings such as sub\-drawings or details, and object counting\.
- •Reasoning:Higher\-order processing, includingSpatial\(2D\-to\-2/3D reasoning\),Alignment\(cross\-referencing between visual and textual content\), andVisual\(compositional analysis\) reasoning\. An example can include understanding spatial composition of elements within the model \(e\.g\. this element A is on the left of element B, element C exists between grid line D and E\)\.
- •Knowledge:Retrieval of external domain facts or latent knowledge \(e\.g\., standard acronyms, drawing standards, industry best practice, specifications or other engineering standards\) not explicitly visible in the pixel data\.
Table 6:TheDrawingVQADomain\-Capability Map\. We define seven domain dimensions and identify the primary MLLM cognitive capabilities typically required to solve tasks within each category\.
## Appendix FExperiments details
### F\.1Evaluation Prompts and Parsing
To evaluate the model’s performance, we structured the input prompts to encourage Chain\-of\-Thought \(CoT\) reasoning and specified a strict output format\. Depending on the question type, we appended the following instructions to the input\.
Multiple Choice:
The following is a multiple choice question\.Think step by step and then output the answer in the format of\\\\backslash“The answer is \(X\)\\\\backslash” at the end\.\{\{\[QUESTION\]\}\}Options:\{\{\[CHOICES\]\}\}
Open\-Ended:
The following is an open\-ended question \(with an explicit numeric answer\)\.Think step by step and then output the answer in the format of\\\\backslash“The answer is \(X\)\\\\backslash” at the end\.\{\{\[QUESTION\]\}\}
Despite these explicit formatting instructions, the models occasionally generated heterogeneous output formats\. Common variations included:
- •Standard sentences: “The answer is A\.” or “The answer is \(A\)”
- •LaTeX formatting: “The final answer is\\boxed\{A\}\.”
- •Markdown emphasis: “The answer is \*\*A: XX\*\*\.”
- •Minimalist output: “A” or “A: XX”
To robustly extract the predicted label across these variations, we implemented a regular expression pattern designed to capture the answer key while ignoring surrounding formatting or punctuation\. Other variations, except the above list, are not handled in the parsing process\.
### F\.2Model details
Table[7](https://arxiv.org/html/2607.15418#A6.T7)details the specific versions and API endpoints utilized in our experiments to ensure full reproducibility of our results\.
Table 7:MLLM API endpoints and sources\.ModelAPI EndpointSourceGPT\-4ogpt\-4o\-2024\-08\-06OpenAI APIo3o3\-2025\-04\-16OpenAI APIGemini\-2\.5\-progemini\-2\.5\-flashGoogle Gemini APIGemini\-2\.5\-flashgemini\-2\.5\-proGoogle Gemini APIClaude\-4\.5\-Sonnetclaude\-sonnet\-4\-5\-20250929Claude APIClaude\-4\.5\-Haikuclaude\-haiku\-4\-5\-20251001Claude APILLaVA\-OneVision\-1\.5\-8B\-Instructlmms\-lab/LLaVA\-OneVision\-1\.5\-8B\-InstructHuggingFace, local inferenceLLaVA\-v1\.6\-Mistral\-7Bllava\-hf/llava\-v1\.6\-mistral\-7b\-hfHuggingFace, local inferenceLLaVA\-v1\.6\-34Bllava\-hf/llava\-v1\.6\-34b\-hfHuggingFace, local inferenceLlama\-3\.2\-11B\-Vision\-Instructmeta\-llama/Llama\-3\.2\-11B\-Vision\-InstructHuggingFace, local inferenceLlama\-4\-Scout\-17B\-16E\-Instructmeta\-llama/Llama\-4\-Scout\-17B\-16E\-InstructHuggingFace, local inferenceQwen3\-VL\-8B\-InstructQwen/Qwen3\-VL\-8B\-InstructHuggingFace, local inferenceQwen3\-VL\-32B\-InstructQwen/Qwen3\-VL\-32B\-InstructHuggingFace, local inferenceQwen3\-VL\-30B\-A3B\-InstructQwen/Qwen3\-VL\-30B\-A3B\-InstructHuggingFace, local inferenceInternVL3\.5\-8BOpenGVLab/InternVL3\_5\-8B\-HFHuggingFace, local inferenceInternVL3\.5\-38BOpenGVLab/InternVL3\_5\-38B\-HFHuggingFace, local inferenceInternVL3\.5\-30B\-A3BOpenGVLab/InternVL3\_5\-30B\-A3B\-HFHuggingFace, local inferencePhi\-4\-multimodal\-instructmicrosoft/Phi\-4\-multimodal\-instructHuggingFace, local inference
### F\.3Human baseline
As part of an effort to ensure that the tests being conducted on various SOTA MLLMs are reasonable and not superficial or impractical, a human baselines for the same questions were created to compare their performance against the latest models\.
Similar to how various models may have variant parameter sizes or access to different knowledge, a similar delineation was made to the human baseline tests based on their industry experience in years\. We recruited 52 participants within the Civil Engineering or Construction Management industry with 3 levels of expertise in the following manner: Undergraduate \(Years of experience = 0 years\); Young Professionals such as graduate students or those who have just started working in the industry \(0≤\\leqYears of experience≤\\leq2 years\); Professionals \(3≤\\leqYears of Experience\)\. This delineation will also aim to help create human benchmark checkpoints that are comparable with model performance\.
As many of the participants were students and working professionals, the questionnaire for the baselines was shortened to 20 questions per person, but the questions were asked in the exact same format as to the MLLMs, with one image \+ multiple choices for the closed\-ended questions, and one image \+ free response for the open\-ended questions\. The following rules were used to set up the human benchmark test:
- •Questions were provided in a google form\.
- •Participants did not view any questions prior to completing the form\.
- •Any internet, LLM or other resource access were prohibited, to preserve testing hygiene, where only the context window of the drawing image, as well as their prior experiences in the industry, are utilized\.
- •Time limit was set up to be 20 minutes for 20 questions, to ensure all participants had the same control variable, regardless of their experiences\.
At the end of the benchmark test, all the participants were asked to rank the difficulty of the QA set, as well as the realistic nature of the QA set to the practical industry application\. The results collected by the participants are shown below\.
Table 8:DrawingVQA Human Benchmark Test Feedback Results \(%\)\.Table 9:Participant votes for the dataset that best requires reasoning skills on real\-world AEC industry context and engineering practices\.
### F\.4Dense and MoE models
To understand the trade\-off between inference efficiency and reasoning accuracy on IFC construction drawings, we compared Dense architectures against Mixture\-of\-Experts \(MoE\) variants within the same model families\. MoE models are designed to reduce computational cost by activating only a subset of parameters per token\.
Table[10](https://arxiv.org/html/2607.15418#A6.T10)summarizes the performance across the Qwen, InternVL, Llama, and LLaVA families\. We observe a distinct trend whereDense architectures generally outperform their MoE counterparts on the DrawingVQA benchmark\.
- •Qwen Family:The Qwen3\-VL\-8B \(Dense\) achieved53\.3%accuracy, surpassing the Qwen3\-VL\-30B\-A3B \(MoE\) which scored47\.8%, despite the MoE model having nearly4×4\\timesthe total parameters\.
- •InternVL Family:A similar pattern emerges, where the Dense 8B model \(50\.0%\) significantly outperforms the MoE 30B variant \(41\.3%\)\.
- •Scaling Behavior:While the Llama\-4\-Scout MoE achieved a competitive 53\.3%, it required a massive total parameter budget of 109B to match the performance of the significantly smaller Qwen\-8B Dense model\.
These results suggest that for high\-fidelity visual tasks requiring global context interpretation—such as reading complex construction drawings—the sparsity of MoE models may be a limiting factor compared to the dense connectivity of standard Transformers\.
Table 10:Performance comparison between Dense and Mixture\-of\-Experts \(MoE\) models across different parameter scales\.†\{\\dagger\}These numbers are from each model’s report and official website\.
### F\.5No\-Image Ablations
To quantify the extent to which models rely on visual information versus language priors, we conducted a no\-image ablation study, a standard diagnostic in VQA evaluations\[[8](https://arxiv.org/html/2607.15418#biba.bib8),[9](https://arxiv.org/html/2607.15418#biba.bib9)\]\. In this setting, the MLLMs are provided with the textual question alone, without the accompanying visual context\. Following\[[8](https://arxiv.org/html/2607.15418#biba.bib8)\], we prepended the prompt with the following sentence:
> If an image is mentioned, ignore this information and try your best to answer the question\.
A rigorous VQA benchmark is “vision\-centric”, ensuring that questions cannot be solved solely through common sense reasoning or textual artifacts\.DrawingVQAwas explicitly curated to reflect realistic engineering problems encountered by professionals on job sites, which inherently require visual verification\.
As presented in Table[11](https://arxiv.org/html/2607.15418#A6.T11), the significant performance drop \(Δ\\Delta\) observed across high\-performing models provides strong empirical evidence thatDrawingVQAis vision\-dependent\. Qualitative analysis reveals that when text\-only models answer correctly, they often rely on hallucinations \(inventing scenarios\), bias toward specific answers \(e\.g\., defaulting to “True” on binary questions\), or exploit domain priors \(e\.g\., guessing typical quantities of drawing notes\)\. However, the substantial performance gap \(around or even below that random guess\) confirms that reasoning on this benchmark necessitates visual grounding\.
Notably, refusal rates offer insight into model safety and grounding capabilities\. Models such asGemini\-2\.5\-flashandLLaVA\-OneVisionexhibited high refusal rates \(55\.4% and 81\.5%, respectively\), correctly identifying that the questions were unanswerable without the visual drawing context\.
Table 11:Text\-Only Ablation Study\. Value is a percentage\.
### F\.6Breakdown results on Dual Category Mapping
To identify specific cognitive bottlenecks in construction drawing interpretation, we introduce our dual\-category mapping to the error instances of the top\-performing model, Gemini\-2\.5\-Pro\. Figure[6](https://arxiv.org/html/2607.15418#A6.F6)illustrates the distribution of failure causes across seven domain\-specific tasks\. The analysis reveals that failure modes are highly context\-dependent\. General administrative tasks \(“Admin\.”\) requireOCRandVisual Perceptioncapabilities from MLLMs, indicating the model gets answers by reading explicit information in drawings\. In contrast, technical tasks such as “Domain Element Identification” and “Dimensional Understanding \(Cross\-Referencing\)” exhibit a shift towardReasoning\(spatial, visual, and alignment\), OCR, andVisual Perceptioncapabilities required\. “QTO” \(Quantity Take\-off\) shows sensitivity toOCRandVisual alignmenterrors, highlighting the challenge of accurately grounding text within complex graphical geometries\. On the other hand, “Compliance” requires knowledge beyond explicit information within the drawings\. This heterogeneity indicates that a single optimization strategy is insufficient; distinct construction tasks require targeted improvements in specific multimodal capabilities\.
Figure 6:Dual\-Category Error Breakdown\.
### F\.7More experiments on Image resolution
We investigate the impact of image resolution on the performance of MLLMs within the context of engineering drawings\. In our dataset construction, we distinguish between Region of Interest \(ROI\) images and Full Drawing Sheets\. While ROI images are typically user\-generated crops that do not require standardization, the Full Drawing Sheet serves as the global context and must be rasterized at a resolution that balances legibility with computational efficiency\.
To determine the optimal setting, we generated variants of the full drawing sheets at pixel densities ranging from 25 PPI to 125 PPI\. Table[12](https://arxiv.org/html/2607.15418#A6.T12)details the resulting pixel dimensions\. Note that the100 PPIsetting yields a median resolution of4800×36004800\\times 3600pixels\. This resolution substantially exceeds standard 4K UHD resolution \(3840×21603840\\times 2160\), ensuring that fine\-grained details—such as dimension text, line weights, and hatch patterns—are preserved without the artifacts common in lower\-resolution rasterization\.
On top of this, we conducted an ablation study across varying resolutions to assess robustness\. As shown in Table[13](https://arxiv.org/html/2607.15418#A6.T13), model performance generally correlates with increased resolution\. However, the accuracy gains plateau beyond 100 PPI, with some models \(e\.g\., Gemini\-2\.5\-flash, Qwen3\) showing peak performance at 100 PPI rather than 125 PPI\. The 100 PPI setting achieved the highest average accuracy of55\.7%across all tested MLLMs\. Based on these results, we selected 100 PPI as the default resolution for the DrawingVQA benchmark, as it offers the best trade\-off between high\-fidelity visual detail and model performance\.
Table 12:Summary of Image Dimensions for Different PPI Variants on a Full Drawing Sheet\.Table 13:Ablation study on image resolution \(Accuracy %\)\. Best results are highlighted in bold\.
### F\.8More experiments on on Option Permutation and Prefix Sensitivity
Prior research indicates that Large Language Models \(LLMs\) often exhibit sensitivity to the ordering of choices and the assignment of option labels \(e\.g\., A, B, C, D\) in multiple\-choice settings, a phenomenon known as position bias or selection bias\[[14](https://arxiv.org/html/2607.15418#biba.bib14),[13](https://arxiv.org/html/2607.15418#biba.bib13)\]\.
To evaluate the robustness of the models onDrawingVQA, we introduced an ablation setting named “Prefix v2”\. In this setting, we shuffle the order of the answer options and randomize the assignment of alphabetical prefixes\. As shown in Table[14](https://arxiv.org/html/2607.15418#A6.T14), this randomization does not result in a significant difference \(Avg\. 49\.4% and 49\.5%\)\.
We attribute this result to the design of the originalDrawingVQA\. The original dataset features carefully curated options and their sequences that are intentionally designed to challenge humans’ reasoning capabilities\.
Table 14:Ablation Study on Random Option Prefixes\. Values represent accuracy percentages\. “Prefix v2” denotes results with shuffled option orders and randomized prefixes\.
### F\.9Ablation Study on Input Modality: The PDF Hybrid Setting
In standard construction industry workflows, stakeholders typically exchange drawings as vectorized PDFs \(e\.g\., exported from BIM/CAD platforms\) rather than rasterized, scanned images\. To align our evaluation with this real\-world practice, we introduced a PDF Hybrid experimental setting\. In this configuration, whenever a question references a full drawing sheet, the model is provided with the native vectorized PDF file instead of a raster image\. Conversely, if a question targets a specific Region of Interest \(ROI\), the input remains a cropped raster screenshot\. For questions containing both a full sheet and an ROI, the input includes both the raster crop with the vectorized full\-sheet PDF\. This hybrid approach mirrors practical scenarios where engineers might upload a full PDF sheet but query a specific visual detail\. We limited this study to proprietary MLLMs \(GPT, Gemini, Claude series\), as most current open\-weights models do not natively support PDF document ingestion\.
Table[15](https://arxiv.org/html/2607.15418#A6.T15)presents the comparative performance\. Theoretically, vectorized PDFs should offer superior fidelity by providing explicit text layers and sharp vector paths, bypassing the resolution constraints of rasterization\. However, our results indicate that this benefit is inconsistent and minor\. While GPT\-4o and o3 got performance gain \(\+5\.4%\), the Gemini and Claude families either stagnated or suffered performance degradation \(e\.g\., Gemini\-2\.5\-flash dropped by 15\.2%\)\. This suggests that MLLMs struggle to spatially ground the explicit information contained in vector PDFs\.
In queries such as“How many columns are located along vertical grid line Q?”, MLLMs in the PDF setting frequently fail despite the text “Q” and adjacent lines being machine\-readable\. The models appear to perceive the text and lines as disjointed lists of data rather than a spatially coherent map, failing to align the identifier “Q” with the specific vertical vector line and its intersecting column elements\.
Furthermore, the hybrid modality \(pairing a vector PDF full\-sheet with a raster ROI crop\) introduces amodality gap\. The models struggle to perform visual referencing when the context \(full sheet\) is represented in vector space while the detail \(crop\) exists in pixel space\.
This confirms thatDrawingVQAis a robust, visual\-centric benchmark; simply providing machine\-readable text via PDFs does not bypass the core requirement for complex spatial and visual reasoning\. Future research should focus on bridging this gap, potentially by developing architectures that can better align vector primitives with raster visual features\.
Table 15:Ablation Study on PDF Hybrid setting\.
## Appendix GError Analysis
We reviewed the Chain\-of\-Thought \(CoT\) responses on the best models that were available from our benchmark study, which was Gemini\-2\.5\-pro and QWen3\-VL\-8B\-Instruct\. The following figures illustrates the reasoning the model made to evaluate the options that were given to make the final determination\. The figures aim to illustrate where in their Chain\-of\-Thought went wrong\. For all the following figures, light green highlights and comments are added to indicate when models have performed well\. The highlight and comments are added in red when model makes an error or false claims\.
### G\.1Visual Perception errors
Figure 7:Visual Perception and Visual Grounding Errors for Structural Beam Identification based on Grid Lines \(by Gemini\-2\.5\-pro\)\.
### G\.2Visual and Text Perception errors
Figure 8:A Mix of Visual and Text Perception Errors for Column Schedule Tabular Identifications \(by Gemini\-2\.5\-pro\)\.
### G\.3Text Perception \(OCR\) errors
Figure 9:Text Perception \(OCR\) Errors on Beam Annotations \(by Qwen3\-VL\-8B\-Instruct\)\.
### G\.4Knowledge errors
Figure 10:Knowledge and perception error on reading section detail tags \(by Gemini\-2\.5\-pro\)\.
### G\.5Reasoning errors
Figure 11:Visual Perception and Reasoning Error on Recognizing Drain Tiles \(by Gemini\-2\.5\-pro\)\.Figure 12:Knowledge, Reasoning and Perception Error while calculating length of elements \(by Gemini\-2\.5\-pro\)\.
### G\.6Instruction Adherence and Output Formatting
We observed a distinct disparity in instruction following between proprietary and open\-sourced models\. Commercial MLLMs \(GPT, Gemini, Claude\) demonstrated robust adherence to formatting constraints with no syntax errors\. In contrast, a few smaller open\-sourced models and certain MoE variants exhibited occasional anomalies:
- •Prefix Omission:Outputting raw text without the required option label, despite the explicit prompt instruction to use the format “The answer is \(A\)” \(see Section[F\.1](https://arxiv.org/html/2607.15418#A6.SS1)\)\.
- •Language and Refusal Hallucinations:Reverting to non\-target languages \(e\.g\., generating Chinese characters such as 无法确定\) or producing generic refusal phrases like “I don’t know” instead of selecting a valid option\.
### G\.7Parametric Bias vs\. Visual Grounding
Beyond formatting adherence, we investigated the models’ ability to prioritize explicit visual evidence over general domain knowledge \(parametric priors\)\. MLLMs often exhibit a strong bias toward standard engineering patterns, which can lead to hallucinations when a specific drawing deviates from typical conventions\.
To evaluate this, we test an adversarial example where the visual detail contradicts common construction norms\. Figure[14](https://arxiv.org/html/2607.15418#A7.F14)illustrates a representative case involving a slab anchor detail\. In standard structural engineering practice, such anchors are typically welded “all\-around”\. However, in this specific test case, the standard “all\-around” circle symbol was intentionally removed from the welding notation\. The original example is shown in Figure[13](https://arxiv.org/html/2607.15418#A7.F13)\.
Despite the explicit visual absence of this symbol, models selected the “all\-around weld” option\. This indicates that the models are over\-relying on their training priors regarding how slab anchors areusuallydetailed, rather than grounding their reasoning in the objective visual syntax provided in the drawing\. This prior knowledge bias remains a significant challenge for automated checking systems, which must detect non\-standard or erroneous deviations rather than assuming standard compliance\.
Figure 13:Original example on “all\-around” welding example \(Gemini\-2\.5\-pro\)\.Figure 14:Adversarial example to exclude “all\-around” circle symbol to highlight parametric knowledge bias in MLLM \(Gemini\-2\.5\-pro\)\.
### G\.8Visual Attention
To investigate the interpretability of the model’s reasoning process, we extract the attention weights from the final layer of the Qwen3\-VL\-8B\-Instruct decoder\[[3](https://arxiv.org/html/2607.15418#biba.bib3)\]\. To ensure semantic coherence, we aggregate the attention maps for multi\-token words \(e\.g\., merging “P”, “\.”, and “8” into a single “P\.8” concept\) by averaging their respective weights\.
The visualization results in Figure[15](https://arxiv.org/html/2607.15418#A7.F15)reveal a disconnect between the model’s textual output and its visual grounding\. While the model attempts to localize general geometric features—showing some attention to linear drawing outlines when processing the token “grid line” or “P\.8”—it fails to accurately ground specific beam entities \(i\.e\., attention sink\)\. Crucially, when processing the target location “P\.8”, the attention map does not focus on the “P\.8” label in the drawing; instead, the focus drifts to irrelevant regions such as grid line symbol F, G, J, P, P\.1, G, or R\. This visual misalignment indicates that the model is effectively hallucinating the entity’s location \(i\.e\., spatial reasoning\) rather than looking at the right pixel area\. Consequently, this failure in visual grounding leads to a logic error: the model incorrectly predicts option, having “W24”, whereas the ground truth for the structural member at P\.8 is “W21x50”\.
We hypothesize that this phenomenon stems from a domain gap: MLLMs are predominantly trained on natural images and may lack the capability to interpret the sparse, symbolic, geometry representations found in 2D technical drawings\. These findings bring attention to future work focused on improving cross\-modal alignment and grounding specifically for engineering artifacts and 2D schematic views\.




Figure 15:Visual attention maps of Qwen3\-VL\-8B\-Instruct\. The maps illustrate where the model “looks” when processing each text token\.
## Supplementary References
- \[1\]R\. Xiong, Y\. Wang, S\. Gunhan, Y\. Zhu, and C\. Berryman\.Can ai master construction management \(cm\)? benchmarking state\-of\-the\-art large language models on cm certification exams\.arXiv:2504\.08779, 2025\.
- \[2\]A\. C\. Doris, D\. Grandi, R\. Tomich, M\. F\. Alam, M\. Ataei, H\. Cheong, and F\. Ahmed\.Designqa: A multimodal benchmark for evaluating large language models’ understanding of engineering documentation\.arXiv:2404\.07917, 2024\.
- \[3\]S\. Kang, J\. Kim, J\. Kim, and S\. J\. Hwang\.See what you are told: Visual attention sink in large multimodal models\.arXiv:2503\.03321, 2025\.
- \[4\]Meta\.Llama\-3\.2\-11b\-vision\-instruct\.Technical report, 2024\.
- \[5\]H\. Liu, C\. Li, Y\. Li, and Y\. J\. Lee\.Improved baselines with visual instruction tuning\.arXiv:2310\.03744, 2023\.
- \[6\]Meta AI\.The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, April 2025\.
- \[7\]Microsoft et al\.Phi\-4\-mini technical report: Compact yet powerful multimodal language models via mixture\-of\-loras\.arXiv:2503\.01743, 2025\.
- \[8\]J\. Burgess et al\.Microvqa: A multimodal reasoning benchmark for microscopy\-based scientific research\.InCVPR, pages 19552–19564, 2025\.
- \[9\]X\. Yue et al\.MMMU: A massive multi\-discipline multimodal understanding and reasoning benchmark for expert AGI\.InCVPR, pages 9556–9567, 2024\.
- \[10\]X\. An et al\.LLaVA\-onevision\-1\.5: Fully open framework for democratized multimodal training\.arXiv:2509\.23661, 2025\.
- \[11\]Qwen Team\.Qwen3 technical report\.arXiv:2505\.09388, 2025\.
- \[12\]W\. Wang et al\.InternVL3\.5: Advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv:2508\.18265, 2025\.
- \[13\]Z\. Yang, P\. Jian, and C\. Li\.Option symbol matters: Investigating and mitigating multiple\-choice option symbol bias of large language models\.InNAACL, pages 1902–1917, 2025\.
- \[14\]C\. Zheng, H\. Zhou, F\. Meng, J\. Zhou, and M\. Huang\.Large language models are not robust multiple choice selectors\.arXiv:2309\.03882, 2024\.Similar Articles
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
This paper introduces MechVQA, a dataset with 3.3k high-density mechanical engineering drawings and 21k question-answer pairs, along with the MechVL model that outperforms existing baselines by 7.57 percentage points on the MechVQA total score, advancing multimodal LLM understanding of mechanical drawings.
Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation
This paper introduces EngVQA, a multimodal benchmark for evaluating engineering reasoning in vision-language models, along with an 8-stage automatic evaluation framework that enables fine-grained analysis of reasoning failures. It reveals substantial limitations in current VLMs' engineering reasoning capabilities.
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
Introduces TableVista, a comprehensive benchmark for evaluating foundation models on multimodal table reasoning under visual and structural complexity, comprising 3,000 problems expanded into 30,000 multimodal samples. Evaluation of 29 models reveals performance degradation on complex layouts and vision-only settings.
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Introduces WorldBench, a visually diverse multimodal reasoning benchmark that reveals significant limitations in current multimodal large language models' visual understanding.
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
HyperGVL introduces the first benchmark for evaluating Large Vision-Language Models on hypergraph understanding and reasoning, featuring 84,000 QA samples across 12 tasks and real-world applications. The paper also proposes WiseHyGR, a generalizable router that enhances LVLM performance through adaptive hypergraph representations.