PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

arXiv cs.CL Papers

Summary

This paper introduces PlanBench-V, the first comprehensive benchmark for evaluating Vision-Language Models on spatial planning map interpretation, including an expert-annotated dataset and a four-dimension evaluation framework. Experiments show significant progress but highlight persistent challenges in implementation-oriented tasks.

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for decision-making, public communication, and institutional coordination. Their interpretation, however, requires fine-grained visual perception, spatial reasoning, and policy-informed professional judgment, creating major challenges for both human learners and AI systems. With the rapid progress of Vision-Language Models (VLMs), their use in urban planning analysis is gaining attention, yet existing multimodal benchmarks mainly target general visual understanding and overlook the domain-specific cognitive processes of planning practice. To address this gap, we introduce PlanBench-V, the first comprehensive benchmark for evaluating VLMs in spatial planning map interpretation. We first build the Spatial Planning Map Database (SPMD), an expert-annotated dataset of 223 planning maps and 1629 question-answer pairs curated by professional planners, covering diverse geographic regions and cartographic styles. We then propose a theory-informed evaluation framework assessing four progressive capabilities: Perception, Reasoning, Association, and Implementation, corresponding to the cognitive pipeline of planning map interpretation. Extensive experiments across two generations of VLMs show clear progress but persistent limitations. The best 2026 agentic reasoning model, Qwen3.6-Plus, substantially outperforms the best 2025 model, GPT-4o, by 27%. Nevertheless, all models still struggle with implementation-oriented tasks requiring evaluative judgment, policy sensitivity, and constraint-aware decision-making. These findings reveal fundamental limitations of current VLMs in professional planning contexts and highlight the need for domain-adaptive multimodal reasoning frameworks. Code and data are available at https://plangpt.github.io.
Original Article
View Cached Full Text

Cached at: 06/05/26, 08:07 AM

# A Spatial Planning Map Benchmark for Vision-Language Models
Source: [https://arxiv.org/html/2606.05744](https://arxiv.org/html/2606.05744)
He Zhu\*,1,2Junyou Su1,2Wen Wang1,2 Yijie Deng1,2Wenjia Zhang†,1,3 1Behavioral and Spatial AI Lab∗Equal contribution\.†Corresponding author\. Email: wenjiazhang@tongji\.edu\.cnTongji University 2Behavioral and Spatial AI LabPeking University 3College of Architecture and Urban PlanningTongji University

###### Abstract

Spatial planning maps play a central role in territorial governance by translating planning objectives, regulatory frameworks, and spatial strategies into visual representations\. These maps serve not only as technical instruments for decision\-making but also as key media for public communication and institutional coordination\. However, interpreting spatial planning maps requires a combination of fine\-grained visual perception, spatial reasoning, and policy\-informed professional judgment, posing substantial challenges for both human learners and artificial intelligence systems\. With the rapid development of Vision\-Language Models \(VLMs\), their potential application in urban planning analysis has attracted increasing attention\. Nevertheless, existing multimodal benchmarks largely focus on general visual understanding and fail to capture the domain\-specific cognitive processes embedded in planning practice\. In this work, we introducePlanBench\-V, the first comprehensive benchmark designed to evaluate VLM performance in spatial planning map interpretation\. First, we construct the Spatial Planning Map Database \(SPMD\), an expert\-annotated dataset containing 223 planning maps and 1629 question–answer pairs curated by professional planners, covering diverse geographic regions and cartographic styles\. Second, we propose a theory\-informed evaluation framework that assesses model capabilities across four progressive dimensions: Perception, Reasoning, Association, and Implementation, reflecting the full cognitive pipeline of planning map interpretation\. Extensive experiments across two generations of VLMs reveal a clear trajectory of progress, yet also persistent challenges\. While the best 2026 agentic reasoning model, Qwen3\.6\-Plus, substantially outperforms the best 2025 model, GPT\-4o, by a margin of 27%, all models continue to struggle with implementation\-oriented tasks that require evaluative judgment, policy sensitivity, and constraint\-aware decision\-making\. Our findings highlight fundamental limitations of existing VLMs in professional planning contexts and underscore the need for domain\-adaptive multimodal reasoning frameworks\. Code and data are available athttps://plangpt\.github\.io/\.

###### keywords:

spatial planning maps; vision\-language models; multimodal benchmark

## 1Introduction

Spatial Planning maps serve as essential tools in urban development, functioning as socioeconomic and thematic visualizations that document existing conditions, illustrate future scenarios, and guide policy implementation\. Through specialized cartographic expressions, planners communicate vision, priorities, and constraints using symbols and annotations that convey land use allocations, infrastructure layouts, and functional zoning\. These maps employ geographic space as a structured canvas, depicting planning elements with specific scales and orientations to ensure clarity in spatial arrangements\. Unlike general\-purpose maps, planning maps feature distinct representation elements and vary from large\-scale master plans to detailed urban designs, each with unique stylistic features\(Lynch[1984](https://arxiv.org/html/2606.05744#bib.bib20), Steinitz[1995](https://arxiv.org/html/2606.05744#bib.bib21), Healey[1997](https://arxiv.org/html/2606.05744#bib.bib22)\)\. This study focuses specifically on*statutory*spatial planning maps, as broader categories such as conceptual diagrams or urban design illustrations are not always tied to explicit, regulated planning intentions\. Enabling VLMs to effectively interpret these specialized planning maps would significantly enhance both professional practice and educational contexts in urban planning\.

While VLMs have demonstrated impressive capabilities in general vision\-language tasks, their effectiveness in highly specialized domains remains largely unexplored\. Planning maps embody a high degree of complexity and professional specificity, integrating fine\-grained visual elements \(such as symbols, legends, and color codes\), intricate spatial structures and layout relationships, and planning semantics closely tied to regulatory and policy frameworks\. However, current VLMs often struggle to fully recognize essential cartographic elements like map legends, geographic entities, and planning boundaries\. They also exhibit limited ability to parse complex spatial relations or infer the embedded planning logic and spatial strategies\. Furthermore, VLMs generally lack a deep understanding of domain\-specific terminologies and formal planning expressions, which impedes accurate interpretation and normative reasoning\.

In addition, spatial planning has increasingly emphasized human\-centered governance and public participation\. The shift from technocratic planning toward participatory approaches highlights the growing importance of ensuring that planning information is accessible to diverse audiences, including residents, developers, policymakers, and researchers\. Yet current VLMs provide insufficient support for non\-expert users, thereby limiting public engagement and transparency\. Their inability to effectively convey planning intentions across stakeholder groups also weakens interdisciplinary collaboration and hampers the development of intelligent, data\-driven planning support systems\. If VLMs are to function as mediators between technical planning documents and broader audiences, they must be capable of accurately interpreting planning maps while respecting institutional and regulatory contexts\. The absence of systematic evaluation frameworks hinders progress toward this goal\.

To address this gap, we introducePlanBench\-V, a benchmark specifically designed to assess the planning map understanding capabilities of VLMs\. PlanBench\-V consists of two main components: a high\-quality dataset and a domain\-informed evaluation framework\.

The dataset comprises 223 spatial planning maps, derived from \(1\) official master plans in China; \(2\)Chinese Certified Urban\-Rural Planner Qualification Examination \(CURPQE\); and \(3\) international spatial planning maps, to ensure both accuracy and diversity in visual styles\. Complementing this, we construct 1629 question\-answer pairs manually annotated by professional urban planners, providing a rigorous basis for evaluating reasoning grounded in real\-world planning practice\.

We developed a comprehensive benchmark to systematically assess the capabilities of VLMs in understanding spatial planning maps\. As illustrated in Figure[1](https://arxiv.org/html/2606.05744#S1.F1), this framework is structured across four key dimensions:

- •Perception: evaluates models’ ability to identify visual components, including layout configurations, textual annotations, basic geographic features, and drawing elements\.
- •Reasoning: focuses on how well the model can derive structured insights and perform domain\-specific reasoning from the complex visual and semantic information embedded in planning maps\.
- •Association: assesses the ability to collect and relate background policies, regulations, and planning indicators relevant to planning maps\.
- •Implementation: addresses the capacity for comparing, critiquing, and optimizing planning proposals\.

![Refer to caption](https://arxiv.org/html/2606.05744v1/images/benchmark.png)Figure 1:Overview of the proposed PlanBench\-VIn summary, our contributions are as follows:

1. 1\.We introducePlanBench\-V, the first benchmark explicitly designed to evaluate vision\-language models in the context of spatial planning map interpretation, grounded in real\-world planning documents and expert knowledge\.
2. 2\.We propose a four\-dimensional assessment framework—Perception, Reasoning, Association, and Implementation—that reflects the progressive cognitive structure of planning practice, bridging foundational map literacy and applied professional judgment\.
3. 3\.Through extensive evaluation across two generations of VLMs, we reveal substantial performance gains driven by agentic reasoning, yet also persistent gaps in implementation\-oriented tasks that require policy sensitivity and evaluative judgment, offering guidance for future domain\-adaptive multimodal model development\.

## 2Related Work

### 2\.1Spatial Planning Maps

Spatial Planning maps represent a specialized form of cartographic visualization that has evolved alongside urban planning practice\. They are not merely technical drawings of space; rather, they are symbolic representations that communicate planning intentions, values, and ideologies within specific political and institutional contexts\(Harley[1989](https://arxiv.org/html/2606.05744#bib.bib18),[2002](https://arxiv.org/html/2606.05744#bib.bib19), Albrechtset al\.[2003](https://arxiv.org/html/2606.05744#bib.bib16), Dühr[2007](https://arxiv.org/html/2606.05744#bib.bib12)\)\. These maps are integral to spatial planning processes, functioning as tools of communication, negotiation, and persuasion\.

Spatial planning maps can be broadly categorized according to their scale, purpose, and legal status\. At the national and regional level, strategic spatial maps provide high\-level visions for spatial development, often emphasizing economic corridors, ecological networks, and urban hierarchies\(Faludi[2000](https://arxiv.org/html/2606.05744#bib.bib15), Albrechtset al\.[2003](https://arxiv.org/html/2606.05744#bib.bib16)\)\. At the municipal level, zoning maps delineate specific regulations regarding land use, density, and building form\(Burgess and Carmona[2009](https://arxiv.org/html/2606.05744#bib.bib14)\)\. Between these extremes, a range of non\-statutory or illustrative maps exists, including conceptual diagrams, master plans, and urban design frameworks that aim to foster public engagement and design communication\(Dühr[2009](https://arxiv.org/html/2606.05744#bib.bib13), Liet al\.[2025](https://arxiv.org/html/2606.05744#bib.bib11)\)\.

The cartographic design of these maps presents distinctive challenges in encoding abstract policies into visual form\. This process involves careful decisions about symbol systems, visual hierarchy, and spatial composition\. The semiotic dimension, encompassing visual grammar, symbolization, and layout, serves as the foundation for making complex planning intentions interpretable to various audiences\(Bertin[1983](https://arxiv.org/html/2606.05744#bib.bib9), Dühr[2007](https://arxiv.org/html/2606.05744#bib.bib12)\)\. Complementing this, analytical typologies classify maps into descriptive, evaluative, and prescriptive forms, each serving specific functions: documenting existing conditions, assessing potentials, or guiding future spatial organization\(Moroni and Lorini[2017](https://arxiv.org/html/2606.05744#bib.bib10)\)\.

The interpretation of spatial planning maps relies on a combination of perceptual and cognitive mechanisms that are grounded in both innate visual processing and acquired knowledge structures\. At the perceptual level, users extract immediate visual cues, such as shape, size, proximity, and contrast, that are essential for recognizing zoning patterns, infrastructure alignments, and land\-use hierarchies\(Palmer[1999](https://arxiv.org/html/2606.05744#bib.bib2), Rautenbachet al\.[2017](https://arxiv.org/html/2606.05744#bib.bib50)\)\. Gestalt principles like proximity, continuity, and figure\-ground differentiation play a central role in how viewers form holistic spatial understandings from complex visual data\(Wertheimer[1938](https://arxiv.org/html/2606.05744#bib.bib4), Koffka[1935](https://arxiv.org/html/2606.05744#bib.bib3)\)\. Beyond perception, cognitive processes such as schema activation and pattern recognition enable users to interpret map symbols, like colors, lines, and textures, based on prior knowledge and contextual expectations\(Mayer[2005](https://arxiv.org/html/2606.05744#bib.bib5), Arnheim[1969](https://arxiv.org/html/2606.05744#bib.bib6)\)\. By leveraging both visual and cognitive principles, planners can ensure that spatial representations are not only technically accurate but also intuitively accessible across diverse stakeholder groups\(Kress and van Leeuwen[2006](https://arxiv.org/html/2606.05744#bib.bib8), Alexander[1977](https://arxiv.org/html/2606.05744#bib.bib7)\)\.

Despite the extensive understanding of spatial planning maps, gaps remain in the cognitive processes involved in map interpretation, particularly regarding how different stakeholders decode and utilize these maps\.

### 2\.2VLMs and Benchmarks in Specialized Domains

VLMs have demonstrated remarkable progress in general multimodal tasks, including image understanding\(Hurstet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib200), DeepMind[2023](https://arxiv.org/html/2606.05744#bib.bib226)\), visual reasoning\(Zhuet al\.[2025c](https://arxiv.org/html/2606.05744#bib.bib63), Guoet al\.[2025](https://arxiv.org/html/2606.05744#bib.bib183)\), and multimodal dialogue\(Liuet al\.[2023](https://arxiv.org/html/2606.05744#bib.bib106), Wanget al\.[2024a](https://arxiv.org/html/2606.05744#bib.bib96)\)\. Recent research has successfully extended these models to specialized domains such as medical imaging\(Liet al\.[2023](https://arxiv.org/html/2606.05744#bib.bib56), Laiet al\.[2025](https://arxiv.org/html/2606.05744#bib.bib57), Panet al\.[2025](https://arxiv.org/html/2606.05744#bib.bib58)\), geographical information systems\(Zhanget al\.[2024b](https://arxiv.org/html/2606.05744#bib.bib59),[a](https://arxiv.org/html/2606.05744#bib.bib60)\), and mathematical reasoning\(Chenet al\.[2025](https://arxiv.org/html/2606.05744#bib.bib61), Shenet al\.[2025](https://arxiv.org/html/2606.05744#bib.bib62)\)\. The development of comprehensive benchmarks has been crucial for advancing multimodal AI capabilities\. Established datasets like VQA\(Agrawalet al\.[2016](https://arxiv.org/html/2606.05744#bib.bib35)\), COCO\-Captions\(Linet al\.[2015](https://arxiv.org/html/2606.05744#bib.bib36)\), and Visual Genome\(Krishnaet al\.[2017](https://arxiv.org/html/2606.05744#bib.bib37)\)have driven progress in general vision\-language understanding\. More recent specialized benchmarks have emerged for specific domains, including scientific diagrams\(Li and Tajbakhsh[2023](https://arxiv.org/html/2606.05744#bib.bib40), Robertset al\.[2024](https://arxiv.org/html/2606.05744#bib.bib41)\), document understanding\(Mathewet al\.[2021](https://arxiv.org/html/2606.05744#bib.bib39), Wanget al\.[2024b](https://arxiv.org/html/2606.05744#bib.bib38)\), aesthetics\(Huanget al\.[2024](https://arxiv.org/html/2606.05744#bib.bib64), Zhouet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib66), Linet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib65)\), autonomous driving\(Qianet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib67), Simaet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib68)\), and other fields\.

In the geographical and cartographic domain, MapQA\(Changet al\.[2022](https://arxiv.org/html/2606.05744#bib.bib33)\)and the Charting New Territories dataset\(Robertset al\.[2023](https://arxiv.org/html/2606.05744#bib.bib34)\)provided an initial evaluation for map\-based question answering, such as localization and identification\. However, these resources primarily address basic geographic interpretation rather than the complex domain\-specific requirements of planning map analysis\.

We identify urban planning as a critical domain that could significantly benefit from specialized VLMs tointerpret complex planning maps—a task where even leading commercial models exhibit substantial limitations in recognizing specialized elements and applying the cartographic interpretation skills essential for planning practices\. Recent efforts toward domain\-tailored systems, such as PlanGPT\(Zhuet al\.[2025a](https://arxiv.org/html/2606.05744#bib.bib31)\), which integrates customized retrieval, domain\-specific knowledge activation, and tool orchestration to support urban planning workflows, and its multimodal extension PlanGPT\-VL\(Zhuet al\.[2025b](https://arxiv.org/html/2606.05744#bib.bib32)\), illustrate the value of planning\-aware model design\. However, these systems mainly target text generation or general visual understanding rather than the structured interpretation of planning maps, leaving the evaluation of map\-grounded perception, reasoning, association, and implementation largely open—a gap that PlanBench\-V is designed to fill\.

## 3Constructing Database

Understanding planning maps is an inherently abstract and complex task due to their highly specialized visual and semantic nature\(Zhuet al\.[2025b](https://arxiv.org/html/2606.05744#bib.bib32)\)\. To ensure the reliability of our benchmark, the construction of a high\-quality, professionally annotated planning map dataset is essential\. This SPMD was meticulously curated by experts with domain\-specific knowledge in urban planning and geography\.

Our data synthesis pipeline consists of three core stages: parsing, question generation, and answer generation\.

First, in the parsing stage, we employed custom scripts to extract planning maps from official planning documents sourced from municipal governments, planning agencies, and academic institutions\. Chinese planning documents were primarily collected from the official websites of provincial and municipal natural resource bureaus and from theCURPQE\. The spatial planning maps from other countries encompass a diverse range of cases as shown in Figure[2](https://arxiv.org/html/2606.05744#S3.F2)\. The dataset includes three general and two thematic planning maps from Japan’sTokyo 2040 Plan; four maps from the UK’sLondon 2021 Master Planand two from theLondon Greenway; and seven maps from the USNew York 2050 Master Plan, five thematic plans, and 50 additional maps from US cities, including San Francisco, Oakland, and Chicago\. The Netherlands contributes seven maps from theThe MosaicMaster Plan, Canada provides one site plan from theGranville Island Redevelopment Projectin Vancouver, and Germany offers seven maps focused on village revitalization in Bavaria\. Additionally, one planning map each is included from Singapore and Mexico\. Given that many documents contained extraneous visual elements such as icons, background images, and real\-world photographs, we conducted manual screening to isolate clean planning maps\. Where necessary, we enhanced low\-resolution maps through sharpening or source tracing to ensure image clarity and legibility of embedded text\.

![Refer to caption](https://arxiv.org/html/2606.05744v1/images/data_map.jpg)Figure 2:Geographical Distribution of Planning Cases in SPMDBuilding on these parsed maps, a panel of experts with professional backgrounds in urban planning and geography, including 8 graduate\-level students trained in advanced urban design and 1 practitioner with over four years of experience at a planning institute, formulated a set of questions that targeted key dimensions of visual–spatial reasoning in planning scenarios\. These questions were explicitly designed to evaluate model capabilities across four core dimensions of planning map understanding: perception, reasoning, association, and implementation\.

To minimize hallucination and ensure factual correctness, reference answers were generated through a combination of automated processes and expert validation\. Specifically, answers initially generated by Qwen2\.5\-VL\-72B\-Instruct were reviewed, corrected, and rewritten by professional planners to produce authoritative references\. To improve the ecological validity and generalizability of the benchmark, the dataset incorporates diverse planning styles and visual conventions, reflecting a variety of regional practices and map\-making traditions\. These efforts collectively enable more robust and reliable evaluation of VLMs’ capabilities in real\-world, expert\-level planning contexts\.

In total, we collected 223 images and 1629 QA pairs, as summarized in Table[1](https://arxiv.org/html/2606.05744#S3.T1)\. Specifically, these VQA pairs assess capabilities in perception \(25\.7%\), reasoning \(54\.5%\), association \(18\.0%\), and implementation \(1\.8%\)\. The visual materials cover diverse planning contexts, drawing on both Chinese sources \(China Master Plans andCURPQE\) and international cases from Asian, European, and North\-American cities\.

Table 1:Statistical Information of SPMDCapability LevelTaskTotalCMPaCURPQEOthersbPerceptionElement Recognition31716552100Caption10169320ReasoningClassification29414747100Spatial Relationship Reasoning3001495299Domain\-Specific Reasoning29414153100AssociationPolicy Association29414350101ImplementationScheme Evaluation200200Decision\-making9090\\tabnote

aCMP stands for China Master Plans;bSpatial planning maps from Other countries, including Japan, United Kingdom, United States, Netherlands, Canada, Germany, Singapore, and Mexico\.

## 4Developing the Metrics

Our investigation begins with a comparative analysis of established benchmarks across professional image domains, including Engineering Documentation\(Doriset al\.[2024](https://arxiv.org/html/2606.05744#bib.bib43)\), Graphic Design\(Linet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib65)\), and Image Aesthetics Perception\(Huanget al\.[2024](https://arxiv.org/html/2606.05744#bib.bib64)\)\. This review reveals the distinct nature of spatial planning maps, which occupy a unique position at the intersection of functional abstraction, symbolic representation, and domain\-specific conventions\.

Unlike mechanical design drawings, which are strictly technical and follow standardized engineering protocols, or artistic images that prioritize aesthetic expression and subjective interpretation, planning maps balance functionality and minimal aesthetics\. They employ simplified, codified, and standardized visual styles to depict land use, spatial layouts, transportation systems, and public facility allocations\. Furthermore, unlike photographs that can be intuitively understood based on visual familiarity, planning maps require both perceptual recognition and abstract reasoning rooted in planning theory and regulatory frameworks, making their interpretation a highly specialized cognitive task\.

To deconstruct this specialized task, we examined several foundational frameworks\. Map reading skills are outlined in Map Literacy frameworks\(Rautenbachet al\.[2017](https://arxiv.org/html/2606.05744#bib.bib50)\), which identifies core competencies including: recognizing symbology, orientation and locating, measuring and estimating, calculating and explaining, and extracting knowledge\. These skills form the mechanical basis for decoding the map’s surface information\. Transcending this, the ACRL Framework for Visual Literacy in Higher Education \(2022\) provides a higher\-order cognitive lens, emphasizing critical interpretation through its four frames: perceiving visuals as information, practicing discernment and criticality, participating in the visual information landscape, and pursuing social justice through visual practice\.

However, the ultimate measure of competence in urban planning lies in the application of this knowledge to solve real\-world problems\. TheCURPQEin China exemplifies this, particularly through its subjective questions on ’Urban Planning Practice’\. These questions assess the core professional capabilities of a contemporary planner: the ability to solve complex problems and to apply established rules, technical standards, and legal principles to practical scenarios\. This represents the synthesis of theoretical knowledge and interpretive skill into actionable professional judgment\.

Based on this analysis, it is evident that a holistic assessment of planning map comprehension requires a framework that bridges foundational literacy with applied professional competence\. We therefore propose a comprehensive evaluation framework constructed upon the urban planning knowledge system\. The framework’s logic follows a progressive path from theoretical foundations, through perceptual and interpretive mechanisms, to applied problem\-solving capabilities\. It specifically comprises the following four core dimensions:

### 4\.1Perception

The questions of the perception subset consist of 2 categories as follows:

Element Recognitionevaluates models’ ability to identify layout configurations, textual annotations, basic geographic features, and drawing elements in planning maps\. It facilitates the establishment of semantic alignment between image content and natural language\. This dimension includes a total of 317 VQA items, of which 165 are from the China Master Plan, 52 are from theCURPQEand 100 from other countries\.

\\tbl

Example Questions on elements of Territorial Spatial Master Plan Maps Refer to Specification for the Mapping of Municipal Territorial Spatial Master Plans in China

\\tbl

Example Questions on spatial relationships of Territorial Spatial Master Plan Maps

\\tbl

Example Questions on domain\-specific reasoning of Territorial Spatial Master Plan Maps

Caption, in this study, refers to extracting as many details from the image as possible\. The model generates descriptions based solely on the image itself, rather than identifying elements in response to a specific question\. The quality of the caption reflects whether the model has truly ”seen” the image clearly\. This dimension includes a total of 101 VQA items, of which 69 are from the China Master Plan and 32 are from theCURPQE\.

An example Question would be:”Please describe this planning map in detail\.”

### 4\.2Reasoning

This dimension encompasses classification, spatial relationship reasoning, and domain\-specific reasoning\.

Classificationfocuses on the ability to recognize different types of planning maps\. Due to the fact that this benchmark focuses on statutory spatial planning, other forms such as conceptual plans, urban design, district\-scale strategic planning, and thematic studies are currently not included\. Based on China’s “five\-level, three\-category” planning system, the maps are categorized into master plans, detailed plans \(including both regulatory and site plans\), and specialized planning documents\. For maps without explicitly labeled types, such as those appearing in exams, experts annotated their types based on the scale of the map \(e\.g\., national, provincial, municipal, or township level\) and the visualized planning content\. Maps presenting broad spatial structures without detailed regulatory indicators are generally considered master plans\. Those containing control indicators such as building density, height limits, floor area ratio, green space ratio, and redline boundaries are identified as regulatory detailed plans\. In contrast, maps that display site layouts or architectural schemes are classified as site plans, while maps focusing on specific domains or policy areas are categorized as specialized plans\. This dimension includes a total of 294 VQA items, of which 147 are from the China Master Plan, 47 are from theCURPQEand 100 from other countries\.

Spatial Relationship Reasoning, as shown in Table[4\.1](https://arxiv.org/html/2606.05744#S4.SS1), assesses the understanding of spatial relationships between geographic elements in planning maps, including: topological spatial relations, sequential spatial relations, and metric spatial relations\. This dimension includes a total of 300 VQA items, of which 149 are from the China Master Plan, 52 are from theCURPQEand 99 from other countries\.

Domain\-specific Reasoning, as shown in Table[4\.1](https://arxiv.org/html/2606.05744#S4.SS1), encompasses four key aspects: spatial layout, functional organization, transportation system, and environmental ecology, each reflecting a critical dimension of professional interpretation in territorial spatial planning\. This dimension includes a total of 294 VQA items, of which 141 are from the China Master Plan, 53 are from theCURPQEand 100 from other countries\.

### 4\.3Association

This dimension assesses the ability to collect and relate background policies and contextual documents relevant to planning maps\. At a fine scale, it examines policy, regulations, and planning indicators\. In the third part ofPlanBench\-V, Association refers to the model’s ability to retrieve, relate, and reason over background policies and contextual documents that are relevant to the content of planning maps\. Unlike tasks that focus on visual identification or structural interpretation, association requires models to connect spatial elements with appropriate policy frameworks, including administrative hierarchies, regulatory classifications, and land\-use guidelines\. This task evaluates whether the model can ”understand” not only what is shown on the map, but also the broader institutional and legal context in which the design operates\.

At a finer level, this dimension includes tasks involving national and local planning policies, special\-purpose plans, and technical reports\. The model is expected to recognize references to zoning regulations, ecological protections, urban renewal standards, and policy constraints implicitly indicated by visual symbols or layout patterns\. A strong performance reflects the model’s capacity to bridge visual and textual modalities and to align spatial content with normative planning knowledge\.

This dimension includes a total of 294 VQA items, of which 143 are from the China Master Plan, 50 are from theCURPQEand 101 from other countries\.

An example question would be: ”Based on the content of this planning map, which relevant policies or guidelines should be considered when evaluating its implementation?”

### 4\.4Implementation

This dimension addresses the capacity for comparing, critiquing, and optimizing planning proposals\. Unlike tasks focused on recognition or factual recall, Implementation requires models to make evaluative judgments under real\-world constraints, such as balancing ecological protection, land\-use efficiency, and development goals\. This dimension reflects whether a model can move beyond descriptive outputs and generate insights that are strategic, selective, and professionally grounded\.

Given the integrative nature of urban planning, models are expected not only to identify problems in design proposals but also to reason through trade\-offs and suggest actionable revisions\. Strong performance in this dimension indicates a grasp of planning logic, an ability to critique spatial solutions, and the use of terminology aligned with professional practice\.

The questions of the Implementation subset consist of 2 categories as follows:

Scheme Evaluationmeasures the model’s ability to evaluate the strengths and weaknesses of a planning proposal based on the visual input and spatial context\. Responses are expected to reflect planning\-specific criteria such as land use distribution, spatial continuity, accessibility, and compatibility\. This subset includes 20 VQA items, all derived from theCURPQE, and presents both schematic diagrams and explanatory texts\. Example question: “What are the strengths and weaknesses of this planning scheme?”

Decision makingtests whether the model can make value\-driven choices in a constrained design scenario\. The model must compare multiple alternatives or identify optimal directions based on planning principles and contextual information\. It includes 9 VQA items, all fromCURPQE, often accompanied by prompts requiring the selection or justification of preferred options\.

Example question: “Which planning direction is more suitable for the site, and why?”

## 5Experiments

### 5\.1Experimental Setup

We conducted extensive experiments on 17 VLMs to evaluate their capabilities in interpreting planning maps\. The first round includes GPT\-4o and GPT\-4o\-mini together with nine state\-of\-the\-art variants from the popular Intern\(Chenet al\.[2023](https://arxiv.org/html/2606.05744#bib.bib29)\)and Qwen\-VL\(Baiet al\.[2023](https://arxiv.org/html/2606.05744#bib.bib28)\)model families, while the second round adds six agentic reasoning models\. To ensure completeness and fairness, all VLMs were evaluated using their original released weights without any dataset\-specific fine\-tuning\.

Qwen and InternVL, especially 2\.5 and 3\-series models, were selected due to their strong performance as open\-source models in existing multimodal benchmarks\. MMMU\(Yueet al\.[2024](https://arxiv.org/html/2606.05744#bib.bib30)\)consists of 11\.5k multimodal questions derived from university\-level course content, and MMBench v1\.1 covers a wide range of disciplines from fundamental sciences to engineering applications\. These benchmarks motivated the inclusion of Qwen2\.5\-VL\-72B, InternVL2\.5\-78B\-MPO, and InternVL3 variants as representative open\-source baselines\.

To systematically analyze the impact of model scaling laws, our selection spans a full range from lightweight to flagship models\. Within the first\-round Qwen series, we included models with 2B, 3B, 7B, and 72B parameters\. In the InternVL3 series, we selected models with 8B, 9B, and 14B parameters\. This broad parameter spectrum provides a solid empirical basis for investigating the trade\-off between performance and resource consumption, as well as how different capability dimensions \(e\.g\., perception, reasoning\) evolve with model size\. Crucially, to assess performance under practical deployment scenarios, we also included models that employ AWQ quantization, namely Qwen2\-VL\-72B\-Instruct\-AWQ and Qwen2\.5\-VL\-72B\-Instruct\-AWQ\. These quantized entries provide a practical view of compressed flagship\-scale models alongside smaller non\-quantized models\.

Our evaluation methodology addresses the inherent complexity of planning problems by accommodating multiple valid approaches to the same challenge\. Rather than enforcing a single correct answer, our framework evaluates responses based on adherence to planning principles, logical consistency, evidence\-based reasoning, and consideration of diverse stakeholder perspectives\. For bilingual evaluation, we developed specialized protocols that account for cross\-cultural variations in planning terminology, regulatory frameworks, and professional practices\.

### 5\.2Exploration of Prompt Refining

Given that the benchmark tasks primarily consist of subjective, judgment\-intensive questions, we adopted the LLM\-as\-a\-judge evaluation framework enhanced with structured scoring rubrics and automated prompt optimization\. To establish a human performance baseline and validate the reliability of automated evaluation, we recruited three graduate students specializing in urban planning to answer 40 questions sampled across all eight task types \(five questions per type\)\. Their responses were independently scored by four LLM judges—GPT\-4o\-mini, GPT\-4o, Claude Opus 4\.7, and GPT\-5\.4—using identical rubrics\.

#### 5\.2\.1Human Performance Baseline and Judge Calibration

A clear strictness gradient emerges: GPT\-4o\-mini is the most lenient judge \(1\.377/2, 68\.9%\), followed by GPT\-4o \(1\.342/2, 67\.1%\), Claude Opus 4\.7 \(1\.289/2, 64\.5%\), and GPT\-5\.4 \(1\.255/2, 62\.8%\)\. The per\-item Spearman rank correlations among all judge pairs exceed 0\.75 \(p<0\.001p<0\.001for all pairs\), indicating strong agreement on the relative quality ordering of human responses\. The four judges naturally partition into two clusters: GPT\-4o\-mini and GPT\-4o form a more lenient group \(ρ=0\.894\\rho=0\.894\), while Claude Opus 4\.7 and GPT\-5\.4 form a stricter group \(ρ=0\.921\\rho=0\.921\)\. GPT\-4o occupies a favorable intermediate position, exhibiting high cross\-cluster correlations \(0\.856 and 0\.790 with Claude Opus 4\.7 and GPT\-5\.4, respectively\)\.

At the per\-type level, correlations are moderately lower, with the weakest agreement observed between GPT\-4o\-mini and GPT\-5\.4 \(ρ=0\.548\\rho=0\.548, not significant\), driven primarily by task categories such as Decision\-making and Classification, where the two models apply substantially different evaluative standards\. These results confirm that, while LLM judges are well\-calibrated for within\-model ranking, absolute score interpretation should account for judge\-specific strictness\.

Critically, the adoption of structured reference answers with annotated critical points—rather than open\-ended rubrics—substantially narrows the gap between lenient and strict judges\. When scoring prompts are anchored to predefined evaluation criteria, even the most cost\-efficient model, GPT\-4o\-mini, achieves per\-item correlations above 0\.89 with GPT\-4o and above 0\.75 with the strictest judge, GPT\-5\.4\. This establishes a practical cost–accuracy tradeoff: deploying GPT\-4o\-mini with well\-designed critical\-point rubrics yields evaluation fidelity comparable to that of substantially larger judges, at a fraction of the inference cost\.

#### 5\.2\.2Prompt Optimization and Scoring Protocol

Our evaluation system was trained on expert\-annotated planning assessments and incorporates structured scoring rubrics to enable consistent and reliable evaluation across diverse planning subdomains\. To mitigate potential biases in subjective scoring, we developed a prompt optimization mechanism that iteratively aligns automated judgments with ratings provided by human experts\. Experimental results demonstrate that the proposed method achieves an average pairwise error below 0\.1 between human evaluators and the automated system, indicating high reliability even for complex, subjective planning tasks\.

We conducted evaluations using GPT\-4o\-mini \(temperature=0\) as the primary evaluation model\. To address variability introduced by API connection instability, we performed multiple evaluation runs per test case and averaged the resulting scores\. For text\-based tasks, we employed exact match or semantic similarity metrics for objective questions, while relying on the LLM\-as\-judge methodology with structured rubrics for subjective questions\. Visual tasks were analogously partitioned into caption\-based questions, which target descriptive analysis, and non\-caption questions, which target spatial reasoning and planning interpretation\.

We began by designing and testing multiple versions of evaluation prompts, each varying in phrasing, structure, and the granularity of evaluation dimensions\. These prompt variants were applied across a diverse set of planning QA tasks, and their outputs were compared against expert\-annotated ground truths\. Empirically, we observed that stronger\-performing VLMs tend to be more robust to prompt variation, yielding relatively stable scores across different evaluation formats\.

However, we also identified several challenges during prompt optimization\. We experimented with a range of scoring prompts spanning from minimal formulations to highly specific ones that emphasize geographic location, functional zoning, infrastructure layout, and element\-to\-legend alignment\. When an excessive number of evaluation dimensions were introduced—for instance, style, fluency, and technical vocabulary usage—the scoring output from the LLM\-as\-judge became less stable and exhibited weaker correlation with expert judgment\. Among these, the “language style” dimension exhibited particularly high variance, as models often struggled to distinguish between stylistic appropriateness and technical correctness\. Given the critical importance of factual accuracy in planning contexts, we ultimately excluded style\-related factors from the final rubric to preserve evaluation precision\. The incorporation of stylistic metrics is left for future refinement\.

We also observed that concise prompts, when combined with few\-shot scoring examples or structured reference answers such as annotated critical points, can significantly enhance the judgment accuracy and analytical capacity of the evaluation model\.

To summarize, the prompt refinement process revealed that:

- •Evaluation prompts with fewer, well\-defined dimensions produce more stable and reliable scores;
- •Models with stronger grounding and reasoning capabilities exhibit greater resilience to prompt variation;
- •Structured reference answers with annotated critical points enable small, cost\-efficient judges such as GPT\-4o\-mini to achieve evaluation fidelity comparable to that of substantially larger models, establishing a favorable cost–accuracy tradeoff for large\-scale benchmark evaluation\.

These findings informed the final prompt templates used in the benchmark, striking a balance among evaluation fidelity, interpretability, and computational efficiency\.

### 5\.3VLM Performance in Planbench

To characterize both the current capability of VLMs and the trajectory of progress in planning map interpretation, we conducted two rounds of evaluation on the same 300\-item stratified subset\. The first round \(2025\) assessed 11 models representing the state of the art in vision\-language pretraining at that time\. The second round \(2026\) re\-evaluated the identical items with six models that incorporate enhanced agentic reasoning capabilities\. All models were scored using the same LLM\-as\-judge protocol described in Section[5\.2](https://arxiv.org/html/2606.05744#S5.SS2)\. Results for both rounds are presented jointly in Table[2](https://arxiv.org/html/2606.05744#S5.T2)and Figure[3](https://arxiv.org/html/2606.05744#S5.F3), with a separator line distinguishing the two generations\.

![Refer to caption](https://arxiv.org/html/2606.05744v1/images/result3.png)Figure 3:Performance of 17 Models on Planbench\-V Vision\-Language Tasks across Two GenerationsTable 2:Performance of Models on PlanBench\-V Vision\-Language Tasks\. Overall scores are computed on a 300\-item stratified subset under the LLM\-as\-judge protocol described in Section[5\.2](https://arxiv.org/html/2606.05744#S5.SS2), and therefore are not the unweighted mean of the eight sub\-task columns\.ModelPerceptionReasoningAssociationImplementationOverallRankElem\.Recog\.CaptionClassifi\-cationSpatialReasoningDomainReasoningSchemeEval\.DecisionmakingFirst Round \(2025\)GPT\-4o1\.0511\.8781\.4061\.3051\.5641\.5271\.2231\.4291\.3421Qwen2\.5\-VL\-72Ba1\.2991\.8251\.4061\.2481\.2631\.2531\.1531\.0901\.2882InternVL3\-9B1\.1731\.8781\.2971\.2601\.4351\.2970\.9210\.9031\.2713Qwen2\.5\-VL\-7B1\.1011\.6281\.0890\.8651\.1101\.0690\.8021\.0541\.0504InternVL3\-14B0\.9311\.5801\.1770\.7930\.9981\.0980\.7090\.9170\.9805Qwen2\-VL\-72Ba1\.0101\.3671\.1250\.7460\.9671\.1140\.6700\.6320\.9636Qwen2\-VL\-7B0\.9021\.3861\.0310\.7160\.9430\.9790\.8570\.6570\.9107InternVL3\-8B0\.9921\.7831\.0260\.7980\.7510\.9260\.6311\.0730\.9098Qwen2\.5\-VL\-3B0\.8621\.5540\.9530\.6910\.8700\.9700\.6970\.9360\.8769GPT\-4o\-mini0\.6640\.8901\.0210\.7891\.0300\.9630\.6361\.1750\.86610Qwen2\-VL\-2B0\.7440\.9250\.9480\.5000\.6560\.9260\.5370\.7920\.73111Second Round \(2026\)Qwen3\.6\-Plus1\.7511\.9501\.6721\.6701\.7441\.6191\.6191\.7101\.7011Gemini\-2\.5\-Pro1\.4081\.7751\.6561\.4441\.4251\.4681\.4391\.5251\.4722GPT\-5\.41\.2331\.9001\.5621\.3831\.4861\.4381\.5861\.5081\.4313Kimi\-K2\.61\.5181\.5161\.3441\.4141\.3761\.3651\.4401\.3171\.4174Claude\-Opus\-4\.71\.1861\.8251\.3201\.2951\.5581\.4931\.4341\.3211\.3845Qwen3\.6\-Flash1\.3181\.6361\.2971\.3531\.4761\.3511\.0461\.4411\.3536
\\tabnote

aEvaluated using the AWQ 4\-bit quantized release \(Qwen2\-VL\-72B\-Instruct\-AWQandQwen2\.5\-VL\-72B\-Instruct\-AWQ\); displayed without the\-AWQsuffix for table readability\.

#### 5\.3\.1First Round \(2025\)

The first round evaluated 11 VLMs spanning a broad parameter range from 2B to 72B\. The best\-performing model was GPT\-4o \(Overall 1\.342\), followed by Qwen2\.5\-VL\-72B\-AWQ \(1\.288\) and InternVL3\-9B \(1\.271\)\. Scaling offered diminishing returns: InternVL3\-9B outperformed its larger counterpart InternVL3\-14B \(0\.980\), and flagship\-scale models did not uniformly dominate smaller models\. Detailed per\-dimension analysis is provided below\.

#### 5\.3\.2Evaluation on Perception

ThePerceptioncapability is assessed through two subtasks:Element RecognitionandCaptioning\. Among all evaluated models in the first round, Qwen2\.5\-VL\-72B\-AWQ achieved the strongest Element Recognition score \(1\.299\), while InternVL3\-9B and GPT\-4o tied for the highest Caption score \(1\.878\)\. This result demonstrates robust visual grounding and scene interpretation with clear, domain\-relevant terminology\.

In contrast, GPT\-4o\-mini \(0\.664 in Element Recognition, 0\.890 in Captioning\) was the weakest performer on the two perception subtasks among the complete evaluations\. The weakest small models struggled to correctly identify legend symbols and misattributed multiple design elements, particularly in crowded visual scenes\. Their captions also exhibited vagueness and lacked domain\-specific language\. Notably, Qwen2\-VL\-2B showed limited overall capability with an average score of 0\.731\.

Overall, perception performance scales with model size up to a point, but several common failure patterns persist even in large models: color misclassification, hallucination of objects absent from the image, and degraded reliability on crowded or multi\-element scenes\.

#### 5\.3\.3Evaluation on Reasoning

TheReasoningcapability consists of three subtasks:Classification,Spatial Relationship Reasoning, andDomain\-Specific Reasoning\. This dimension assesses the ability of a model to categorize urban elements accurately, understand their spatial relations, and reason with planning knowledge\.

The top performer in this dimension was GPT\-4o, achieving an overall Reasoning score of 1\.425\. It attained 1\.406 in Classification, 1\.305 in Spatial Reasoning, and 1\.564 in Domain\-Specific Reasoning—the highest score among first\-round models in the latter category\. These results demonstrate the robust capacity of GPT\-4o to conduct structured inferences based on visual layouts and implicit planning constraints\.

Following closely were InternVL3\-9B \(1\.331\) and Qwen2\.5\-VL\-72B\-AWQ \(1\.306\)\. InternVL3\-9B performed particularly strongly in Domain\-Specific Reasoning \(1\.435\) and Classification \(1\.297\), while maintaining solid Spatial Reasoning \(1\.260\)\. Qwen2\.5\-VL\-72B\-AWQ achieved the highest Classification score \(1\.406, tied with GPT\-4o\) and exhibited balanced performance across all three subtasks, although its Spatial Reasoning \(1\.248\) lagged slightly behind that of InternVL3\-9B\.

In contrast, the weakest performer was Qwen2\-VL\-2B, with a Reasoning score of 0\.701, followed by Qwen2\.5\-VL\-3B \(0\.838\)\. Qwen2\.5\-VL\-3B struggled most notably in Spatial Reasoning \(0\.691\), indicating fundamental challenges with interpreting spatial relationships in complex urban layouts\. Qwen2\-VL\-7B also showed limitations with an overall score of 0\.897, particularly in Spatial Reasoning \(0\.716\) and Classification \(1\.031\)\.

The reasoning tasks revealed that even high\-scoring models encountered difficulty with multi\-step logic and condition\-based classification\. Qwen2\.5\-VL\-72B\-Instruct\-AWQ frequently produced overgeneralized answers or missed key spatial and policy conditions, leading to incorrect or superficial classifications\.

#### 5\.3\.4Evaluation on Association

TheAssociationdimension focuses on the ability of a model to link visual elements and spatial design with relevant policy knowledge, specifically evaluated through the Policy Association task\. This dimension tests whether a model can identify and align visual evidence with appropriate regulatory frameworks and planning categories, a critical capability in professional design evaluation\.

The best performer in this dimension was GPT\-4o, achieving a score of 1\.527\. It was closely followed by InternVL3\-9B \(1\.297\) and Qwen2\.5\-VL\-72B\-AWQ \(1\.253\)\. These models demonstrated the capacity to cite relevant planning documents, recognize administrative hierarchies such as national versus local regulations, and align visual cues with textual policy language\. For instance, in urban renewal scenarios, these models correctly referenced floor\-area\-ratio regulations or ecological redline zones based on spatial overlays in the input images\.

Despite this, even the strongest models exhibited limited policy awareness in terms of structural specificity\. Several answers listed relevant policies but failed to clarify their jurisdictional scope or functional classification\. A recurring issue was category confusion, in which local planning guidelines, special\-purpose plans, and national strategic frameworks were intermingled without proper distinction\. This lack of differentiation weakens the reliability of the association chain and limits its applicability in scenario\-based reasoning\.

The lowest\-scoring models in this dimension were InternVL3\-8B and Qwen2\-VL\-2B \(both 0\.926\), followed by GPT\-4o\-mini \(0\.963\) and Qwen2\.5\-VL\-3B \(0\.970\)\. These models frequently hallucinated plausible\-sounding regulations with no basis in the image or failed to distinguish between visually similar zones, for instance confusing green buffer zones with agricultural protection areas\. In addition, they often emphasized advantages such as innovation or aesthetics but neglected to identify regulatory risks or constraints, reflecting a one\-sided interpretation of planning implications\.

Association tasks required models to connect visual evidence with planning policies\. Despite high surface\-level fluency, models lackedpolicy awarenessandstructural coherence: Qwen2\.5\-VL\-72B, for instance, often listed multiple policy items but failed to distinguishadministrative levelsorplanning categories, resulting in conceptual overlaps and ambiguity\. Models also tended to highlight strengths such as innovation or leadership while neglecting limitations, and produced syntactically coherent but unsupported reasoning chains\.

#### 5\.3\.5Evaluation on Implementation

TheImplementationdimension assesses how well models can generate actionable feedback and evaluation statements within professional design contexts\. It comprises two tasks:Scheme EvaluationandDecision\-Making, each requiring models to analyze spatial proposals and produce concise, domain\-appropriate responses aligned with planning logic\.

The best performer was GPT\-4o, achieving 1\.223 in Scheme Evaluation and 1\.429 in Decision\-Making\. Its answers were well\-structured, referenced key urban design metrics such as density and accessibility, and reflected trade\-offs among aesthetics, ecology, and policy feasibility\.

Following closely were Qwen2\.5\-VL\-72B\-AWQ \(1\.153 and 1\.090\) and InternVL3\-9B \(0\.921 and 0\.903\)\. Qwen2\.5\-VL\-72B\-AWQ demonstrated balanced performance across both tasks, while InternVL3\-9B exhibited particular strength in Scheme Evaluation but relatively weaker capability in Decision\-Making\.

Nevertheless, common issues persisted across models\. Qwen2\.5\-VL\-72B\-Instruct, despite its large parameter count, produced verbose and unfocused responses that often lacked prioritization or avoided value judgments altogether\. Many models failed to identify constraints such as zoning incompatibilities or ecological conflicts—central to real\-world planning evaluations—and Qwen2\-VL\-2B recorded the lowest overall average \(0\.731\), consistently failing to deliver actionable, terminology\-grounded advice\.

#### 5\.3\.6Second Round \(2026\)

The second round evaluated six models that incorporate enhanced agentic reasoning capabilities\. Results are shown in the lower panel of Table[2](https://arxiv.org/html/2606.05744#S5.T2)\. The standout performer is Qwen3\.6\-Plus, which achieves an Overall score of 1\.701, decisively surpassing all other models in both generations\. Its dominance is broad: it ranks first in every task category, with particularly commanding leads in Element Recognition \(1\.751 vs\. the next\-best Kimi\-K2\.6 at 1\.518\) and Scheme Evaluation \(1\.619 vs\. the next\-best GPT\-5\.4 at 1\.586\)\.

The remaining 2026 models cluster within a narrower band \(1\.353–1\.472\)\. Gemini\-2\.5\-Pro \(1\.472\) and GPT\-5\.4 \(1\.431\) form the second tier\. Gemini\-2\.5\-Pro excels in Classification \(1\.656\), while GPT\-5\.4 achieves the second\-highest Caption score \(1\.900, after Qwen3\.6\-Plus at 1\.950\) and the second\-highest Scheme Evaluation score \(1\.586\), but is held back by comparatively weak Element Recognition \(1\.233\)\. Kimi\-K2\.6 \(1\.417\) is notable for strong Element Recognition \(1\.518\)\. Claude\-Opus\-4\.7 \(1\.384\) presents a polarized profile: strong on Domain\-Specific Reasoning \(1\.558\) and Policy Association \(1\.493\), yet the weakest in Element Recognition \(1\.186\)\. At the lower end, Qwen3\.6\-Flash \(1\.353\) performs creditably for a lightweight model, but its Scheme Evaluation score \(1\.046\) is the weakest among all 2026 models\.

#### 5\.3\.7Cross\-Generation Comparison

Comparing the two rounds reveals a clear step\-change in capability\. The best 2026 model \(Qwen3\.6\-Plus, 1\.701\) outperforms the best 2025 model \(GPT\-4o, 1\.342\) by a margin of 0\.359, an improvement of approximately 27%\. Strikingly, even the lightest 2026 model, Qwen3\.6\-Flash \(1\.353\), exceeds the top 2025 score\. This pattern confirms that the agentic reasoning enhancements present in the 2026 generation yield a substantive and uniform upward shift across the entire capability spectrum\.

The gains are not uniform across dimensions\. The most dramatic improvement occurs in Scheme Evaluation, where the 2026 generation mean \(1\.427\) is 77\.7% higher than that of the 2025 generation \(0\.803\), reflecting the capacity of agentic models to structure multi\-criteria judgments and produce actionable planning critiques\. Spatial Reasoning \(0\.883→\\rightarrow1\.426, \+0\.543\) and Decision\-making \(0\.969→\\rightarrow1\.470, \+0\.501\) also show large absolute gains\. Domain\-Specific Reasoning \(1\.053→\\rightarrow1\.511, \+0\.458\) and Element Recognition \(0\.975→\\rightarrow1\.402, \+0\.427\) improve substantially, while Policy Association \(1\.102→\\rightarrow1\.456, \+0\.354\), Classification \(1\.134→\\rightarrow1\.475, \+0\.341\), and Captioning \(1\.518→\\rightarrow1\.767, \+0\.249\) show smaller absolute gains\.

### 5\.4Image Resolution Sensitivity Analysis

A natural concern when evaluating VLMs on planning maps is whether variation in image resolution systematically biases model scores, given that the benchmark draws images from heterogeneous sources whose total pixel counts vary by roughly 46\-fold\. To investigate this, we conducted a resolution sensitivity analysis on the evaluation subset, which contains 63 unique images with resolutions ranging from 352×\\times373 to 2,057×\\times2,953 pixels\.

For this analysis, we scored the subset across three representative models \(Claude Opus 4\.7, Gemini 2\.5 Pro, and GPT\-5\.4\) and computed Spearman rank correlations between per\-item scores and six resolution metrics: log10\(total pixels\), total pixels,pixels\\sqrt\{\\text\{pixels\}\}, width, height, and aspect ratio\. As shown in Figure[4](https://arxiv.org/html/2606.05744#S5.F4), none of the models exhibited a meaningful correlation between image resolution and score\. The strongest per\-model correlation was observed for Claude Opus 4\.7 on width \(ρ=\+0\.057\\rho=\+0\.057,p=0\.329p=0\.329\), and all log10\(pixels\) correlations remained negligible \(\|ρ\|<0\.06\|\\rho\|<0\.06,p\>0\.30p\>0\.30across all three models\)\. We further partitioned the images into five quantile bins by log10\(pixels\) and conducted one\-way ANOVA tests for each model; no model showed significant between\-bin score differences \(allp\>0\.05p\>0\.05\)\. Aggregating scores to the image level yielded a Spearmanρ\\rhoof−0\.222\-0\.222\(p=0\.080p=0\.080\)\.

![Refer to caption](https://arxiv.org/html/2606.05744v1/images/boxplot_by_resolution_bin.png)Figure 4:Image resolution sensitivity analysis\. Boxplots of per\-model scores aggregated into five quantile resolution bins\. No meaningful or statistically significant difference is observed across bins for any model \(ANOVA, allp\>0\.05p\>0\.05\)\.These results collectively indicate that image resolution does not act as a confounding factor in the evaluation of PlanBench\-V\. The natural variation in map image dimensions within the benchmark is sufficiently narrow and uncorrelated with task difficulty that it does not introduce systematic bias into model rankings\. This finding supports the validity of cross\-model comparisons reported in Section[5\.3](https://arxiv.org/html/2606.05744#S5.SS3)without requiring resolution normalization\.

## 6Conclusion

This study presents PlanBench\-V, the first comprehensive benchmark for evaluating the capabilities of VLMs in spatial planning map interpretation\. We constructed the SPMD with 223 planning maps and 1,629 expert\-annotated QA pairs, organized across four progressive dimensions—Perception, Reasoning, Association, and Implementation—that reflect the full cognitive pipeline of planning practice\. Through two rounds of evaluation spanning two generations of models, we characterized both the current state of VLM capability and the trajectory of progress\. The best 2025 model, GPT\-4o, achieved an Overall score of 1\.342; the best 2026 agentic reasoning model, Qwen3\.6\-Plus, improved this to 1\.701, a margin of 0\.359 or approximately 27%\. Strikingly, even the lightest 2026 model exceeds the best 2025 score, confirming that agentic reasoning yields a substantive generation\-level uplift\.

However, despite this progress, implementation\-oriented tasks remain a central challenge\. Even the best 2026 models can still struggle to generate concise, actionable, and well\-structured evaluations, often producing verbose outputs with limited domain specificity\. The absolute gains are markedly uneven: Scheme Evaluation, Spatial Reasoning, Decision\-making, Domain\-Specific Reasoning, and Element Recognition show the largest leaps, while Captioning, Classification, and Policy Association improve less\. This pattern suggests that general\-purpose reasoning enhancements alone cannot compensate for the lack of structured domain knowledge in planning\-specific tasks\.

A plausible explanation is that implementation\-oriented tasks impose a substantially higher cognitive burden than conventional VQA problems\. They demand long\-horizon reasoning within a single query, requiring the simultaneous integration of visual perception, spatial relationship understanding, domain\-specific planning knowledge, and policy\-context constraints\. The resulting challenge arises not from the number of questions but from the depth of reasoning and the heterogeneity of information that must be coordinated within a continuous inference process\. Even the strongest agentic models exhibit characteristic failure patterns, including unfocused verbosity, drifting attention, insufficient policy grounding, and inconsistencies across reasoning stages\.

These observations confirm that the performance bottleneck in planning map interpretation is not merely a consequence of missing domain knowledge\. Rather, it reflects a fundamental mismatch between prevailing vision\-language learning paradigms and the cognitive demands of professional planning practice, which requires normative judgment under legal frameworks, hierarchical governance structures, and competing public objectives\.

Future research should therefore pursue domain\-adaptive multimodal learning strategies that explicitly encode planning logic, regulatory constraints, and spatial decision\-making processes\. Integrating structured policy knowledge, constraint\-aware reasoning mechanisms, and theory\-informed supervision may prove essential for bridging the gap between general\-purpose multimodal intelligence and expert\-level urban planning cognition\. By articulating these challenges across two model generations, PlanBench\-V provides not only an evaluation tool but also a conceptual foundation for advancing vision\-language models toward more reliable and responsible applications in spatial governance\.

## References

- A\. Agrawal, J\. Lu, S\. Antol, M\. Mitchell, C\. L\. Zitnick, D\. Batra, and D\. Parikh \(2016\)VQA: Visual Question Answering\.arXiv\.Note:arXiv:1505\.00468 \[cs\]External Links:[Link](http://arxiv.org/abs/1505.00468),[Document](https://dx.doi.org/10.48550/arXiv.1505.00468)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- L\. Albrechts, P\. Healey, and K\. R\. Kunzmann \(2003\)Strategic spatial planning and regional governance in europe\.Journal of the American Planning Association69\(2\),pp\. 113–129\.External Links:[Document](https://dx.doi.org/10.1080/01944360308976301),ISSN 0194\-4363Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p2.1)\.
- C\. Alexander \(1977\)A pattern language: towns, buildings, construction\.Oxford University Press\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- R\. Arnheim \(1969\)Visual thinking\.University of California Press\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- J\. Bai, S\. Bai, S\. Yang, S\. Wang, S\. Tan, P\. Wang, J\. Lin, C\. Zhou, and J\. Zhou \(2023\)Qwen\-vl: a versatile vision\-language model for understanding, localization, text reading, and beyond\.External Links:2308\.12966Cited by:[§5\.1](https://arxiv.org/html/2606.05744#S5.SS1.p1.1)\.
- J\. Bertin \(1983\)Semiology of graphics: diagrams, networks, maps\.University of Wisconsin Press\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p3.1)\.
- R\. Burgess and M\. I\. Carmona \(2009\)The shift from master planning to strategic planning\.InPlanning through projects: Moving from master planning to strategic planning \- 30 cities,M\. I\. Carmona, R\. Burgess, and M\. S\. Badenhorst \(Eds\.\),pp\. 12–42\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p2.1)\.
- S\. Chang, D\. Palzer, J\. Li, E\. Fosler\-Lussier, and N\. Xiao \(2022\)MapQA: a dataset for question answering on choropleth maps\.External Links:2211\.08545,[Link](https://arxiv.org/abs/2211.08545)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p2.1)\.
- S\. Chen, J\. Zhang, T\. Zhu, W\. Liu, S\. Gao, M\. Xiong, M\. Li, and J\. He \(2025\)Bring reason to vision: understanding perception and reasoning through model merging\.arXiv preprint arXiv:2505\.05464\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- Z\. Chen, J\. Wu, W\. Wang, W\. Su, G\. Chen, S\. Xing, M\. Zhong, Q\. Zhang, X\. Zhu, L\. Lu, B\. Li, P\. Luo, T\. Lu, Y\. Qiao, and J\. Dai \(2023\)InternVL: scaling up vision foundation models and aligning for generic visual\-linguistic tasks\.arXiv preprint arXiv:2312\.14238\.Cited by:[§5\.1](https://arxiv.org/html/2606.05744#S5.SS1.p1.1)\.
- G\. DeepMind \(2023\)Gemini\.Note:\\urlhttps://gemini\.google\.comCited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- A\. C\. Doris, D\. Grandi, R\. Tomich, M\. F\. Alam, M\. Ataei, H\. Cheong, and F\. Ahmed \(2024\)DesignQA: A Multimodal Benchmark for Evaluating Large Language Models’ Understanding of Engineering Documentation\.arXiv\.External Links:2404\.07917,[Document](https://dx.doi.org/10.48550/arXiv.2404.07917)Cited by:[§4](https://arxiv.org/html/2606.05744#S4.p1.1)\.
- S\. Dühr \(2007\)The visual language of spatial planning: exploring cartographic representations for spatial planning in europe\.1st edition,Routledge,London\.External Links:[Document](https://dx.doi.org/10.4324/9780203965818),ISBN 9780203965818,[Link](https://doi.org/10.4324/9780203965818)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p3.1)\.
- S\. Dühr \(2009\)Visualising spatial policy in europe\.InPlanning Cultures in Europe,F\. Othengrafen and J\. Knieling \(Eds\.\),pp\. 24\.External Links:ISBN 9781315246727,[Document](https://dx.doi.org/10.4324/9781315246727)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p2.1)\.
- A\. Faludi \(2000\)The european spatial development perspective–what next?\.European Planning Studies8\(2\),pp\. 237–250\.External Links:[Document](https://dx.doi.org/10.1080/713666411)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p2.1)\.
- D\. Guo, F\. Wu, F\. Zhu, F\. Leng, G\. Shi, H\. Chen, H\. Fan, J\. Wang, J\. Jiang, J\. Wang,et al\.\(2025\)Seed1\. 5\-vl technical report\.arXiv preprint arXiv:2505\.07062\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- J\. B\. Harley \(1989\)Deconstructing the map\.Cartographica: The International Journal for Geographic Information and Geovisualization26\(2\),pp\. 1–20\.External Links:[Document](https://dx.doi.org/10.3138/J674-7727-1757-1412)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p1.1)\.
- J\. B\. Harley \(2002\)The new nature of maps: essays in the history of cartography\.Johns Hopkins University Press,Baltimore\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p1.1)\.
- P\. Healey \(1997\)Collaborative planning: shaping places in fragmented societies\.UBC Press\.Cited by:[§1](https://arxiv.org/html/2606.05744#S1.p1.1)\.
- Y\. Huang, Q\. Yuan, X\. Sheng, Z\. Yang, H\. Wu, P\. Chen, Y\. Yang, L\. Li, and W\. Lin \(2024\)Aesbench: an expert benchmark for multimodal large language models on image aesthetics perception\.arXiv preprint arXiv:2401\.08276\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.05744#S4.p1.1)\.
- A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.\(2024\)Gpt\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- K\. Koffka \(1935\)Principles of gestalt psychology\.Harcourt, Brace and Company\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- G\. Kress and T\. van Leeuwen \(2006\)Reading images: the grammar of visual design\.2nd edition,Routledge\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- R\. Krishna, Y\. Zhu, O\. Groth, J\. Johnson, K\. Hata, J\. Kravitz, S\. Chen, Y\. Kalantidis, L\. Li, D\. A\. Shamma, M\. S\. Bernstein, and L\. Fei\-Fei \(2017\)Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations\.International Journal of Computer Vision123\(1\),pp\. 32–73\.External Links:ISSN 1573\-1405,[Link](https://doi.org/10.1007/s11263-016-0981-7),[Document](https://dx.doi.org/10.1007/s11263-016-0981-7)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- Y\. Lai, J\. Zhong, M\. Li, S\. Zhao, and X\. Yang \(2025\)Med\-r1: reinforcement learning for generalizable medical reasoning in vision\-language models\.arXiv preprint arXiv:2503\.13939\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- C\. Li, C\. Wong, S\. Zhang, N\. Usuyama, H\. Liu, J\. Yang, T\. Naumann, H\. Poon, and J\. Gao \(2023\)Llava\-med: training a large language\-and\-vision assistant for biomedicine in one day\.Advances in Neural Information Processing Systems36,pp\. 28541–28564\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- S\. Li and N\. Tajbakhsh \(2023\)SciGraphQA: a large\-scale synthetic multi\-turn question\-answering dataset for scientific graphs\.External Links:2308\.03349,[Link](https://arxiv.org/abs/2308.03349)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- Z\. Li, Y\. Lin, P\. Hooimeijer, J\. Monstadt, and J\. He \(2025\)The communicative turn in planning? examining community planner’s role as a third actor in beijing, china\.Cities159,pp\. 105785\.External Links:ISSN 0264\-2751,[Document](https://dx.doi.org/10.1016/j.cities.2025.105785),[Link](https://www.sciencedirect.com/science/article/pii/S026427512500085X)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p2.1)\.
- J\. Lin, D\. Huang, T\. Zhao, D\. Zhan, and C\. Lin \(2024\)Designprobe: a graphic design benchmark for multimodal large language models\.arXiv preprint arXiv:2404\.14801\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.05744#S4.p1.1)\.
- T\. Lin, M\. Maire, S\. Belongie, L\. Bourdev, R\. Girshick, J\. Hays, P\. Perona, D\. Ramanan, C\. L\. Zitnick, and P\. Dollár \(2015\)Microsoft COCO: Common Objects in Context\.arXiv\.Note:arXiv:1405\.0312 \[cs\]External Links:[Link](http://arxiv.org/abs/1405.0312),[Document](https://dx.doi.org/10.48550/arXiv.1405.0312)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee \(2023\)Visual instruction tuning\.External Links:2304\.08485,[Link](https://arxiv.org/abs/2304.08485)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- K\. Lynch \(1984\)Good city form\.MIT press\.Cited by:[§1](https://arxiv.org/html/2606.05744#S1.p1.1)\.
- M\. Mathew, R\. Tito, D\. Karatzas, R\. Manmatha, and C\. V\. Jawahar \(2021\)Document visual question answering challenge 2020\.External Links:2008\.08899,[Link](https://arxiv.org/abs/2008.08899)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- R\. E\. Mayer \(2005\)The cambridge handbook of multimedia learning\.Cambridge University Press\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- S\. Moroni and G\. Lorini \(2017\)Graphic rules in planning: a critical exploration of normative drawings starting from zoning maps and form\-based codes\.Planning Theory16\(3\),pp\. 318–338\.External Links:[Document](https://dx.doi.org/10.1177/1473095216672342)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p3.1)\.
- S\. E\. Palmer \(1999\)Vision science: photons to phenomenology\.MIT Press\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- J\. Pan, C\. Liu, J\. Wu, F\. Liu, J\. Zhu, H\. B\. Li, C\. Chen, C\. Ouyang, and D\. Rueckert \(2025\)Medvlm\-r1: incentivizing medical reasoning capability of vision\-language models \(vlms\) via reinforcement learning\.arXiv preprint arXiv:2502\.19634\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- T\. Qian, J\. Chen, L\. Zhuo, Y\. Jiao, and Y\. Jiang \(2024\)Nuscenes\-qa: a multi\-modal visual question answering benchmark for autonomous driving scenario\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 4542–4550\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- V\. Rautenbach, S\. Coetzee, and A\. Çöltekin \(2017\)Development and evaluation of a specialized task taxonomy for spatial planning – A map literacy experiment with topographic maps\.ISPRS Journal of Photogrammetry and Remote Sensing127,pp\. 16–26\.External Links:[Document](https://dx.doi.org/10.1016/j.isprsjprs.2016.06.013)Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1),[§4](https://arxiv.org/html/2606.05744#S4.p3.1)\.
- J\. Roberts, K\. Han, N\. Houlsby, and S\. Albanie \(2024\)SciFIBench: benchmarking large multimodal models for scientific figure interpretation\.External Links:2405\.08807,[Link](https://arxiv.org/abs/2405.08807)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- J\. Roberts, T\. Lüddecke, R\. Sheikh, K\. Han, and S\. Albanie \(2023\)Charting New Territories: Exploring the geographic and geospatial capabilities of multimodal LLMs\.arXiv preprint arXiv:2311\.14656\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p2.1)\.
- H\. Shen, P\. Liu, J\. Li, C\. Fang, Y\. Ma, J\. Liao, Q\. Shen, Z\. Zhang, K\. Zhao, Q\. Zhang,et al\.\(2025\)Vlm\-r1: a stable and generalizable r1\-style large vision\-language model\.arXiv preprint arXiv:2504\.07615\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- C\. Sima, K\. Renz, K\. Chitta, L\. Chen, H\. Zhang, C\. Xie, J\. Beißwenger, P\. Luo, A\. Geiger, and H\. Li \(2024\)Drivelm: driving with graph visual question answering\.InEuropean Conference on Computer Vision,pp\. 256–274\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- C\. Steinitz \(1995\)A framework for theory and practice in landscape planning\.Process architecture127,pp\. 12–31\.Cited by:[§1](https://arxiv.org/html/2606.05744#S1.p1.1)\.
- P\. Wang, S\. Bai, S\. Tan, S\. Wang, Z\. Fan, J\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, Y\. Fan, K\. Dang, M\. Du, X\. Ren, R\. Men, D\. Liu, C\. Zhou, J\. Zhou, and J\. Lin \(2024a\)Qwen2\-vl: enhancing vision\-language model’s perception of the world at any resolution\.External Links:2409\.12191,[Link](https://arxiv.org/abs/2409.12191)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- W\. Wang, S\. Zhang, Y\. Ren, Y\. Duan, T\. Li, S\. Liu, M\. Hu, Z\. Chen, K\. Zhang, L\. Lu, X\. Zhu, P\. Luo, Y\. Qiao, J\. Dai, W\. Shao, and W\. Wang \(2024b\)Needle In A Multimodal Haystack\.arXiv\.Note:arXiv:2406\.07230 \[cs\]External Links:[Link](http://arxiv.org/abs/2406.07230),[Document](https://dx.doi.org/10.48550/arXiv.2406.07230)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- M\. Wertheimer \(1938\)Gestalt theory\.Social Research5\(1\),pp\. 81–121\.Cited by:[§2\.1](https://arxiv.org/html/2606.05744#S2.SS1.p4.1)\.
- X\. Yue, Y\. Ben\-David, S\. Nova, X\. Shi, K\. K\. Lin, B\. Al\-Ghamdi, P\. Kim, R\. Roggener, D\. Glickman, S\. Lim, R\. Pryzant, W\. Yih, Y\. Carmon, W\. Wang, and C\.J\. R\. Shi \(2024\)MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert agi\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),External Links:2311\.16502Cited by:[§5\.1](https://arxiv.org/html/2606.05744#S5.SS1.p2.1)\.
- Y\. Zhang, Z\. He, J\. Li, J\. Lin, Q\. Guan, and W\. Yu \(2024a\)MapGPT: an autonomous framework for mapping by integrating large language model and cartographic tools\.Cartography and Geographic Information Science51\(6\),pp\. 717–743\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- Y\. Zhang, Z\. Wang, Z\. He, J\. Li, G\. Mai, J\. Lin, C\. Wei, and W\. Yu \(2024b\)BB\-geogpt: a framework for learning a large language model for geographic information science\.Information Processing & Management61\(5\),pp\. 103808\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- Z\. Zhou, Q\. Wang, B\. Lin, Y\. Su, R\. Chen, X\. Tao, A\. Zheng, L\. Yuan, P\. Wan, and D\. Zhang \(2024\)Uniaa: a unified multi\-modal image aesthetic assessment baseline and benchmark\.arXiv preprint arXiv:2404\.09619\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.
- H\. Zhu, G\. Chen, and W\. Zhang \(2025a\)PlanGPT: enhancing urban planning with a tailored agent framework\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 6: Industry Track\),G\. Rehm and Y\. Li \(Eds\.\),Vienna, Austria,pp\. 764–783\.External Links:[Link](https://aclanthology.org/2025.acl-industry.54/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-industry.54),ISBN 979\-8\-89176\-288\-6Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p3.1)\.
- H\. Zhu, J\. Su, M\. Chen, W\. Wang, Y\. Deng, G\. Chen, and W\. Zhang \(2025b\)PlanGPT\-vl: enhancing urban planning with domain\-specific vision\-language models\.External Links:2505\.14481,[Link](https://arxiv.org/abs/2505.14481)Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p3.1),[§3](https://arxiv.org/html/2606.05744#S3.p1.1)\.
- J\. Zhu, W\. Wang, Z\. Chen, Z\. Liu, S\. Ye, L\. Gu, Y\. Duan, H\. Tian, W\. Su, J\. Shao,et al\.\(2025c\)InternVL3: exploring advanced training and test\-time recipes for open\-source multimodal models\.arXiv preprint arXiv:2504\.10479\.Cited by:[§2\.2](https://arxiv.org/html/2606.05744#S2.SS2.p1.1)\.

Similar Articles

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

arXiv cs.AI

PlanningBench is a framework for generating scalable, diverse, and verifiable planning data to evaluate and train large language models, featuring a constraint-driven synthesis pipeline with adaptive difficulty control and quality filtering. Experiments show that frontier LLMs struggle with coupled constraints, and reinforcement learning on PlanningBench data improves performance on unseen planning tasks.