WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Hugging Face Daily Papers Papers

Summary

Introduces WorldBench, a visually diverse multimodal reasoning benchmark that reveals significant limitations in current multimodal large language models' visual understanding.

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves higher visual diversity than any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.
Original Article
View Cached Full Text

Cached at: 06/08/26, 03:29 AM

Paper page - WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Source: https://huggingface.co/papers/2606.06538 Authors:

,

,

,

,

,

,

,

,

,

,

Abstract

WorldBench is introduced as a visually diverse reasoning benchmark for evaluating multimodal large language models, revealing significant limitations in current models’ visual understanding capabilities.

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existingmultimodal benchmarksexpand task types without capturing thevisual diversityneeded to handle open-ended visual inputs. We present WorldBench, a challenging and visually diversereasoning benchmarkto evaluateMultimodal Large Language Models(MLLMs). We build a taxonomy of thousands ofvisual conceptsacross multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves highervisual diversitythan any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance ofvisual diversityin buildingmultimodal benchmarks.

View arXiv pageView PDFProject pageAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2606.06538 in a model README.md to link it from this page.

Datasets citing this paper1

#### zlab-princeton/WorldBench Viewer• Updatedabout 2 hours ago • 2k • 21

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.06538 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles