WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Summary
Introduces WorldBench, a visually diverse multimodal reasoning benchmark that reveals significant limitations in current multimodal large language models' visual understanding.
View Cached Full Text
Cached at: 06/08/26, 03:29 AM
Paper page - WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Source: https://huggingface.co/papers/2606.06538 Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
WorldBench is introduced as a visually diverse reasoning benchmark for evaluating multimodal large language models, revealing significant limitations in current models’ visual understanding capabilities.
In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existingmultimodal benchmarksexpand task types without capturing thevisual diversityneeded to handle open-ended visual inputs. We present WorldBench, a challenging and visually diversereasoning benchmarkto evaluateMultimodal Large Language Models(MLLMs). We build a taxonomy of thousands ofvisual conceptsacross multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves highervisual diversitythan any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance ofvisual diversityin buildingmultimodal benchmarks.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.06538 in a model README.md to link it from this page.
Datasets citing this paper1
#### zlab-princeton/WorldBench Viewer• Updatedabout 2 hours ago • 2k • 21
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.06538 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Introduces MultivationBench, a benchmark for evaluating multimodal large language models' sequential motivation reasoning using story-driven visual narratives based on Maslow's hierarchy and Reiss's desires. Results show all tested models struggle with dynamic motivation inference.
WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild
WildTableBench introduces the first question-answering benchmark for real-world table images, revealing that existing multimodal foundation models struggle significantly with structural perception and numerical reasoning, with only one model exceeding 50% accuracy.
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
BEAR-Bench introduces a bilingual English-and-Russian benchmark for evaluating multimodal models' reasoning on text-rich professional documents, assessing 16 models and highlighting performance gaps.
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
This paper introduces The Unwritten Benchmark, a new challenge to evaluate abstract perceptual reasoning in multimodal AI models, revealing a significant performance gap between humans and current models like GPT-4o and Gemini 2.5-Pro.
M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.