PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Summary
PerceptionBench is a benchmark designed to evaluate atomic visual perception capabilities of Multimodal Large Language Models (MLLMs), using a bottom-up taxonomy of ten atomic perceptual capabilities. Results across 16 frontier MLLMs show no model reaches 60% accuracy, indicating visual perception remains largely unsolved.
View Cached Full Text
Cached at: 07/29/26, 03:50 AM
Paper page - PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Source: https://huggingface.co/papers/2607.24957 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
WeintroducePerceptionBench,abenchmarkspecificallydesignedtoevaluatetheatomicvisualperceptioncapabilitiesofMultimodalLargeLanguageModels(MLLMs).Existingbenchmarksoftenfailtoisolateperception:holisticevaluationsconflateperceptualerrorswithfailuresinreasoningordomainknowledge,whileapplication-drivenbenchmarksonlycovernarrow,fragmenteddomainsshapedbyheuristicdesigns.Toaddresstheselimitations,PerceptionBenchadoptsabottom-upapproach:bydiagnosingtheearliestfailurepointsintheresponsesoffrontierMLLMsacross42existingbenchmarks,weconstructanerrortaxonomywhoseperceptionbranchdefinestenatomicperceptualcapabilities.Guidedbythistaxonomy,weconstruct3,000verifiedquestionswithshort,unambiguousanswers,eachisolatingasinglecapability,withdifficultystemmingfromperceptionratherthanreasoningorknowledge.BenchmarkresultsacrosssixteenfrontierMLLMsrevealthatatomicperceptionremainslargelyunsolved---nomodelreaches60\%accuracy,perception-relatedhallucinationistheweakestcapabilityonaverage,andsimilaroverallscoresconcealsharplydivergentcapabilityprofiles.PerceptionBenchthusprovidesacapability-levelstandardformeasuringanddiagnosingthevisualperceptionboundariesofMLLMs.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.24957
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.24957 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.24957 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.24957 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Can Multimodal Large Language Models Understand OCT?
This paper introduces OCT-Bench, a comprehensive benchmark for evaluating multimodal large language models on OCT image understanding, comprising 10,076 questions across 20 fine-grained tasks in perception, cognition, and reasoning. Evaluation of 20 models shows current MLLMs fall short of reliable OCT understanding.
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
MuseBench is a comprehensive benchmark introduced to evaluate multimodal large language models on nuanced, intent-level understanding of audiovisual arts, revealing that even the best model achieves only 48.29% accuracy compared to 87.18% for human experts.
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Artifact-Bench is a comprehensive benchmark that evaluates multimodal large language models on detecting and analyzing artifacts in AI-generated videos, revealing significant limitations and misalignment with human perception.
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models
MCBench is a new benchmark for assessing the safety of omnimodal large language models across vision, audio, and text modalities. It includes 1196 scenarios and finds current models struggle with cross-modal safety reasoning.
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Introduces WorldBench, a visually diverse multimodal reasoning benchmark that reveals significant limitations in current multimodal large language models' visual understanding.