VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Summary
Introduces VIABench, a comprehensive video benchmark for evaluating multimodal large language models in real-world visual assistance for blind and visually impaired individuals, covering 761 videos and 14,526 annotations across three tasks.
View Cached Full Text
Cached at: 07/20/26, 09:42 AM
Paper page - VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Source: https://huggingface.co/papers/2607.14660 We introduceVIABench, a comprehensive, time-aligned video benchmark for evaluating multimodal large language models (MLLMs) in real-world visual assistance scenarios for blind and visually impaired people.
VIABench contains761 videos,14,526 manually curated annotations, and46.9 hours of footage. It covers three complementary tasks:Proactive Reminder,Visual Question Answering, andVision-Guided Interaction. We also proposeToken-Level Prompt Activation Decoding (TPAD), a two-stage framework for evaluating proactive assistance in both online and offline settings.
Our evaluation shows that current MLLMs still struggle with reliable real-world assistance, especially in anticipating navigation-critical events and responding in real time. We hope VIABench encourages progress toward safer and more useful visual assistants.
Code and data:https://github.com/MCG-NJU/VIABench
Similar Articles
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Introduces WorldBench, a visually diverse multimodal reasoning benchmark that reveals significant limitations in current multimodal large language models' visual understanding.
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.
Benchmarking Visual State Tracking in Multimodal Video Understanding
Introduces VSTAT, a benchmark for evaluating visual state tracking in multimodal large language models (MLLMs) using 834 clips and 1,500 questions. Current MLLMs perform poorly compared to humans, failing at visual perception rather than reasoning.
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
MSAVBench is the first comprehensive benchmark and adaptive evaluation framework for multi-shot audio-video generation, assessing 19 models across diverse tasks and achieving high alignment with human judgment.
Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
BloomBench is a cognitively grounded bilingual (English-Arabic) multimodal benchmark for Vision-Language Models, systematically evaluating six cognitive levels based on Bloom's Taxonomy. Experiments reveal significant cognitive asymmetries and cross-lingual performance gaps in current models.