OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Hugging Face Daily Papers 05/18/26, 12:00 AM Papers

benchmark video-understanding multimodal proactive streaming-video large-language-models

Summary

OmniPro is the first benchmark for evaluating proactive streaming video understanding in omni-modal large language models, featuring 2,700 samples covering diverse tasks and dual-mode evaluation protocols.

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed-timestamp protocols instead of true proactive evaluation, and cover only a limited range of tasks, preventing reliable assessment and differentiation of omni-proactive streaming models. We present OmniPro, the first benchmark to jointly evaluate omni-modal perception, proactive responding, and diverse video understanding tasks. It comprises 2,700 human-verified samples spanning 9 sub-tasks and 3 cognitive levels, covering 6 basic video understanding capabilities. Notably, 84% of samples require audio signals (speech or non-speech), and each sample is annotated with modality-isolation labels to enable fine-grained multimodal analysis. We further introduce a dual-mode evaluation protocol: Probe mode assesses content understanding by querying the model before and after each ground-truth trigger, while Online mode evaluates full proactive ability by requiring models to autonomously decide when to respond in streaming input. Evaluating 11 representative models reveals three key findings: (1) audio provides consistent gains but with highly variable utilization across models, (2) performance degrades significantly over time, indicating limited long-horizon robustness, and (3) non-speech audio perception remains the weakest dimension.

Original Article

View Cached Full Text

Cached at: 05/22/26, 10:19 AM

Paper page - OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Source: https://huggingface.co/papers/2605.18577

Abstract

OmniPro is introduced as the first benchmark for evaluating omni-modal large language models’ proactive streaming video understanding, featuring diverse tasks and dual-mode evaluation protocols.

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability ofomni-modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed-timestamp protocols instead of true proactive evaluation, and cover only a limited range of tasks, preventing reliable assessment and differentiation of omni-proactive streaming models. We present OmniPro, the first benchmark to jointly evaluate omni-modal perception, proactive responding, and diverse video understanding tasks. It comprises 2,700 human-verified samples spanning 9 sub-tasks and 3 cognitive levels, covering 6 basic video understanding capabilities. Notably, 84% of samples require audio signals (speech or non-speech), and each sample is annotated with modality-isolation labels to enable fine-grainedmultimodal analysis. We further introduce adual-mode evaluation protocol:Probe modeassesses content understanding by querying the model before and after each ground-truth trigger, whileOnline modeevaluates full proactive ability by requiring models to autonomously decide when to respond in streaming input. Evaluating 11 representative models reveals three key findings: (1) audio provides consistent gains but with highly variable utilization across models, (2) performance degrades significantly over time, indicating limited long-horizon robustness, and (3) non-speech audio perception remains the weakest dimension.

View arXiv page View PDF Project page GitHub5 Add to collection

Get this paper in your agent:

hf papers read 2605\.18577

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.18577 in a model README.md to link it from this page.

Datasets citing this paper1

#### RuixiangZhao/OmniPro Viewer• Updated3 days ago • 2.7k • 977 • 2

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.18577 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Paper page - OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Abstract

Models citing this paper0

Datasets citing this paper1

Spaces citing this paper0

Collections including this paper0

Similar Articles

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

Submit Feedback

Similar Articles

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization