OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Summary
The paper introduces OmniVBench, a comprehensive benchmark for omni reference-to-video generation, and the Omni-R2VDataset, a large-scale training dataset, to evaluate and improve R2V models.
View Cached Full Text
Cached at: 09/21/26, 03:20 AM
Paper page - OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Source: https://huggingface.co/papers/2609.22069 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Reference-to-video(R2V)generationisevolvingtowardincreasinglygeneralandversatilereferencecontrol,givingrisetotheemergingparadigmofomniR2Vgeneration.However,existingbenchmarksfallshortoftheseemergingcapabilities:theirtestcasescoverlimitedreferencetypesandcompositions,andtheirevaluationprotocolslargelyassessholisticreferenceconsistency,overlookingwhetherreferencefactorsareproperlypreserved,disentangled,androuted.Meanwhile,thehighcostofconstructingomniR2Vtrainingdatamakessuitabletrainingresourcesscarce.Toaddressthesegaps,weintroduceOmniVBenchandtheOmni-R2VDatasetforevaluatingandtrainingomniR2Vmodels.OmniVBenchexpandsR2Vevaluationacrossbroaderreferencetypes,fine-grainedcontroltasks,andricherreferencecompositions,covering7taskfamiliesand18fine-grainedtasksspanningcontent,motion,style,structure,narrative,andmulti-referencesettings.Weintroducefactor-groundedevaluationwith12,172case-specificchecklistitems,assessingwhetherintendedreferencefactorsarefaithfullypreserved,correctlydisentangledandboundtotheirtargets,andproperlyrealizedaccordingtotheinstruction.WefurtherintroducetheOmni-R2VDataset,bringingindustrial-gradetrainingresourcesfordiverseR2Vtaskstothebroaderresearchcommunity.Drawingprimarilyonalarge-scalecorpusofprofessionalvideofootage,itcomprises340Kprocessedtrainingsamplesspanningdiversereferencetypesandmulti-referencecompositions.Wedeveloptask-specificpipelinesforreference-targetpairconstruction,offeringapracticalandscalablerecipeforomniR2Vdataconstruction.Extensiveevaluationofadvancedopen-andclosed-sourceR2VmodelsrevealsclearperformancegapsacrosstaskfamiliesandevaluationdimensionsonOmniVBench,highlightingremaininglimitationsofcurrentR2Vmodels.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.22069
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.22069 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.22069 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.22069 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
OmniPro is the first benchmark for evaluating proactive streaming video understanding in omni-modal large language models, featuring 2,700 samples covering diverse tasks and dual-mode evaluation protocols.
OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains
OmniVideo-100K introduces an automated data engine with entity-anchored scripting and clue-guided QA generation to improve audio-visual reasoning and temporal consistency, achieving significant performance gains across multiple benchmarks.
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Introduces VIABench, a comprehensive video benchmark for evaluating multimodal large language models in real-world visual assistance for blind and visually impaired individuals, covering 761 videos and 14,526 annotations across three tasks.
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
OmniAssistBench is a benchmark for evaluating omni-modal large language models as real-time video assistants, revealing that current models struggle with visual prompts, context retention, and timely responses.
A Very Big Video Reasoning Suite
This paper introduces the Very Big Video Reasoning (VBVR) dataset and benchmark, a large-scale resource with over one million video clips across 200 reasoning tasks, enabling systematic study of spatiotemporal reasoning and showing early signs of emergent generalization.