Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Hugging Face Daily Papers Papers

Summary

Introduces Sci-VBench, a benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains, finding that visual realism has not translated into reliable scientific and causal correctness.

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
Original Article
View Cached Full Text

Cached at: 08/11/26, 06:20 AM

Paper page - Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Source: https://huggingface.co/papers/2608.09873

Abstract

WeintroduceSci-VBench,acomprehensivebenchmarkforevaluatingknowledge-andreasoning-intensivevideogenerationacrossscientificdomains.Itcontains1,253expert-annotatedexamplesspanning60subjectsacrossfourcoredisciplines:NaturalScience,Healthcare,Humanities&SocialSciences,andEngineering.Eachexamplerequiresmodelstogeneratetemporallyrichvideosthatdemandscientificreasoningandknowledge-groundedsynthesis,goingbeyondsurface-levelvisualplausibility.Wefurtherestablisharubric-basedevaluationprotocol.Ouranalysisshowsthat,underthisprotocol,bothnon-experthumanevaluatorsandMLLM-as-Judgesystemscanachieverelativelyhighagreementwithexpertjudgments,supportingreproducibleevaluationatscale.Webenchmark16frontierproprietaryandopen-sourcemodelsandfindthat,whileautomaticperceptual-qualityscoresclustertightlyacrosssystems,performanceonPromptGroundingandScientificandCausalCorrectnessvariessubstantially,withapronouncedproprietary-open-sourcegap.Thesefindingsshowthatadvancesinvisualrealismhavenotyettranslatedintoreliablemodelingofscientificandcausaldynamics.

View arXiv pageView PDFGitHub3Add to collection

Get this paper in your agent:

hf papers read 2608\.09873

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.09873 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.09873 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.09873 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

Hugging Face Daily Papers

VideoKR introduces a large-scale video reasoning dataset and benchmark designed to enhance knowledge-intensive video understanding through expert-domain content and human-in-the-loop example generation. The dataset contains 315K video reasoning examples over 145K expert-domain videos.

A Very Big Video Reasoning Suite

Papers with Code Trending

This paper introduces the Very Big Video Reasoning (VBVR) dataset and benchmark, a large-scale resource with over one million video clips across 200 reasoning tasks, enabling systematic study of spatiotemporal reasoning and showing early signs of emergent generalization.

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

Hugging Face Daily Papers

Introduces SVI-Bench, a large-scale benchmark for strategic video intelligence using team sports, designed to evaluate models on dynamic scene understanding, causal reasoning, strategic simulation, and agentic synthesis. The benchmark reveals a capability cliff where models perform well on perceptual tasks but sharply degrade on higher-level strategic reasoning.