WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Summary
WorldExam is a new hierarchical benchmark for evaluating world models in controllable video generation, spanning visual quality, control adherence, spatial consistency, and world reactivity. Tests on 20 models show that high visual quality and instruction fulfillment do not guarantee inherent reactivity.
View Cached Full Text
Cached at: 08/04/26, 05:37 AM
Paper page - WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Source: https://huggingface.co/papers/2608.02603 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Controllablevideogenerationmodelsareincreasinglybeingdevelopedasworldmodels.Accordingly,evaluatingtheminthisroleextendsbeyondtheapparentappearanceofgeneratedvideostotheinherentreactivityoftheworldstheydepict:theabilitytoinferfromthescenestatehowtheworldshouldreactandtogenerateplausibleconsequencesnotexplicitlydescribedintheinput.Yetexistingbenchmarksmainlyassessvisualqualityorexplicitinstructionfulfillmentbycheckingwhetherrequestedactionsandinteractionoutcomesarerealized,leavinginherentreactivityunderexamined.WeintroduceWorldExam,ahierarchicaldiagnosticbenchmarkspanningfourlevels:VisualQuality,ControlAdherence,SpatialConsistency,andWorldReactivity.Itcomprises1,474casesacrosseightdedicatedtasksandsupportsunifiedevaluationofcamera-,action-,andlanguage-drivenmodelparadigms.TheWorldReactivitylevelevaluatesscene-conditionedreactionsandgoal-directedbehaviorsbeyondwhatisexplicitlyspecifiedintheinput.Evaluationof20representativemodelsrevealsaclearcapabilitysplit.Camera-drivenmodelsexcelatcameracontrol,buttheirinterfacesdonotsupportdynamicinteraction;action-drivenmodelscontrolsubjectsmorepreciselybutoftenleavetheworldunresponsive;andlanguage-drivenmodelsperformbetteroninteractionbutfollowcomplexcontrolslessfaithfully.Nomodelcombinesbroadtaskcoveragewithconsistentlystrongperformance,showingthathighvisualqualityandexplicitinstructionfulfillmentdonotguaranteeinherentreactivity.
View arXiv pageView PDFProject pageGitHub6Add to collection
Get this paper in your agent:
hf papers read 2608\.02603
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.02603 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.02603 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.02603 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
WBench is a comprehensive multi-turn benchmark for evaluating interactive world models across five dimensions using 289 test cases and 1,058 interaction turns, providing automatic sub-metrics and diagnostic insights. It reveals that no single model excels across all dimensions.
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
This paper introduces WorldReasonBench and WorldRewardBench, new benchmarks designed to evaluate video generation models' ability to reason about world-state evolution and physical consistency. The research highlights a gap between visual plausibility and true logical reasoning in current commercial video generators.
How Should World Models Be Evaluated? A Decision-Making-Centric Position
This paper surveys evaluation methods for world models and argues for a decision-making-centric framework that prioritizes counterfactual reasoning, planning, and policy optimization over visual quality. It introduces an L0–L7 evaluation ladder and a benchmark protocol to align evaluation with claimed utility.
WorldReward: Reward Modeling for Camera-Conditioned World Models
WorldReward introduces a vision-language reward model for camera-conditioned world models that unifies action-consistency and visual-quality evaluation through chunk decomposition and preference aggregation, outperforming existing methods like GPT-5.5 on benchmarks.
HappyWorld-Bench
HappyWorld-Bench is a comprehensive benchmark that evaluates the reliability of world models under interaction and modification across video, spatial, and embodied tracks.