WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Hugging Face Daily Papers Papers

Summary

WorldExam is a new hierarchical benchmark for evaluating world models in controllable video generation, spanning visual quality, control adherence, spatial consistency, and world reactivity. Tests on 20 models show that high visual quality and instruction fulfillment do not guarantee inherent reactivity.

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.
Original Article
View Cached Full Text

Cached at: 08/04/26, 05:37 AM

Paper page - WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Source: https://huggingface.co/papers/2608.02603 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Controllablevideogenerationmodelsareincreasinglybeingdevelopedasworldmodels.Accordingly,evaluatingtheminthisroleextendsbeyondtheapparentappearanceofgeneratedvideostotheinherentreactivityoftheworldstheydepict:theabilitytoinferfromthescenestatehowtheworldshouldreactandtogenerateplausibleconsequencesnotexplicitlydescribedintheinput.Yetexistingbenchmarksmainlyassessvisualqualityorexplicitinstructionfulfillmentbycheckingwhetherrequestedactionsandinteractionoutcomesarerealized,leavinginherentreactivityunderexamined.WeintroduceWorldExam,ahierarchicaldiagnosticbenchmarkspanningfourlevels:VisualQuality,ControlAdherence,SpatialConsistency,andWorldReactivity.Itcomprises1,474casesacrosseightdedicatedtasksandsupportsunifiedevaluationofcamera-,action-,andlanguage-drivenmodelparadigms.TheWorldReactivitylevelevaluatesscene-conditionedreactionsandgoal-directedbehaviorsbeyondwhatisexplicitlyspecifiedintheinput.Evaluationof20representativemodelsrevealsaclearcapabilitysplit.Camera-drivenmodelsexcelatcameracontrol,buttheirinterfacesdonotsupportdynamicinteraction;action-drivenmodelscontrolsubjectsmorepreciselybutoftenleavetheworldunresponsive;andlanguage-drivenmodelsperformbetteroninteractionbutfollowcomplexcontrolslessfaithfully.Nomodelcombinesbroadtaskcoveragewithconsistentlystrongperformance,showingthathighvisualqualityandexplicitinstructionfulfillmentdonotguaranteeinherentreactivity.

View arXiv pageView PDFProject pageGitHub6Add to collection

Get this paper in your agent:

hf papers read 2608\.02603

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.02603 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.02603 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.02603 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

How Should World Models Be Evaluated? A Decision-Making-Centric Position

arXiv cs.LG

This paper surveys evaluation methods for world models and argues for a decision-making-centric framework that prioritizes counterfactual reasoning, planning, and policy optimization over visual quality. It introduces an L0–L7 evaluation ladder and a benchmark protocol to align evaluation with claimed utility.

WorldReward: Reward Modeling for Camera-Conditioned World Models

Hugging Face Daily Papers

WorldReward introduces a vision-language reward model for camera-conditioned world models that unifies action-consistency and visual-quality evaluation through chunk decomposition and preference aggregation, outperforming existing methods like GPT-5.5 on benchmarks.

HappyWorld-Bench

Hugging Face Daily Papers

HappyWorld-Bench is a comprehensive benchmark that evaluates the reliability of world models under interaction and modification across video, spatial, and embodied tracks.