Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Summary
Introduces Latent Dynamics Reasoning (LDR), a video world model that integrates kinematic dynamics in a structured latent space, enabling extrapolation of learned dynamics far beyond training distributions while using far fewer parameters and running much faster than video diffusion baselines.
View Cached Full Text
Cached at: 08/13/26, 07:24 PM
Paper page - Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Source: https://huggingface.co/papers/2608.09926
Abstract
Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference.
The world evolves following its dynamics, i.e., its laws of motion. However, leadingvideo diffusionmodels largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduceLatent Dynamics Reasoning(LDR). LDR casts the latent transition as an explicitkinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on astructured latentrather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-boxphysics benchmarkspanning five tasks (uniform motion, parabola, collision, bouncing, looming), focusing on out-of-distribution scenarios that reveal whether a model has truly learned the underlying dynamics. LDR extrapolates the learned dynamics far better: the gap between its in- and out-of-distribution error is over 20times smaller than thevideo diffusionbaseline’s, under both single- and joint-task training at 256^2 resolution, while using 26times fewer parameters and running 143times faster. LDR can even generalize under severe shift: for example, trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left. To our knowledge, this is the first videoworld modelthat extrapolates learned dynamics beyond its training distribution. Project page: https://lat-dyn-reason.github.io/
View arXiv pageView PDFProject pageGitHub28Add to collection
Get this paper in your agent:
hf papers read 2608\.09926
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### haodongli/LDR Updatedabout 1 hour ago • 8
Datasets citing this paper1
#### haodongli/LDR Updatedabout 1 hour ago • 115 • 7
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Beyond Pixels: From Video Priors to 4D Worlds
This paper introduces Latent-to-4D, a method for direct 4D scene generation from video diffusion latents without retraining across generators, achieving better geometry and temporal stability than cascaded approaches.
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
This paper proposes a Latent World Model (LWM) for robot navigation that predicts action-conditioned latent feature compatibility, enabling policy learning from unlabeled video data and reinforcement learning without additional environment interaction, outperforming existing methods.
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
This paper argues that large language models struggle with causal reasoning and long-horizon planning due to a mismatch between sequence prediction and reasoning over latent environment dynamics, and introduces the Latent Dynamics Inference perspective along with the Flux environment to study these limitations.
Latent Spatial Memory for Video World Models
This paper introduces latent spatial memory for video world models, storing 3D scene information directly in diffusion latent space to avoid costly pixel-space reconstruction. The proposed Mirage framework achieves up to 10.57x faster generation and 55x memory reduction while achieving state-of-the-art performance on WorldScore and RealEstate10K.
Learning Visual Feature-Based World Models via Residual Latent Action
This paper introduces RLA-WM, a visual feature-based world model that leverages residual latent actions and flow matching to efficiently predict future visual states. The method outperforms existing video-diffusion and feature-based approaches while enabling novel robot learning techniques from offline, actionless demonstration videos.