Tag
Introduces Quantum-Structured World Models (QSWMs), a quantum-inspired framework for predictive world modeling with structured latent states, and evaluates them on elementary cellular automata against classical baselines.
SJEPA introduces a reconstruction-free JEPA framework that learns hybrid symbolic-neural latent dynamics, aiming for the simplest adequate predictive representation. Experiments show it discovers simpler symbolic dynamics with lower rollout error than post-hoc fitting, while controlling symbolic-neural allocation under grammar misspecification.
This paper introduces MASS, a method for multiplayer world models that disentangles world dynamics from view rendering using an authoritative shared state, enabling scalable and consistent multi-agent simulation with up to 1,024 concurrent players.
A tweet commenting that Google needs to catch up on frontier coding to stay competitive, while Demis Hassabis focuses on fundamental research like world models for long-term goals.
Wayve announces GAIA-4, a multimodal world model powering closed-loop simulation for safety-critical evaluation of end-to-end autonomous driving models, enabling counterfactual replay of cyclist and pedestrian interactions.
This preprint introduces hierarchical self-supervised world models for music co-creation agents, with fast CPU-friendly models and a live demo for MIDI inpainting and generation.
WorldCycle proposes a self-verifiable reinforcement learning method for long-horizon video world models, using reversible action cycles as free supervision to reduce state-returning drift by up to 44% and boost composite-action accuracy nearly 4x. It also introduces CycleBench to evaluate world models as simulators.
WorldExam is a new hierarchical benchmark for evaluating world models in controllable video generation, spanning visual quality, control adherence, spatial consistency, and world reactivity. Tests on 20 models show that high visual quality and instruction fulfillment do not guarantee inherent reactivity.
Introduces DENSEWORLD, a 1,000-hour dataset of crowded Global South urban scenes, and FactorJEPA, a JEPA variant that factorizes future prediction into layout, agents, and interactions, improving accuracy and robustness under occlusion and heterogeneity.
This paper proposes SG-WAM, a self-guided framework for learning geometry-aware action-conditioned world models directly in policy-derived representation space. It achieves state-of-the-art success rates on LIBERO and LIBERO-Plus benchmarks, outperforming strong baselines in real-world evaluations.
CG-World is a large-scale world-state dataset and protocol derived from industrial computer graphics pipelines, explicitly recording multimodal world states, interventions, and counterfactual branches to support world model research. It demonstrates improvements in geometry-conditioned video generation, action prediction, and closed-loop transfer of vision-language-action policies.
This paper presents a method for learning implicit causal world models from multi-agent demonstrations, enabling agents to infer causal structures from observed behavior.
This paper proposes QQWorld, a quantile-quantile matching objective that replaces the Epps-Pulley objective in LeWorldModel for better regularization of latent distributions, improving planning success in control environments.
NVIDIA released cosmos-framework, an end-to-end open-source framework for training and serving world models including the Cosmos3 model family, supporting distributed training and inference with multiple backends.
Proposes temporal-distance JEPA (TD-JEPA) which mines directed temporal cost from offline trajectories to improve latent world model predictive control, achieving higher success rates on robotic environments.
VisualPatchWorld introduces a method for learning world dynamics as code, enabling inspectable and editable simulators from data. It achieves strong planning success in navigation and manipulation tasks.
INTACT is an end-to-end unified JEPA that learns the intent-to-action mapping directly, enabling search-free world model control. It achieves 95.33% direct macro success rate across four visual-control tasks with zero test-time search and ~300x lower planning latency.
NVIDIA introduces Cosmos-H-Dreams, a real-time action-conditioned generative simulator for surgical robotics, distilled from the larger Cosmos-H-Surgical-Simulator. It runs on a single RTX PRO 6000 GPU using the FlashDreams inference library, enabling interactive closed-loop control for training and evaluation.
This paper empirically investigates how multi-horizon latent consistency affects transition geometry in world models, using an expansion proxy on Moving-MNIST, Pendulum, CartPole, and KTH Actions. It finds that soft consistency can push passive video dynamics toward contraction but not action-conditioned domains.
An analysis of Yann LeCun's bet that intelligence starts with world models via JEPA, not language, supported by AMI Labs' $1.03 billion funding. The article explains why next-pixel prediction fails and how JEPA predicts in latent space to avoid blurry futures.