Tag
This paper introduces GE-Act 2.0, a world-action model pretrained from scratch to enable scalable zero-shot robotic manipulation with improved success rates across diverse tasks and conditions.
GameWAM introduces the first world-action model for native closed-loop gameplay and GUI control in video games, jointly generating visual observations and executable actions with competitive task success and revealing a source-sensitivity failure mode.
GameWAM introduces the first unified world-action model for native video-game control, jointly predicting future visuals and executable actions using block-causal flow matching and mode-specific distributions.
The paper introduces DECOWAM, a decoupled whole-body world-action model for legged mobile manipulation that improves video and action prediction performance over existing models like FastWAM through dedicated conditional interfaces and a new dataset.
Dyna Robotics introduces Dyna-2, a world-action model pre-trained on one million hours of human video, revealing scaling laws and suggesting the embodiment gap is an adaptation problem rather than a knowledge problem.
Dyna-2 is a world-action model pre-trained on over a million hours of human video, showing scaling laws on human data and a human-to-robot transfer scaling law for zero-shot robot performance.
Dyna Robotics introduces Dyna-2, a world-action model pre-trained on one million hours of human video, discovering new scaling laws for robot manipulation.
SimWAM is a simple yet effective World Action Model for end-to-end autonomous driving that uses video generation purely as a training signal, achieving state-of-the-art 91.5 PDMS on NAVSIM while reducing inference latency.
ST-WAM proposes a semantic-temporal world action model that uses DINOv3 features as a shared semantic representation to improve robot manipulation robustness under visual distribution shifts, achieving 98.7% on LIBERO and 92.8% on RoboTwin 2.0, with significant gains over Fast-WAM in zero-shot settings.
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.
World Pilot enhances Vision-Language-Action models by incorporating dynamic scene evolution and trajectory priors from a World-Action Model, achieving state-of-the-art zero-shot performance on manipulation tasks.
AHA-WAM is an asynchronous world-action model that uses dual Diffusion Transformers to decouple world prediction from action execution, achieving efficient long-horizon planning and real-time control. It achieves state-of-the-art performance on robotic manipulation tasks with up to 92.8% success on RoboTwin and 78.3% on real-world tasks, while reaching 24.17 Hz closed-loop control.
WALL-WM advances video-action learning by using semantic events as learning units instead of fixed action chunks, enabling more flexible and scalable vision-language-action training and inference.
Introduces two projects related to robot world models: Awesome-WAM (OpenMOSS) includes papers such as World Action Models and DreamDojo; awesome-physical-ai curates a collection of papers on VLA models, world models, and embodied foundation models (including NVIDIA Cosmos Predict2.5).
NVIDIA's head of robotics, Jim Fan, gave a public talk, advocating that robots should directly replicate the successful path of large language models. He proposed directions such as World Action Model (WAM), a data revolution based on human first-person video, and neural simulation, and predicted a 95% probability of achieving the endgame of general-purpose physical robots by 2040.