Tag
This paper finds that latent world action models fail to generalize and proposes Simple-WAM, which improves generalization performance while maintaining efficiency by simplifying future modeling into a single forward pass.
EVO-WAM is a framework that adapts world action models to unseen robotics tasks by learning from generated video-action trajectories, significantly improving success rates without additional expert demonstrations.
AnyStep-WAM introduces a framework for tunable-budget prediction and adaptive inference in world-action models, reducing denoising steps by significant margins while maintaining task success rates in robotic manipulation.
Rolling-WAM introduces a method to distribute denoising across replanning cycles in world action models for robotic manipulation, achieving a 4.5x speedup in replanning while maintaining competitive performance.
DeltaWAM introduces delta-based world-action models for bimanual manipulation, enhancing efficiency and performance by predicting visual changes and actions. It demonstrates improved success rates and reduced computational overhead.
OpenWAM presents an open, modular framework for world-action model pretraining to systematically identify design principles, with pretrained models showing strong performance in simulation and real-world robotics.
Scaling video pre-training to 120K hours boosts zero-shot success in World Action Models from 36.1% to 77.8% on real robots, enabling faster action prediction.
ZimaBlue introduces a scalable framework for learning generalizable world action models from large-scale egocentric video, substantially improving zero-shot robotic manipulation through a three-stage curriculum and slow-fast architecture.
LAWA is a world action model that uses latent actions to enable efficient future imagination for robot control, achieving state-of-the-art performance with reduced inference latency. It improves over baselines in generalization and efficiency without generating future observations.
RISE introduces an adaptive framework for imagination rollouts in world action models, using a counterfactual driving dataset to improve planning performance while reducing unnecessary computation.
BadWAM introduces a framework for adversarial attacks on World-Action Models (WAMs), breaking the alignment between imagination and action via small visual perturbations. The attacks significantly reduce task success rates, exposing a vulnerability in this class of models.
GigaWorld-Policy-0.5 is an enhanced World Action Model for robot control that improves training and inference efficiency through a Mixed Action-Conditioned World Modeling strategy and a Mixture-of-Transformers architecture, achieving 85ms latency on a local RTX 4090.
Embodied.cpp is a portable C++ inference runtime that enables efficient deployment of vision-language-action and world-action models across heterogeneous edge devices and robots through modular execution layers and optimized inference.
A survey paper on World Action Models, covering recent advances in AI action and world models.
This survey provides a comprehensive overview of World Action Models (WAMs), predictive-action systems that generate future states for decision-making, and organizes existing works by their required outputs and design choices.
ImageWAM proposes replacing video generation with pretrained image editing models in world action models for robot control, achieving superior performance while reducing FLOPs to 1/6 and latency to 1/4 of video-based approaches.
Light-WAM is a lightweight world action model for efficient robot manipulation that uses a compact video backbone and downsampled latent space for future-video supervision, achieving high performance with low inference latency.
Flash-WAM introduces a modality-aware distillation method for world-action models, achieving real-time inference by compressing diffusion to a single step per modality, resulting in 23x speedup.
Curated GitHub list of Vision-Language-Action and World Action Models research for robotics foundation models.
In his talk at Sequoia AI Ascent, Dr. Jim Fan presents a roadmap for achieving Physical AGI parallel to LLM success, introducing concepts like video world models, World Action Models (WAM), and the Dexterity Scaling Law, and sharing predictions for the near future.