Tag
Dyna-2 is a world-action model pre-trained on over a million hours of human video, showing scaling laws on human data and a human-to-robot transfer scaling law for zero-shot robot performance.
Pantograph introduces Pan, a 4B parameter goal-conditioned model trained on internet video, capable of performing diverse tasks in Minecraft such as fighting mobs, building structures, and exploring. The method uses hindsight relabeling to learn goal-directed behavior during pretraining.
LingBot-Video presents a DiT-based video pretraining framework with Mixture-of-Experts architecture, specialized data augmentation, and multi-dimensional reward system for embodied intelligence applications.
This paper introduces Orca, a world foundation model that learns a unified latent space from multimodal data using next-state-prediction, outperforming specialized baselines on downstream tasks like text generation, image prediction, and embodied action generation.
OpenAI introduced Video PreTraining (VPT), a semi-supervised method that trains neural networks to play Minecraft by learning from 70,000 hours of unlabeled human gameplay video combined with a small labeled dataset. The model learns complex sequential tasks using the native human interface (keyboard and mouse) and demonstrates capabilities like crafting diamond tools and pillar jumping, representing progress toward general computer-using agents.