@HuggingPapers: Scaling video pre-training to 120K hours ZimaBlue frames video scaling as a route to generalizable World Action Models.…
Summary
Scaling video pre-training to 120K hours boosts zero-shot success in World Action Models from 36.1% to 77.8% on real robots, enabling faster action prediction.
View Cached Full Text
Cached at: 09/03/26, 08:14 PM
Scaling video pre-training to 120K hours
ZimaBlue frames video scaling as a route to generalizable World Action Models. Zero-shot success jumps from 36.1% to 77.8% on real robots, enabling 30 Hz action prediction. https://t.co/qoauPCL21J
Similar Articles
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
ZimaBlue introduces a scalable framework for learning generalizable world action models from large-scale egocentric video, substantially improving zero-shot robotic manipulation through a three-stage curriculum and slow-fast architecture.
Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models (29 minute read)
Dyna-2 is a world-action model pre-trained on over a million hours of human video, showing scaling laws on human data and a human-to-robot transfer scaling law for zero-shot robot performance.
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi introduces Xiaomi-Robotics-1, a vision-language-action foundation model trained on over 100,000 hours of real-world manipulation trajectories, demonstrating clear scaling laws and achieving high success rates on real-world tasks with minimal fine-tuning data.
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
This paper finds that egocentric human video, when processed with a filtering and labeling pipeline, can outperform teleoperated real-robot data for pretraining embodied foundation models, achieving lower validation loss and higher success rates on real-robot tasks.
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Zero-WAM is a causal video-action model that enables zero-shot robotic manipulation of unseen tasks by conditioning on in-context human video guidance, with the HumanGen dataset and a future-chunk prediction objective to improve generalization.